Model training method and device, information determination method and device, equipment and storage medium
By building and training the task generation and matching model, the inaccurate data prediction of campus equipment management and the effectiveness of personnel scheduling management are solved, and the intelligent management of the park is realized.
Patent Information
- Application Number
- CN202410102675.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, there are problems in the management of park equipment, where data prediction is inaccurate and personnel scheduling management cannot be effectively implemented.
By building an initial task generation model and an initial task matching model, using the target reinforcement learning algorithm for training, the target task generation model and the target task matching model are generated, and the data from the sample park is used for equipment management and personnel scheduling.
The accuracy of equipment management in the park and the effectiveness of personnel scheduling have been achieved, and the intelligence level and management quality of park management have been improved.
Smart Images

Figure CN120374310A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular, to a model training, information determination method, apparatus, device, and storage medium. Background Art
[0002] With the development of technology, property management has changed from traditional manual management to intelligent management. Based on this, the emerging intelligent park can adopt advanced information technology and Internet of Things technology to realize the intelligent management and operation of various facilities and services in the park. An intelligent park is generally a collection of multiple types of buildings and multiple intelligent devices (such as a logistics industrial park, a school, an intelligent community). For the property management of the park, the management of park equipment (inspection of the operating status of various equipment, maintenance of faults, etc.) and personnel scheduling (including fixed building patrol tasks and temporarily assigned tasks) constitute the two main management objects of property management personnel. In traditional park management, manual methods (manual inspections, manual building patrols, manual meter reading, etc.) are generally used for the management of energy-consuming equipment and personnel scheduling; therefore, it is impossible to perform visual processing and statistics on the effectiveness of energy-consuming equipment management and personnel scheduling, and the management method based on experience will make it difficult to effectively optimize the content of property management.
[0003] Based on this, in order to improve the efficiency and accuracy of equipment management, big data and artificial intelligence technologies have begun to be used in related technologies to predict equipment failures. These technologies can predict possible equipment failures by collecting and analyzing the operating data of the equipment, so as to perform maintenance in advance. However, in the implementation process, the inventor found that there are at least the following problems in the prior art: The implementation of the above-mentioned equipment management method requires a large amount of labeled data support, and there will be problems with inaccurate data prediction during cold start; and in personnel scheduling management, purely relying on manual experience or relying on technology (optimization) methods will ignore the problems existing in the real-time environment, resulting in the inability to effectively implement personnel scheduling management. Summary of the Invention
[0004] To solve the above technical problems, the embodiments of this application are expected to provide a model training, information determination method, apparatus, device, and storage medium, which solve the problems of inaccurate data prediction in the equipment management of the park and the inability to effectively implement personnel scheduling management in the park management solution in related technologies.
[0005] The technical solution of this application is implemented as follows:
[0006] A model training method, the method includes:
[0007] Obtain the sample data corresponding to the sample park; wherein, the sample data at least characterizes the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park;
[0008] Construct an initial task generation model, and set first state information, first action information, first policy information, and first reward information for the initial task generation model;
[0009] Construct an initial task matching model, and set second state information, second action information, second policy information, and second reward information for the initial task matching model;
[0010] Based on the sample data, train the set initial task generation model using a first target reinforcement learning algorithm to obtain a target task generation model;
[0011] Based on the sample data, train the set initial task matching model using a second target reinforcement learning algorithm to obtain a target task matching model.
[0012] In the above solution, the obtaining of the sample data corresponding to the sample park includes:
[0013] Obtain the initial data of the sample devices in the sample park, the initial task data of the sample park, and the initial personnel information data corresponding to the sample park;
[0014] Perform time correction on the initial data of the sample devices at a target time interval to obtain the corrected data of the sample devices;
[0015] Screen the corrected data of the sample devices, the initial task data, and the initial personnel information data, and perform format processing on the screened data to obtain the data of the sample devices, the task data, and the personnel information data.
[0016] In the above solution, the setting of the first state information, first action information, first policy information, and first reward information for the initial task generation model includes:
[0017] For the initial task generation model, set a state space based on device operation data, personnel location data, current environmental factors, device task data, user feedback data, and device usage intensity; wherein, the first state information includes the state space;
[0018] Set a first action space for the device based on the actions performed by the device; wherein, the first action information includes the first action space;
[0019] Set an adjustment strategy for the action space based on the environmental state; wherein, the first policy information includes the adjustment strategy;
[0020] Set a task reward function based on the generation result of the task; wherein, the first reward information includes the task reward function.
[0021] In the above solution, the method further includes:
[0022] For the initial task generation model, based on the data characteristics of each data in the state space, set a data running state for each data in the state space.
[0023] In the above solution, the setting of the second state information, the second action information, the second policy information and the second reward information for the initial task matching model includes:
[0024] For the initial task matching model, set a global state space and a local state space; wherein, the global state space represents the supply and demand situation of the park; the local state space represents the supply and demand situation of each intelligent agent in the park; the intelligent agent includes tasks and / or personnel; the second state information includes the global state space and the local state space;
[0025] Set a second action space based on the action selection situation of the intelligent agent, and set a continuous delay times threshold for the delayed actions in the second action space; wherein, the second action information includes the second action space;
[0026] Set a matching strategy for tasks and personnel in the intelligent agent based on the environmental state; wherein, the second policy information includes the matching strategy;
[0027] Set individual reward information corresponding to each intelligent agent and global reward information of the overall intelligent agent based on the matching result of the task; wherein, the second reward information includes the individual reward information and the global reward information.
[0028] In the above solution, the setting of the global state space and the local state space includes:
[0029] Set the global state space based on the number of unsolved tasks, the number of idle personnel, the new task arrival rate and the new idle personnel arrival rate;
[0030] Set the local state space based on the position of the intelligent agent, the time for the task to wait for execution, the idle time of the personnel and the matching distance from the matching object.
[0031] In the above solution, the method further includes:
[0032] For the initial task matching model, if there are newly added tasks and personnel, define a new set of agents;
[0033] Update the environmental state based on the new set of agents;
[0034] Correspondingly, the method further includes:
[0035] For the unmatched agents, set an update strategy for the local state space based on the types of the unmatched agents.
[0036] An information determination method, the method includes:
[0037] Obtain data corresponding to the park to be processed; wherein, the data at least characterizes data of devices in the park to be processed, historical task data of the park to be processed, and personnel information data corresponding to the park to be processed;
[0038] Use a target task generation model to process the data corresponding to the park to be processed to obtain tasks to be executed;
[0039] Use the target task matching model to process the data corresponding to the park to be processed and the tasks to be executed to allocate the tasks to be executed;
[0040] Among them, the target task generation model and the target task matching model can be trained through the above model training method.
[0041] A model training device, the device includes:
[0042] A first acquisition unit, configured to acquire sample data corresponding to a sample park; wherein, the sample data at least characterizes data of sample devices in the sample park, task data of the sample park, and personnel information data corresponding to the sample park;
[0043] A first processing unit, configured to construct an initial task generation model, and set first state information, first action information, first policy information, and first reward information for the initial task generation model;
[0044] A second processing unit, configured to construct an initial task matching model, and set second state information, second action information, second policy information, and second reward information for the initial task matching model;
[0045] A training unit, configured to train the set initial task generation model based on the sample data by using a first target reinforcement learning algorithm to obtain a target task generation model;
[0046] The training unit is further configured to train the set initial task matching model based on the sample data by using a second target reinforcement learning algorithm to obtain a target task matching model.
[0047] An information determination device, the device includes:
[0048] A second acquisition unit, configured to acquire data corresponding to a park to be processed; wherein, the data at least characterizes device data of the park to be processed, historical task data of the park to be processed, and personnel information data corresponding to the park to be processed;
[0049] A third processing unit, configured to process the data corresponding to the park to be processed by using a target task generation model to obtain a to-be-executed task;
[0050] The third processing unit is further configured to process the data corresponding to the park to be processed and the to-be-executed task by using the target task matching model to allocate the to-be-executed task;
[0051] Wherein, the target task generation model and the target task matching model can be trained by using the above model training method.
[0052] A model training device, the device includes: a first processor, a first memory, and a first communication bus;
[0053] The first communication bus is used to implement a communication connection between the first processor and the first memory;
[0054] The first processor is configured to execute a model training program in the first memory to implement the steps of the above model training method.
[0055] An information determination device, the device includes: a second processor, a second memory, and a second communication bus;
[0056] The second communication bus is used to implement a communication connection between the second processor and the second memory;
[0057] The second processor is configured to execute an information determination program in the second memory to implement the steps of the above information determination method.
[0058] A computer-readable storage medium, the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the above model training method or information determination method.
[0059] Since sample data that can obtain data of sample devices corresponding to a sample park, task data of the sample park, and personnel information data corresponding to the sample park can be obtained, an initial task generation model is constructed, and first state information, first action information, first policy information, and first reward information are set for the initial task generation model. An initial task matching model is constructed, and second state information, second action information, second policy information, and second reward information are set for the initial task matching model. The set initial task generation model is trained using a first target reinforcement learning algorithm based on the sample data to obtain a target task generation model. The set initial task matching model is trained using a second target reinforcement learning algorithm based on the sample data to obtain a target task matching model. In this way, based on the data of sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park, the initial task generation model and the initial task matching model with state information, action information, policy information, and reward information set can be trained to obtain a target task generation model and a target task matching model. That is, a method of deep learning can be used to construct models for task generation and task matching based on relevant data of the park. Thus, the obtained target task generation model and target task matching model can be used to generate tasks (i.e., device management) and execute tasks within the park, rather than performing device management (i.e., task generation) and manually executing tasks as in the related art. This solves the problems in the park management solution in the related art that there are inaccurate data predictions for the device management of the park and ineffective implementation for personnel scheduling management. While realizing intelligent park management, the management quality is ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0061] Figure 2 It is a schematic flowchart of another model training method provided by an embodiment of the present application;
[0062] Figure 3 It is a schematic diagram of a system corresponding to a model training method provided by an embodiment of the present application;
[0063] Figure 4 It is a schematic flowchart of an information determination method provided by an embodiment of the present application;
[0064] Figure 5 It is a schematic structural diagram of a model training device provided by an embodiment of the present application;
[0065] Figure 6 It is a schematic structural diagram of an information determination device provided by an embodiment of the present application;
[0066] Figure 7 A structural schematic diagram of a model training device provided by an embodiment of the present application;
[0067] Figure 8 A structural schematic diagram of an information determination device provided by an embodiment of the present application. Specific embodiments
[0068] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application.
[0069] It should be understood that the "embodiments of the present application" or "the foregoing embodiments" mentioned throughout the specification mean that specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the appearances of "in the embodiments of the present application" or "in the foregoing embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. In various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.
[0070] Without special instructions, when an electronic device executes any step in the embodiments of the present application, it can be the processor of the electronic device that executes this step. It is also worth noting that the embodiments of the present application do not limit the order of execution of the following steps by the electronic device. In addition, the methods used to process data in different embodiments can be the same method or different methods. It should also be noted that any step in the embodiments of the present application can be independently executed by the electronic device, that is, when the electronic device executes any step in the following embodiments, it does not need to rely on the execution of other steps.
[0071] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0072] An embodiment of the present application provides a model training method. Referring to Figure 1 As shown, the method may include the following steps:
[0073] Step 101, obtain sample data corresponding to a sample park.
[0074] Among them, the sample data at least characterizes the data of sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park.
[0075] In the embodiments of the present application, the sample park may refer to a park with a relatively large number of energy-consuming devices determined for model training; the data of the sample devices may refer to the operation status data generated by all the devices in the sample park during the historical operation process and the monitoring data on their corresponding sensors; the task data may refer to the data corresponding to the historical tasks in the sample park; and the personnel information data may refer to the basic information data of the employees in the sample park.
[0076] Step 102: Construct an initial task generation model, and set first state information, first action information, first policy information, and first reward information for the initial task generation model.
[0077] In the embodiments of the present application, the initial task generation model may be a model initially created and not yet trained for generating tasks; the first state information, first action information, first policy information, and first reward information are the state information, action information, policy information, and reward information set for the initial task generation model in the network model corresponding to the deep learning algorithm. It should be noted that the first state information, first action information, first policy information, and first reward information may be generated based on device-related information, personnel information, and environmental information.
[0078] Step 103: Construct an initial task matching model, and set second state information, second action information, second policy information, and second reward information for the initial task matching model.
[0079] In the embodiments of the present application, the initial task matching model may be a model initially created and not yet trained for task allocation; the second state information, second action information, second policy information, and second reward information are the state information, action information, policy information, and reward information set for the initial task matching model in the network model corresponding to the deep learning algorithm. It should be noted that the second state information, second action information, second policy information, and second reward information may be generated based on the relevant information of the agent.
[0080] Step 104: Based on the sample data, use the first target reinforcement learning algorithm to train the set initial task generation model to obtain a target task generation model.
[0081] In the embodiments of the present application, sample data can be used as input parameters of the initial task generation model. That is, the sample data can be input into the initial task generation model and processed using the first target reinforcement learning algorithm to implement model training of the initial task generation model, thereby obtaining the target task generation model. Among them, the target task generation model is a model used to generate tasks to be processed. It should be noted that the first target reinforcement learning algorithm can refer to an algorithm for deep reinforcement learning; in a feasible implementation, the first target reinforcement learning algorithm can include the Deep Q-Leaning Network (DQN) algorithm or the Actor-Critic algorithm.
[0082] Step 105: Based on the sample data, train the set initial task matching model using the second target reinforcement learning algorithm to obtain the target task matching model.
[0083] In the embodiments of the present application, sample data can be used as input parameters of the initial task matching model. That is, the sample data can be input into the initial task matching model and processed using the second target reinforcement learning algorithm to implement model training of the initial task matching model, thereby obtaining the target task matching model. Among them, the target task matching model is a model used to allocate tasks to be processed. It should be noted that the second target reinforcement learning algorithm can refer to an algorithm for deep reinforcement learning that is different from the first target reinforcement learning algorithm; in a feasible implementation, the second target reinforcement learning algorithm can include the Spatiotemporal MultiDeep Q-Leaning Network (ST-M-DQN) algorithm and the SpatiotemporalMulti Actor-Critic (ST-M-A2C) algorithm.
[0084] The model training method provided by the embodiments of the present application can train the initial task generation model and the initial task matching model set with status information, action information, policy information, and reward information based on the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park, so as to obtain the target task generation model and the target task matching model. That is, a deep learning method can be used to construct a model for task generation and task matching based on the relevant data of the park. Thus, the obtained target task generation model and target task matching model can be used to generate tasks (i.e., device management) and execute tasks within the park, rather than performing device management (i.e., task generation) and manually executing tasks as in the related art. This solves the problems in the park management solution of the related art that there are inaccurate data predictions for the device management of the park and ineffective implementation for personnel scheduling management, and while realizing the intelligent management of the park, it ensures the management quality.
[0085] Based on the foregoing embodiments, an embodiment of the present application provides a model training method. Referring to Figure 2 as shown, the method may include the following steps:
[0086] Step 201, the model training device obtains the initial data of the sample devices in the sample park, the initial task data of the sample park, and the initial personnel information data corresponding to the sample park.
[0087] In the embodiments of the present application, the initial data of the sample devices may refer to the original operating state data of the devices in the sample park initially obtained by the model training device and the monitoring data on the corresponding sensors; the initial task data may refer to the original data corresponding to the historical tasks in the initially obtained sample park; the initial personnel information data may refer to the original basic information data of the employees in the initially obtained sample park.
[0088] Step 202, the model training device performs time correction on the initial data of the sample devices at a target time interval to obtain the corrected data of the sample devices.
[0089] In the embodiments of the present application, the time correction may be performed on the obtained initial data of the sample devices at a preset time interval a; specifically, the number of data that should be available in a duration N can be calculated respectively according to the interval a, and the data volume in units of months is N m / a, the data volume in units of days is N d / a, the data volume in units of hours is N h / a. Then, it is determined whether the data lengths with the same interval are the same. If they are the same, they are matched according to the unified time index. If they are different, the time index is reconstructed according to the position of the missing values, so as to realize the correction of the initial data of the sample devices.
[0090] Step 203: The model training device screens the calibrated data, initial task data, and initial personnel information data of the sample device, and processes the formatted data after screening to obtain the data of the sample device, task data, and personnel information data.
[0091] In the embodiment of the present application, noise reduction and abnormal data removal processing can be performed on the calibrated data, initial task data, and initial personnel information to achieve screening. It should be noted that the model training device will convert the format of the screened data so that the format of the screened data can be converted into a format suitable for deep learning processing that can be recognized by the machine.
[0092] It should be noted that steps 201 to 203 can be implemented by Figure 3 the intelligent park intelligent property management data center shown; the intelligent park intelligent property management data center may include: a data collection module, a data processing module, and a data storage module.
[0093] Step 204: The model training device constructs an initial task generation model.
[0094] Step 205: The model training device sets a state space for the initial task generation model based on device operation data, personnel location data, current environmental factors, device task data, user feedback data, and device usage intensity.
[0095] Among them, the first state information includes the state space.
[0096] In the embodiment of the present application, the model training device can first determine the device operation data, personnel location data, current environmental factors, device task data (i.e., the historical task data of the device), user feedback data, and device usage intensity, and then combine the device operation data, personnel location data, current environmental factors, device task data, user feedback data, and device usage intensity to obtain the state space. Among them, the device operation data can be represented by D e indicated, the personnel location data can be represented by P p indicated, the current environmental factors can be represented by W t indicated, the device task data can be represented by H t indicated, the user feedback data can be represented by Ge, and the device usage intensity can be represented by Q e indicated, then the state space can be represented by S, S = {D e , P p , W t , H t , G e , Q e}.
[0097] Step 206: The model training device sets the first action space of the device based on the actions performed by the device.
[0098] Among them, the first action information includes the first action space.
[0099] In the embodiments of the present application, the first action space Action can be set according to the actions performed by each sample device history; specifically, the first action space represents all possible actions a that each agent (for each sample device) can perform, where a ∈ Action; in a feasible implementation manner, the action space Action that the air conditioning system in the park may include = {a1: regular inspection, a2: equipment failure, a3: performance optimization,..., a n : no task generation}.
[0100] Step 207: The model training device sets the adjustment strategy for the action space based on the environmental state.
[0101] Among them, the first policy information includes the adjustment strategy.
[0102] In the embodiments of the present application, the adjustment strategy can be defined according to the environmental state of the device; and, the adjustment strategy is used to represent the action adaptation rule. In a feasible implementation manner, during peak periods, the adjustment strategy can determine that tasks for performance optimization will be added to the action space list.
[0103] Step 208: The model training device sets the task reward function based on the generation result of the task.
[0104] Among them, the first reward information includes the task reward function.
[0105] In the embodiments of the present application, the task reward function represents the reward information corresponding to different results of task generation. In a feasible implementation manner, the reward function can be represented by R(s, a), and the generated reward function can be R(s, a) = {R po , R ne , R mi , R no}; where R po indicates that the agent correctly generates the necessary tasks and the task types are accurate; R ne indicates the negative reward for missed detection (not generating the actually required tasks); R mi indicates the negative reward for false alarm (generating unnecessary tasks); the negative reward for type error (wrong task type selection).
[0106] Step 209: The model training device generates a model for the initial task, and sets the data running state for each data in the state space based on the data characteristics of each data in the state space.
[0107] In the embodiments of the present application, according to the data characteristics of data at different time points in the state space, different operating states are labeled for each piece of data in the state space; in a feasible implementation manner, the data operating states may include: peak state, normal use state, and low use state.
[0108] Step 210: The model training device trains the set initial task generation model based on the sample data by using the first target reinforcement learning algorithm to obtain a target task generation model.
[0109] In the embodiments of the present application, obtain the data required for the task generation model from the data center, and construct a task distribution model based on reinforcement learning based on the intelligent property task generation module of the intelligent park; construct a deep neural network, where the input layer receives state variables and the output layer outputs the expected Q value for each action; initialize an experience replay pool for storing the state transitions experienced by the Agent to store (state, action, reward, new state); create a target network, create another neural network with the same structure as the main DQN, which is used to stabilize the learning process; perform DQN training, by observing the current environmental state s, based on the current policy (initially a random policy), select an action a, execute the action a and observe the new state s' and the reward r, store (s, a, r, s') in the experience replay pool, randomly extract a batch of samples from the experience replay pool, use the target network to calculate the Q value of the next state, and update the main network according to the loss function (usually the mean square error); use standardized evaluation metrics, such as accuracy, recall, and F1 score, to evaluate the performance of the model; have the property management personnel correct the decision-making generation results of the model to generate a list of tasks L to be executed d 。
[0110] Step 211: The model training device sets a global state space and a local state space for the initial task matching model.
[0111] Among them, the global state space represents the supply and demand situation of the park; the local state space represents the supply and demand situation of each intelligent agent in the park; the intelligent agent includes tasks and / or personnel; the second state information includes the global state space and the local state space.
[0112] In the embodiments of the present application, the model training device can set the global state space for the task matching model based on the information of the task and the personnel information; and, for the task matching model, set the local state space based on the information of the intelligent agent, the information of the task, and the personnel information. Among them, each task or employee (personnel) request is regarded as an independent intelligent agent. In a feasible implementation manner, if there are 10 unassigned tasks and 5 available employees, there are 15 intelligent agents.
[0113] It should be noted that step 211 can be implemented in the following ways:
[0114] Step 211a: Set the global state space based on the number of unsolved tasks, the number of idle personnel, the new task arrival rate, and the new idle personnel arrival rate.
[0115] In the embodiments of the present application, the state space may include a global state space and a local state space; the state space can be represented by F, and the global state space can be represented by F G denoted; F G ={the number of unsolved tasks, the number of idle personnel, the new task arrival rate, the new idle personnel arrival rate}; among them, the new task arrival rate and the new idle personnel arrival rate can be estimated based on the current situation.
[0116] Step 211b: Set the local state space based on the position of the agent, the time for the task to wait for execution, the idle time of the personnel, and the matching distance from the matching object.
[0117] In the embodiments of the present application, the local state space can be represented by F R denoted; for tasks, F R ={the position of the agent, the time for the task to wait for execution, the matching distance from the matching object}, or, for personnel, F R ={the position of the agent, the idle time of the personnel, the matching distance from the matching object}; among them, the position of the agent can be represented by one-hot encoding, the time for the task to wait for execution can refer to the cumulative waiting time of all tasks, and the matching distance from the matching object can be estimated based on the current situation.
[0118] Step 212: The model training device sets the second action space based on the action selection situation of the agent, and sets a continuous delay times threshold for the delayed actions in the second action space.
[0119] Among them, the second action information includes the second action space.
[0120] In the embodiments of the present application, each agent has two action selections: joining the matching pool or delaying the decision; that is, the second action space includes joining the matching pool and delaying the decision. Among them, joining the matching pool means that the agent chooses to immediately participate in the current task matching process; delaying the decision means that the agent chooses to wait, does not participate in the current round of matching, but delays to the next matching time interval. For task agents, the agent's choice to participate in the matching means being ready to accept the assigned task; for personnel agents, the agent's choice to participate in the matching means being ready to execute the assigned task; the choice to delay may be based on various reasons, such as expecting a more suitable matching opportunity, insufficient current resources, etc.
[0121] In other embodiments of the present application, for task agents (such as maintenance tasks, inspection tasks), the second action space includes a rw = 1 and a rw = 0; where a rw = 1 indicates joining the current matching pool, and a rw = 0 indicates delaying until the next matching time interval. For personnel agents, the second action space includes a ry = 1 and a ry = 0; where a ry = 1 indicates willingness to accept the task assignment in the current round, and a ry = 0 indicates choosing to wait and not participating in the matching in the current round.
[0122] It should be noted that, according to the actual situation, a limit on the number of consecutive delays (consecutive delay count threshold) is set to ensure that the task is executed.
[0123] Step 213, the model training device sets the matching strategy between tasks and personnel in the agent based on the environmental state.
[0124] Among them, the second policy information includes the matching strategy.
[0125] In the embodiments of the present application, the matching strategy can represent an optimization algorithm for tasks and personnel. For example, the optimization algorithm Match(E t , A t ), where E t is the current environmental state, and A t is the set of actions of the agent at time t. The matching strategy can be used to match tasks and personnel to obtain the matching result M t+1 .
[0126] Step 214, the model training device sets the individual reward information corresponding to each agent and the global reward information of the overall agent based on the matching result of the task.
[0127] Among them, the second reward information includes individual reward information and global reward information.
[0128] In the embodiments of the present application, the individual reward information refers to defining a positive reward value V to reward each successfully matched agent. In a feasible implementation, when a task is successfully matched with a suitable person, the agent obtains V points, and a negative reward value -C is defined to punish agents that are not successfully matched or choose to delay; if the task fails to be matched within the current time interval or the person does not accept the task, the agent loses C points.
[0129] In the embodiments of the present application, the global reward information is the average of the rewards of all agents. The global reward information is to encourage cooperation among agents to improve the overall efficiency. If there are N agents, the global reward can be expressed as where R i =(a i , s i ) represents the individual reward of the i-th agent based on the action a i and the state s i .
[0130] It should be noted that at the end of each time interval, the reward of the agent is updated according to the agent's behavior and matching result. After each time interval ends, the average reward of all agents is calculated and distributed to each agent.
[0131] Step 215: For the initial task matching model, if there are newly added tasks and personnel, the model training device defines a new agent set.
[0132] In the embodiments of the present application, the new agent set can be defined as N t+1 ; where the new agent set includes the newly added tasks and personnel.
[0133] Step 216: The model training device updates the environmental state based on the new agent set.
[0134] In the embodiments of the present application, the model training device can define a time variable t to represent the current matching time interval, and t + 1 to represent the next matching time interval. Set the environmental state E t at time. When the time enters t + 1, it is updated to E t+1 according to the actions and interaction results of the agents. And, the new agent set is integrated into the environmental state and updated to E t+1 = E t+1 ∪N t+1 .
[0135] Step 217: For the unmatched agents, the model training device sets an update strategy for the local state space based on the types of the unmatched agents.
[0136] In the embodiments of the present application, the unmatched agents U t can be identified from E t+1 and M t+1 . For each unmatched agent u ∈ U t+1 , its state is updated according to its type (task or personnel).
[0137] It should be noted that the corresponding update strategies are different for different types of agents. Specifically, the state update strategy for task agents is as follows: If a rw= 0, then the waiting time w of the task agent rw increases, and the updated policy is expressed as If a rw = 1, the state of the task agent is updated according to the matching result. The state update policy of the personnel agent is as follows: If a ry = 0, then the position state l of the personnel agent ry changes, and the updated policy is expressed as: If a ry = 1, then the state of the personnel agent is updated to the task execution state.
[0138] Step 218, the model training device trains the set initial task matching model based on the sample data by using the second target reinforcement learning algorithm to obtain the target task matching model.
[0139] In the embodiment of the present application, the data required by the task generation model is obtained from the data center, and a multi-agent task matching model is constructed based on the intelligent park intelligent property task matching and distribution module; the global state and the local state are defined, and the feature engineering technology is applied to transform and optimize the state representation; a training framework is constructed and a simulation environment is built, and the model parameters are initialized to initialize the policy for each agent i; at each matching time interval t, the following training steps are executed: Observe the current state s t , select actions and update the policy according to ST-M-A2C or ST-M-DQN; observe the new state s t+1 and the reward r t , update the model parameters; monitor key performance indicators, such as the successful matching rate, waiting time, etc., to evaluate the effectiveness of the policy; repeat the interaction and policy update process until the policy converges or reaches the predetermined performance indicators; finally, generate the target task matching model.
[0140] It should be noted that steps 210 and 218 can be executed by Figure 3 the intelligent park intelligent property task adaptive learning module shown as Figure 3 shown, the intelligent park intelligent property task adaptive learning module realizes the following functions: task data analysis and visualization, model training and optimization, operation feedback and task adjustment, and knowledge base and model update.
[0141] It should be noted that steps 204 to 210 and steps 211 to 218 can be executed simultaneously; or, steps 211 to 218 are executed after steps 204 to 210 are executed; or, steps 204 to 210 are executed after steps 211 to 218 are executed. Among them, Figure 2 the case where steps 211 to 218 are executed after steps 204 to 210 are executed is taken as an example for illustration.
[0142] In other embodiments of the present application, such as Figure 3 shown, the intelligent property management system for an intelligent park further includes an intelligent park intelligent property task generation module and an intelligent park intelligent property task matching and distribution module; among them, the intelligent park intelligent property task generation module realizes a data-driven operation status anomaly detection function and a human-machine collaborative intelligent property task generation function; the intelligent park intelligent property task matching and distribution module realizes the task matching and distribution function.
[0143] It should be noted that the present application is not only based on basic data, but also integrates various information (such as equipment data, personnel location, environment, etc.) to describe the real-time status of the park, enhancing the real-time performance and accuracy of the model. The present application can generate different task types according to the real-time status of the park, such as cleaning, security inspections, etc., thus realizing the dynamic management of park tasks. In the task distribution stage, the present application performs intelligent matching according to the attributes of the tasks and the status of the staff to ensure that the tasks can be completed in a timely and correct manner, improving the efficiency of park management. The present application is oriented to the field of intelligent park intelligent property management, can realize the full life cycle management from the most concerned task generation to distribution, explores a new model of human-machine collaborative property management, and can better promote the intelligent and smart management of the park.
[0144] It should be noted that the descriptions of the same steps and the same content in this embodiment and other embodiments can be referred to the descriptions in other embodiments, and will not be repeated here.
[0145] The model training method provided by the embodiments of the present application can train an initial task generation model and an initial task matching model set with status information, action information, policy information, and reward information according to the sample device data of the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park, so as to obtain a target task generation model and a target task matching model, that is, a deep learning method can be used to construct models for task generation and task matching based on the relevant data of the park, so that the obtained target task generation model and target task matching model can be used to generate tasks (i.e., equipment management) and execute tasks in the park, rather than performing equipment management (i.e., task generation) and manually executing tasks as in the related art, solving the problems in the park management solution of the related art that there are inaccurate data predictions for the equipment management of the park and ineffective implementation for the personnel scheduling management, and ensuring the management quality while realizing the intelligent management of the park.
[0146] Based on the foregoing embodiments, the embodiments of the present application provide an information determination method. Referring to Figure 4 shown, the method may include the following steps:
[0147] Step 301: The information determination device obtains data corresponding to the park to be processed.
[0148] The data at least represents the equipment data of the park to be processed, the historical task data of the park to be processed, and the personnel information data corresponding to the park to be processed.
[0149] In an embodiment of the present application, the park to be processed may refer to a park determined for intelligent property management; the number of equipment in the park to be processed may refer to the operating status data generated by all equipment in the park to be processed during the historical operation process and the monitoring data on its corresponding sensors; the historical task data of the park to be processed may refer to the data corresponding to the historical tasks in the park to be processed; and the personnel information data may refer to the basic information data of the employees in the park to be processed.
[0150] Step 302: The information determination device uses the target task generation model to process the data corresponding to the park to be processed to obtain the tasks to be executed.
[0151] In an embodiment of the present application, the data corresponding to the park to be processed can be used as input parameters, that is, the data corresponding to the park to be processed is input into the target task generation model. After the target task generation model processes the data corresponding to the park to be processed, it can output the tasks that need to be performed in the park to be processed.
[0152] Step 303: Use the target task matching model to process the data corresponding to the park to be processed and the tasks to be executed, so as to allocate the tasks to be executed.
[0153] In an embodiment of the present application, the data corresponding to the park to be processed and the tasks to be executed can be used as input parameters, that is, the data corresponding to the park to be processed and the tasks to be executed are input into the target task matching model. After the target task matching model processes the data corresponding to the park to be processed and the tasks to be executed, it can output the matching results between the tasks to be executed and the personnel in the park to be processed, thereby realizing the allocation of the tasks to be executed.
[0154] The target task generation model and the target task matching model can be trained in the following ways:
[0155] a1. Obtain sample data corresponding to the sample park; wherein the sample data at least represents data of sample equipment in the sample park, task data of the sample park, and personnel information data corresponding to the sample park.
[0156] a2. Construct an initial task generation model, and set first state information, first action information, first strategy information and first reward information for the initial task generation model.
[0157] a3. Construct an initial task generation model, and set second state information, second action information, second policy information, and second reward information for the initial task generation model.
[0158] a4. Based on the sample data, use the first target reinforcement learning algorithm to train the set initial task generation model to obtain a target task generation model.
[0159] a5. Based on the sample data, use the second target reinforcement learning algorithm to train the set initial task matching model to obtain a target task matching model.
[0160] It should be noted that the descriptions of the same steps and the same content in this embodiment and other embodiments can be referred to the descriptions in other embodiments, and will not be repeated here.
[0161] The information determination method provided by the embodiments of the present application can use the target task generation model and the target task matching model obtained by training the initial task generation model and the initial task matching model set with state information, action information, policy information, and reward information according to the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park. Analyze the data in the park to be processed, and can determine the tasks to be executed in the park to be processed, and allocate the tasks to be executed, solving the problems in the park management solution in the related technology that there are inaccurate data predictions for the equipment management in the park and the personnel scheduling management cannot be effectively implemented. While realizing the intelligent management of the park, the management quality is ensured.
[0162] Based on the foregoing embodiments, an embodiment of the present application provides a model training device, which can be applied to Figures 1 - 2 the model training method provided by the corresponding embodiment, refer to Figure 5 As shown, the model training device 4 may include: a first acquisition unit 41, a first processing unit 42, a second processing unit 43, and a training unit 44, where:
[0163] The first acquisition unit 41 is used to acquire sample data corresponding to the sample park; wherein, the sample data at least represents the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park;
[0164] The first processing unit 42 is used to construct an initial task generation model, and set first state information, first action information, first policy information, and first reward information for the initial task generation model;
[0165] A second processing unit 43, configured to build an initial task matching model, and set second state information, second action information, second policy information, and second reward information for the initial task matching model;
[0166] A training unit 44, configured to train the set initial task generation model based on sample data by using a first target reinforcement learning algorithm to obtain a target task generation model;
[0167] The training unit 44 is further configured to train the set initial task matching model based on sample data by using a second target reinforcement learning algorithm to obtain a target task matching model.
[0168] In other embodiments of the present application, the first acquisition unit 41 is further configured to perform the following steps:
[0169] Acquire initial data of sample devices in a sample park, initial task data of the sample park, and initial personnel information data corresponding to the sample park;
[0170] Perform time correction on the initial data of the sample devices at a target time interval to obtain corrected data of the sample devices;
[0171] Screen the corrected data of the sample devices, the initial task data, and the initial personnel information data, and perform format processing on the screened data to obtain device data, task data, and personnel information data of the sample devices.
[0172] In other embodiments of the present application, the first processing unit 42 is further configured to perform the following steps:
[0173] For the initial task generation model, set a state space based on device operation data, personnel location data, current environmental factors, device task data, user feedback data, and device usage intensity; wherein, the first state information includes the state space;
[0174] Set a first action space for the device based on the actions performed by the device; wherein, the first action information includes the first action space;
[0175] Set an adjustment strategy for the action space based on the environmental state; wherein, the first policy information includes the adjustment strategy;
[0176] Set a task reward function based on the generation result of the task; wherein, the first reward information includes the task reward function.
[0177] In other embodiments of the present application, the first processing unit 42 is further configured to, for the initial task generation model, set a data operation state for each data in the state space based on the data characteristics of each data in the state space.
[0178] In other embodiments of the present application, the second processing unit 43 is further configured to perform the following steps:
[0179] For the initial task matching model, set a global state space and a local state space; wherein, the global state space represents the supply and demand situation of the park; the local state space represents the supply and demand situation of each agent in the park; the agent includes a task and / or a person; the second state information includes the global state space and the local state space;
[0180] Set a second action space based on the action selection situation of the agent, and set a continuous delay count threshold for the delayed actions in the second action space; wherein, the second action information includes the second action space;
[0181] Set a matching strategy for the tasks and persons in the agent based on the environmental state; wherein, the second policy information includes the matching strategy;
[0182] Set the individual reward information corresponding to each agent and the global reward information of the overall agent based on the matching result of the task; wherein, the second reward information includes the individual reward information and the global reward information.
[0183] In other embodiments of the present application, the second processing unit 43 is further configured to perform the following steps:
[0184] Set the global state space based on the number of unresolved tasks, the number of idle persons, the new task arrival rate, and the new idle person arrival rate;
[0185] Set the local state space based on the position of the agent, the time for the task to wait for execution, the idle time of the person, and the matching distance from the matching object.
[0186] In other embodiments of the present application, the second processing unit 43 is further configured to perform the following steps:
[0187] For the initial task matching model, if there are newly added tasks and persons, define a new agent set;
[0188] Update the environmental state based on the new agent set;
[0189] Correspondingly, the second processing unit 43 is further configured to, for the unmatched agents, set an update strategy for the local state space based on the types of the unmatched agents.
[0190] It should be noted that the specific descriptions of the steps executed by each unit can be referred to Figures 1 - 2 In the model training method provided in the corresponding embodiment, details are not described herein again.
[0191] The model training device provided by the embodiments of the present application can train an initial task generation model and an initial task matching model with status information, action information, policy information, and reward information set according to the data of sample devices in a sample park, the task data of the sample park, and the personnel information data corresponding to the sample park, so as to obtain a target task generation model and a target task matching model. That is, a deep learning method can be used to construct models for task generation and task matching based on relevant data of the park. Thus, the obtained target task generation model and target task matching model can be used to generate tasks (i.e., device management) and execute tasks within the park, rather than performing device management (i.e., task generation) and manually executing tasks as in the related art. This solves the problems in the park management solution of the related art that there are inaccurate data predictions for the device management of the park and ineffective implementation for personnel scheduling management, and while realizing the intelligent management of the park, the management quality is ensured.
[0192] Based on the foregoing embodiments, an information determination device is provided in an embodiment of the present application. This information determination device can be applied to Figure 4 the information determination method provided in the corresponding embodiment. Referring to Figure 6 as shown, the information determination device 5 may include: a second acquisition unit 51 and a third processing unit 52, where:
[0193] The second acquisition unit 51 is configured to acquire data corresponding to the park to be processed; where the data at least characterizes the data of the devices in the park to be processed, the historical task data of the park to be processed, and the personnel information data corresponding to the park to be processed;
[0194] The third processing unit 52 is configured to process the data corresponding to the park to be processed by using the target task generation model to obtain a task to be executed;
[0195] The third processing unit 52 is further configured to process the data corresponding to the park to be processed and the task to be executed by using the target task matching model to allocate the task to be executed;
[0196] Among them, the target task generation model and the target task matching model can be trained in the following manner:
[0197] Acquire sample data corresponding to the sample park; where the sample data at least characterizes the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park.
[0198] Construct an initial task generation model, and set first status information, first action information, first policy information, and first reward information for the initial task generation model.
[0199] Construct an initial task generation model, and set second state information, second action information, second policy information, and second reward information for the initial task generation model.
[0200] Based on the sample data, use the first target reinforcement learning algorithm to train the set initial task generation model to obtain a target task generation model.
[0201] Based on the sample data, use the second target reinforcement learning algorithm to train the set initial task matching model to obtain a target task matching model.
[0202] It should be noted that the specific descriptions of the steps executed by each unit can be referred to Figure 4 in the information determination method provided in the corresponding embodiment, which will not be elaborated here.
[0203] The information determination device provided by the embodiment of the present application can use the target task generation model and the target task matching model obtained by training the initial task generation model and the initial task matching model set with state information, action information, policy information, and reward information based on the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park to analyze the data in the park to be processed, and can determine the to-be-executed tasks in the park to be processed and allocate the to-be-executed tasks, solving the problems in the park management solution in the related technology that there are inaccurate data predictions for the equipment management in the park and the personnel scheduling management cannot be effectively implemented, and while realizing the intelligent management of the park, ensuring the management quality.
[0204] Based on the foregoing embodiments, an embodiment of the present application provides a model training device, which can be applied to Figures 1 - 2 the model training method provided in the corresponding embodiment, referring to Figure 7 as shown, the model training device 6 may include: a first processor 61, a first memory 62, and a first communication bus 63, where:
[0205] The first communication bus 63 is used to implement the communication connection between the first processor 61 and the first memory 62;
[0206] The first processor 61 is used to execute the model training program in the first memory 62 to implement the following steps:
[0207] Obtain the sample data corresponding to the sample park; wherein, the sample data at least characterizes the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park;
[0208] Construct an initial task generation model, and set first state information, first action information, first policy information, and first reward information for the initial task generation model;
[0209] Construct an initial task matching model, and set second state information, second action information, second policy information, and second reward information for the initial task matching model;
[0210] Based on the sample data, train the set initial task generation model using a first target reinforcement learning algorithm to obtain a target task generation model;
[0211] Based on the sample data, train the set initial task matching model using a second target reinforcement learning algorithm to obtain a target task matching model.
[0212] In other embodiments of the present application, the first processor 61 is used to execute the model training program in the first memory 62 to obtain sample data corresponding to the sample park, so as to implement the following steps:
[0213] Obtain the initial data of the sample devices in the sample park, the initial task data of the sample park, and the initial personnel information data corresponding to the sample park;
[0214] Perform time correction on the initial data of the sample devices at a target time interval to obtain the corrected data of the sample devices;
[0215] Screen the corrected data of the sample devices, the initial task data, and the initial personnel information data, and perform format processing on the screened data to obtain the data of the sample devices, task data, and personnel information data.
[0216] In other embodiments of the present application, the first processor 61 is used to execute the model training program in the first memory 62 to set first state information, first action information, first policy information, and first reward information for the initial task generation model, so as to implement the following steps:
[0217] For the initial task generation model, based on device operation data, personnel location data, current environmental factors, device task data, user feedback data, and device usage intensity, set a state space; wherein, the first state information includes the state space;
[0218] Set a first action space for the device based on the actions performed by the device; wherein, the first action information includes the first action space;
[0219] Set an adjustment strategy for the action space based on the environmental state; wherein, the first policy information includes the adjustment strategy;
[0220] Set a task reward function based on the generation result of the task; wherein, the first reward information includes the task reward function.
[0221] In other embodiments of the present application, the first processor 61 is used to execute the model training program in the first memory 62, and the following steps can also be implemented:
[0222] Generate a model for the initial task, and set the data running state for each data in the state space based on the data characteristics of each data in the state space.
[0223] In other embodiments of the present application, the first processor 61 is used to execute the model training program in the first memory 62 to set the second state information, second action information, second policy information, and second reward information for the model matching the initial task, so as to implement the following steps:
[0224] For the model matching the initial task, set the global state space and the local state space; wherein, the global state space represents the supply and demand situation in the park; the local state space represents the supply and demand situation of each intelligent agent in the park; the intelligent agent includes tasks and / or personnel; the second state information includes the global state space and the local state space;
[0225] Set the second action space based on the action selection situation of the intelligent agent, and set a continuous delay times threshold for the delayed actions in the second action space; wherein, the second action information includes the second action space;
[0226] Set the matching strategy between the tasks and personnel in the intelligent agent based on the environmental state; wherein, the second policy information includes the matching strategy;
[0227] Set the individual reward information corresponding to each intelligent agent and the global reward information of the overall intelligent agent based on the matching result of the task; wherein, the second reward information includes the individual reward information and the global reward information.
[0228] In other embodiments of the present application, the first processor 61 is used to execute the model training program in the first memory 62 to set the global state space and the local state space, so as to implement the following steps:
[0229] Set the global state space based on the number of unresolved tasks, the number of idle personnel, the new task arrival rate, and the new idle personnel arrival rate;
[0230] Set the local state space based on the position of the intelligent agent, the time for the task to wait for execution, the idle time of the personnel, and the matching distance from the matching object.
[0231] In other embodiments of the present application, the first processor 61 is used to execute the model training program in the first memory 62, and the following steps can also be implemented:
[0232] For the model matching the initial task, if there are newly added tasks and personnel, define a new intelligent agent set;
[0233] Update the environmental state based on the new set of agents;
[0234] Correspondingly, the first processor 61 is used to execute the model training program in the first memory 62, and can also implement the following steps:
[0235] For unmatched agents, set the update strategy for the local state space based on the types of the unmatched agents.
[0236] It should be noted that the specific description of the steps executed by the first processor can refer to Figures 1 - 2 In the model training method provided in the corresponding embodiment, it will not be elaborated here.
[0237] The model training device provided by the embodiments of the present application can train the initial task generation model and the initial task matching model with state information, action information, policy information, and reward information set according to the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park, so as to obtain the target task generation model and the target task matching model. That is, a method of deep learning can be used to construct models for task generation and task matching based on the relevant data of the park, so that the obtained target task generation model and target task matching model can be used to generate tasks (i.e., device management) and execute tasks in the park, rather than managing devices (i.e., task generation) and manually executing tasks as in the related art. This solves the problems in the park management solution of the related art that there are inaccurate data predictions for the device management in the park and ineffective implementation for the personnel scheduling management, and while realizing the intelligent management of the park, the management quality is ensured.
[0238] Based on the foregoing embodiments, the embodiments of the present application provide an information determination device, which can be applied to Figure 4 In the information determination method provided in the corresponding embodiment, refer to Figure 8 As shown, the information determination device 7 may include: a second processor 71, a second memory 72, and a second communication bus 73, where:
[0239] The second communication bus 73 is used to implement the communication connection between the second processor 71 and the second memory 72;
[0240] The second processor 71 is used to execute the information determination program in the second memory 72 to implement the following steps:
[0241] Obtain the data corresponding to the park to be processed; where the data at least represents the data of the devices in the park to be processed, the historical task data of the park to be processed, and the personnel information data corresponding to the park to be processed;
[0242] The target task generation model is used to process the data corresponding to the park to be processed to obtain the tasks to be executed;
[0243] The target task matching model is used to process the data corresponding to the park to be processed and the tasks to be executed, so as to allocate the tasks to be executed;
[0244] The target task generation model and the target task matching model can be trained in the following ways:
[0245] Obtain sample data corresponding to the sample park; wherein the sample data at least represents data of sample equipment in the sample park, task data of the sample park, and personnel information data corresponding to the sample park.
[0246] An initial task generation model is constructed, and first state information, first action information, first strategy information, and first reward information are set for the initial task generation model.
[0247] An initial task matching model is constructed, and second state information, second action information, second strategy information, and second reward information are set for the initial task matching model.
[0248] Based on the sample data, the first target reinforcement learning algorithm is used to train the set initial task generation model to obtain the target task generation model.
[0249] Based on the sample data, the second target reinforcement learning algorithm is used to train the set initial task matching model to obtain the target task matching model.
[0250] It should be noted that the specific description of the steps performed by the second processor can be referred to Figure 4 The information determination method provided in the corresponding embodiment will not be repeated here.
[0251] The information determination device provided in the embodiment of the present application can use the data of sample equipment in the sample park, the task data of the sample park and the personnel information data corresponding to the sample park to train the initial task generation model and the initial task matching model that are set with status information, action information, strategy information and reward information to obtain the target task generation model and the target task matching model, and analyze the data in the park that needs to be processed. It can determine the tasks to be executed in the park that need to be processed and allocate the tasks to be executed, which solves the problems of inaccurate data prediction for equipment management of the park and the inability to effectively implement personnel scheduling management in the park management solutions in the related technologies, and realizes intelligent management of the park while ensuring the management quality.
[0252] Based on the foregoing embodiments, an embodiment of the present application provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement Figure 1 and 2 the model training method provided by the corresponding embodiment, and Figure 4 the steps of the information determination method provided by the corresponding embodiment.
[0253] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0254] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0255] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0256] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0257] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application.
Claims
1. A model training method, characterized in that, The method includes: Obtaining sample data corresponding to a sample park; wherein, the sample data at least characterizes data of sample devices in the sample park, task data of the sample park, and personnel information data corresponding to the sample park; Constructing an initial task generation model, and setting first state information, first action information, first policy information, and first reward information for the initial task generation model; Constructing an initial task matching model, and setting second state information, second action information, second policy information, and second reward information for the initial task matching model; Based on the sample data, training the set initial task generation model using a first target reinforcement learning algorithm to obtain a target task generation model; Based on the sample data, training the set initial task matching model using a second target reinforcement learning algorithm to obtain a target task matching model.
2. The method according to claim 1, characterized in that The obtaining of the sample data corresponding to the sample park includes: Obtaining initial data of sample devices in the sample park, initial task data of the sample park, and initial personnel information data corresponding to the sample park; Performing time correction on the initial data of the sample devices at a target time interval to obtain corrected data of the sample devices; Screening the corrected data of the sample devices, the initial task data, and the initial personnel information data, and performing format processing on the screened data to obtain the data of the sample devices, the task data, and the personnel information data.
3. The method according to claim 1, wherein The setting of the first state information, first action information, first policy information, and first reward information for the initial task generation model includes: For the initial task generation model, setting a state space based on device operation data, personnel location data, current environmental factors, device task data, user feedback data, and device usage intensity; wherein, the first state information includes the state space; Setting a first action space for the device based on the actions performed by the device; wherein, the first action information includes the first action space; Setting an adjustment strategy for the action space based on the environmental state; wherein, the first policy information includes the adjustment strategy; Setting a task reward function based on the generation result of the task; wherein, the first reward information includes the task reward function.
4. The method according to claim 3, wherein The method further includes: For the initial task generation model, setting a data operation state for each data in the state space based on the data characteristics of each data in the state space.
5. The method according to claim 1, wherein The setting of the second state information, second action information, second policy information, and second reward information for the initial task matching model includes: For the initial task matching model, setting a global state space and a local state space; wherein, the global state space characterizes the supply and demand situation in the park; the local state space characterizes the supply and demand situation of each intelligent agent in the park; the intelligent agent includes tasks and / or personnel; the second state information includes the global state space and the local state space; Set a second action space based on the action selection of the agent, and set a continuous delay count threshold for the delayed actions in the second action space; wherein, the second action information includes the second action space. Set the matching strategy between tasks and personnel in the agent based on the environmental state; wherein, the second policy information includes the matching strategy. Set the individual reward information corresponding to each agent and the global reward information of the overall agent based on the matching result of the task; wherein, the second reward information includes the individual reward information and the global reward information.
6. The method according to claim 5, wherein The setting of the global state space and the local state space includes: Set the global state space based on the number of unsolved tasks, the number of idle personnel, the new task arrival rate, and the new idle personnel arrival rate. Set the local state space based on the position of the agent, the time for the task to wait for execution, the idle time of the personnel, and the matching distance from the matching object.
7. The method according to claim 5, characterized in that The method further includes: For the initial task matching model, if there are newly added tasks and personnel, define a new agent set. Update the environmental state based on the new agent set. Correspondingly, the method further includes: For the unmatched agents, set the update strategy for the local state space based on the types of the unmatched agents.
8. An information determination method, characterized in that The method includes: Obtain the data corresponding to the park to be processed; wherein, the data at least characterizes the data of the devices in the park to be processed, the historical task data of the park to be processed, and the personnel information data corresponding to the park to be processed. Process the data corresponding to the park to be processed using the target task generation model to obtain the tasks to be executed. Process the data corresponding to the park to be processed and the tasks to be executed using the target task matching model to allocate the tasks to be executed. Wherein, the target task generation model and the target task matching model can be trained by the model training method according to any one of claims 1 to 7.
9. A model training device, characterized in that, The device includes: A first acquisition unit, configured to acquire the sample data corresponding to the sample park; wherein, the sample data at least characterizes the data of the sample devices in the sample park, the task data of the sample park, and the personnel information data corresponding to the sample park. A first processing unit, configured to construct an initial task generation model, and set first state information, first action information, first policy information, and first reward information for the initial task generation model. A second processing unit, configured to construct an initial task matching model, and set second state information, second action information, second policy information, and second reward information for the initial task matching model. A training unit, configured to train the set initial task generation model using a first target reinforcement learning algorithm based on the sample data to obtain a target task generation model. The training unit is further configured to train the set initial task matching model using a second target reinforcement learning algorithm based on the sample data to obtain a target task matching model.
10. An information determination device, characterized in that, The device includes: A second acquisition unit is used to acquire data corresponding to the park to be processed; wherein the data at least represents the data of the equipment of the park to be processed, the historical task data of the park to be processed, and the personnel information data corresponding to the park to be processed; A third processing unit is used to process the data corresponding to the to-be-processed park using a target task generation model to obtain a task to be executed; The third processing unit is further used to process the data corresponding to the to-be-processed park and the to-be-executed tasks using the target task matching model, so as to allocate the to-be-executed tasks; The target task generation model and the target task matching model may be obtained by training using the model training method described in any one of claims 1 to 7.
11. A model training device, characterized in that, The device comprises: a first processor, a first memory and a first communication bus; The first communication bus is used to realize the communication connection between the first processor and the first memory; The first processor is used to execute the model training program in the first memory to implement the steps of the model training method as described in any one of claims 1 to 7.
12. An information determination device, characterized in that, The device comprises: a second processor, a second memory, and a second communication bus; The second communication bus is used to realize the communication connection between the second processor and the second memory; The second processor is used to execute the information determination program in the second memory to implement the steps of the information determination method as claimed in claim 8.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the model training method as described in any one of claims 1 to 7, or the steps of the information determination method as described in claim 8.
Citation Information
Cited By
Model training method and text generation method
CN120725092A