Decision model adversarial migration method, storage medium and electronic device
Through the decision model confrontation migration method, the decision model of existing projects is migrated to a new project, solving the problem of reinforcement learning training time in new projects of central air-conditioning systems, and achieving rapid adaptation and efficient control.
Patent Information
- Application Number
- CN202411687763.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-07-25
AI Technical Summary
The reinforcement learning decision-making model is difficult to reuse directly in new central air-conditioning system projects, and the training time is long, making it difficult to quickly implement it in practical applications.
The decision model is used to adapt the decision model of existing projects to new projects through feature transfer learning, and the adversarial learning framework is used to adapt, including feature extractors, action classifiers and domain classifiers, and feature mapping and action decision-making are optimized.
The training time of the decision model in the new project is shortened, the applicability and control effect of the model in the new project is improved, and the computing power consumption and time-consuming in the initial debugging stage is reduced.
Smart Images

Figure CN120373410A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of HVAC control, and more particularly, to an adversarial transfer method for decision-making models, a storage medium, and an electronic device. Background Art
[0002] As a major consumer of building energy, the central air-conditioning system has become the research focus of many scholars and technicians to control the central air-conditioning system comfortably and energy-efficiently under the background of the country's promotion of energy conservation and carbon reduction. Since the system form of the central air-conditioning can be flexibly combined and the terminal form changes greatly, it is difficult to directly solve its optimal control parameters through theoretical methods. Traditional control methods use the method of preset rules combined with a schedule, which is too rigid in the perception of the terminal and difficult to consider the control lag from a strategic perspective. With the rise of big data and artificial intelligence algorithms, some researchers predict the terminal load and combine it with the system decision-making model for heuristic optimization, considering the impact of load changes on control in advance through prediction. However, limited by the problem of the accuracy decline of the prediction model in long-term prediction, this method can only consider the distribution and control of the overall load in a short period of time and cannot achieve the goal of optimal sequence control. Some researchers combine the advantages of reinforcement learning in sequential decision-making tasks with the demand for optimal comprehensive sequence energy efficiency in the control of the central air-conditioning system. By establishing a decision-making agent, corresponding decision-making control is carried out on the overall system. To achieve the optimal comprehensive energy efficiency of sequence control on the premise of meeting the terminal experience requirements. However, the implementation of the reinforcement learning method requires interactive training between the agent and the system environment. Since the central air-conditioning system takes a long time to reach a stable state after each control instruction is issued. Therefore, the control interval cannot be too short, which determines that the training of the reinforcement learning decision-making model takes a long time to converge. Due to the difficulty in ensuring the generality of the decision-making model, the differences in equipment and working conditions between different projects, directly applying the decision-making model of existing similar projects to new projects cannot obtain good control effects. Therefore, the decision-making model does not have generality and is difficult to reuse in other new projects. Currently, when most projects actually implement the reinforcement learning algorithm, they need to train each special engineering case from scratch separately. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide an adversarial transfer method for decision-making models, a storage medium, and an electronic device to improve the construction speed of the decision-making model for new projects.
[0004] An adversarial transfer method for a decision-making model includes:
[0005] Obtaining the decision target setting conditions of the target project, denoted as the first decision condition;
[0006] Obtaining the decision project setting conditions of the projects corresponding to the respective decision-making models in the decision-making model database, denoted as the second decision condition;
[0007] Use the decision-making model corresponding to the second decision-making condition that matches the first decision-making condition as the target decision-making model to be transferred; perform adversarial transfer on the target decision-making model;
[0008] Use the target decision-making model after adversarial transfer to learn and adapt to the target project.
[0009] Optionally, in the above decision-making model adversarial transfer method, the obtaining of the target decision-making model to be transferred includes:
[0010] Obtain the decision-making target setting conditions of the target project, denoted as the first decision-making condition;
[0011] Obtain the decision-making project setting conditions of the projects corresponding to each decision-making model in the decision-making model database, denoted as the second decision-making condition;
[0012] Use the decision-making model corresponding to the second decision-making condition that matches the first decision-making condition as the target decision-making model to be transferred.
[0013] Optionally, in the above decision-making model adversarial transfer method, the decision-making target setting conditions include:
[0014] Task objectives, system forms, and environmental conditions;
[0015] The use of the decision-making model corresponding to the second decision-making condition that matches the first decision-making condition as the target decision-making model to be transferred includes:
[0016] Obtain the similarity between the first decision-making condition and each second decision-making condition, and use the decision-making model corresponding to the second decision-making condition with the highest similarity to the first decision-making condition as the target decision-making model to be transferred.
[0017] Optionally, in the above decision-making model adversarial transfer method, performing adversarial transfer on the target decision-making model includes:
[0018] Perform adversarial transfer on the target decision-making model using an adversarial transfer framework, and the use of the adversarial transfer framework includes a feature extractor, an action classifier, and a domain classifier;
[0019] The feature extractor is used for feature extraction and feature mapping of the environmental state data of the transferable project and the target project, and the transferable project is the project corresponding to the target decision-making model;
[0020] The action classifier is used for classifying and making decisions on the actions to be controlled according to the feature data extracted by the feature extractor;
[0021] The domain classifier is used to determine whether the data features extracted by the feature extractor come from the source domain or the target domain to which the decision-making model needs to be transferred.
[0022] Optionally, in the above decision model adversarial transfer method, when performing adversarial transfer on the target decision model, the optimization objective of the adversarial transfer is:
[0023]
[0024] where W and b are the feature extraction layer parameters in the transfer process, V and c are the action classification prediction layer parameters in the transfer process, and u and z are the domain classification layer parameters in the transfer process. represents the action label prediction loss of the i-th data sample. represents the domain label prediction loss of the i-th data sample. Here, n represents the number of running data and action labels on the source domain, n' represents the number of running data on the target domain, and N = n + n' represents the total number of the training set.
[0025] Optionally, in the above decision model adversarial transfer method, before obtaining the decision item setting conditions of the projects corresponding to each decision model in the decision model database, it further includes:
[0026] Obtain the project type of the target project;
[0027] Determine a decision model database that matches the project type.
[0028] Optionally, in the above decision model adversarial transfer method, when the project type is a multi-connected air-conditioning system, the system form is the number of indoor and outdoor units of the multi-connected air-conditioning system, and the environmental condition is the operating environment of the source multi-connected air-conditioning system and the target multi-connected air-conditioning system.
[0029] A decision model adversarial transfer device includes:
[0030] A model selection unit, configured to obtain the decision target setting conditions of the target project, denoted as the first decision condition; obtain the decision item setting conditions of the projects corresponding to each decision model in the decision model database, denoted as the second decision condition; and use the decision model corresponding to the second decision condition that matches the first decision condition as the target decision model to be transferred.
[0031] A transfer unit, configured to perform adversarial transfer on the target decision model;
[0032] A learning adaptation unit, configured to perform learning adaptation on the target project by using the target decision model after adversarial transfer.
[0033] A computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it executes the method described in any one of the above.
[0034] An electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the method described in any one of the above through the computer program.
[0035] Based on the above technical solution, in the above solution provided by the embodiments of the present invention, the decision-making model trained in the existing project is adjusted through adversarial learning and then migrated to a new project with a similar system and working conditions. The decision-making model and data accumulated in the existing project are fully utilized, thereby alleviating the problem that it is difficult to implement the corresponding method in practice due to the long training time of the reinforcement learning decision-making model in the initial stage of debugging. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 is a schematic flowchart of a decision model adversarial migration method according to an embodiment of the present application;
[0039] Figure 2 is a schematic flowchart of a decision model adversarial migration method disclosed in another embodiment of the present application;
[0040] Figure 3 is a schematic structural diagram of a decision model migration framework disclosed in an embodiment of the present application;
[0041] Figure 4 is a schematic structural diagram of a decision model adversarial migration device disclosed in an embodiment of the present application;
[0042] Figure 5 is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0044] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0045] The decision-making agent is a type of AI agent, which is mainly responsible for finding the corresponding decision rules in the knowledge base according to the results generated by the data fusion agent. These rules can be expressed by production rules. After querying the corresponding rules, the instruction set in the rules is called to control the control system agent, so as to obtain an ideal system control effect. It interacts with the environment and makes decisions to achieve the goal. It is an autonomous system that can make decisions based on the information at hand and forms the basis of many artificial intelligence systems.
[0046] In order to solve the problem that it is difficult to implement the reinforcement learning method in the actual central air-conditioning system project and reduce the computing power consumption and time consumption during the initial debugging stage for the intelligent agent strategy training process. According to one aspect of the embodiments of this application, a decision model adversarial transfer method is provided. The decision model transfer method of this smart home device is widely applied to the whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, Intelligence House ecosystem, etc. The present invention proposes to use the domain adversarial decision model transfer method to transfer the decision model trained in the existing project to a new project with a similar system and working conditions after adjustment through adversarial learning. Make full use of the decision models and data accumulated in the existing projects, so as to alleviate the problem that the corresponding method is difficult to implement in practice due to the long training time of the reinforcement learning decision model in the initial debugging stage.
[0047] See Figure 1 , the decision model adversarial transfer method disclosed in this embodiment includes:
[0048] Step S101: Select a transferable decision model.
[0049] In the process of implementing adversarial transfer of the decision-making model of the reinforcement learning agent, it is first necessary to determine which decision-making model to use as the source decision-making model for transfer to a new project. The source decision-making model is a model that has been trained and can be well used in the existing project, and this source decision-making model is the decision-making model selected in this step.
[0050] In the technical solution disclosed in this embodiment, each decision-making model corresponds to a decision-making target setting condition corresponding to it. The decision-making target setting condition determines the task target of the decision-making model, the system form and environmental working conditions of the application system corresponding to the decision-making model. The application system is the system to which the decision-making model is applied. For example, it can be a multi-connected air-conditioning control system. When selecting a transferable decision-making model, the decision-making target setting condition needs to be used as a limiting condition to select the decision-making model required for this transfer from multiple decision-making models. Specifically, when selecting a transferable decision-making model, it is necessary to obtain the decision-making target setting condition corresponding to the target project, denoted as the first decision condition, obtain the decision-making project setting conditions of the projects corresponding to each decision-making model in the decision-making model database, denoted as the second decision condition, calculate the similarity between the first decision condition and the second decision condition, and use the decision-making model corresponding to the second decision condition with the highest similarity to the first decision condition as the selected transferable decision-making model. Specifically, see Figure 2 This process includes:
[0051] Step S201: Determine the second decision condition that has the same task target as the first decision condition, denoted as the candidate second decision condition;
[0052] Taking the target project as a central air-conditioning system as an example, the setting of the task target for the agent may be different on different projects. For example, the task target of a precision factory building project is to achieve accurate terminal control parameters as the primary goal. And for example, the task target of an office project is to save energy as much as possible while achieving relatively comfortable terminal requirements. The control strategies of the agent vary greatly between different task targets. Therefore, it is necessary to ensure that the task targets of the two projects before and after transfer are the same.
[0053] Step S202: Calculate the similarity of the system form and environmental working conditions between the first decision condition and each candidate second decision condition.
[0054] For the system form and environmental conditions of the project, the system form and environmental conditions between the two projects before and after migration need to have a certain degree of similarity, which is also the basis for migration. For example, taking the target project as a multi-split air-conditioning system, the similar system form mainly refers to the similarity in the number of outdoor units and indoor units, and the similar environmental conditions refer to the similarity in the external operating conditions of the target project, such as the same region or different regions with similar climates. The higher the similarity of the system form and environmental conditions between the two projects before and after migration, the better the applicability of the decision model after migration.
[0055] Step S203: Use the decision model corresponding to the candidate second decision condition with the highest similarity calculated in Step S202 as the target decision model to be migrated.
[0056] Step S102: Adversarial transfer of the decision model.
[0057] After selecting the transferable decision model in Step S101, the most important part of the present invention - adversarial transfer of the decision model is carried out. When migrating the target decision model, an appropriate transfer learning method can be selected according to its own needs. Different transfer learning methods have different requirements for the knowledge and data in the source domain and the target domain. The source domain refers to the project of the target decision model, and the target domain refers to the target project to which it is required to be migrated. Among them, the feature transfer learning method only requires unlabeled data in the target domain and is applicable to the case where the system forms of the source domain and the target domain are similar but the data distributions are inconsistent. This transfer method is applicable to the case of migrating the reinforcement learning decision model to be solved in the present invention. Therefore, the feature transfer learning method is used as the transfer method for adversarial transfer of the target decision model in this application. The core content of the feature transfer learning method is: map the data in the target domain and the source domain to a common feature space so that the data in the two domains have similar features in the new space to eliminate the differences between domains. The feature transfer learning method can be implemented through a pre-constructed decision model transfer framework. This framework migrates the decision model on the selected transferable project to the target project through adversarial learning of the data features of the decision model on the selected transferable project and the operating data on the target project.
[0058] The specific structure is as Figure 3As shown in the figure, the decision model transfer framework proposed by the present invention can be composed of three parts: a feature extractor, an action classifier, and a domain classifier. Among them, the feature extractor is mainly used for feature extraction and feature mapping of the environment state data of transferable items and new items. The action classifier is used to classify and make decisions on the actions to be controlled according to the features extracted by the feature extractor from the environment state data of transferable items and new items. The domain classifier is used to distinguish whether the data features come from the source domain (existing decision model) or the target domain to which the decision model needs to be transferred. The combination of the feature extractor and the action classifier is the action decision model of the intelligent agent. The domain classifier is mainly used for adversarial game training with the feature extractor, so as to adapt the existing decision model to the new target domain by means of feature transfer.
[0059] The above-mentioned feature transfer learning method proposed by the present invention can also be realized by adding a gradient reversal layer between the feature extractor and the domain classifier, so as to maximize the domain classification loss while minimizing the action label classification loss. Thus, in the process of confrontation between the feature extractor and the domain classifier, the decision model on the source domain is transferred to the target domain by means of feature transfer learning method.
[0060] When the decision model is the decision model of the central air-conditioning system, considering the effective accuracy of the control actions and the system stability problems in the central air-conditioning system. Usually, in the process of training the intelligent agent using reinforcement learning, the action range of the intelligent agent each time is limited. For example, taking the multi-connected unit system as an example, when the target parameter to be controlled is the high / low pressure setting value of the outdoor unit, the action range of the intelligent agent each time is generally set to [-0.5, 0.5]. In terms of control accuracy, the recognition accuracy of the multi-connected unit system can only reach the level of 0.1. That is to say, the actual effective value of the above action range is 11, that is, the continuous control parameters can be converted into discrete forms. Therefore, the output of the decision model for actions can be approximately regarded as a process of classifying environmental features, and finally the actions to be taken are output in the form of an action classifier. Thus, in the adversarial transfer decision model framework proposed by the present invention, the decision model of the intelligent agent is divided into two parts: a feature extractor and an action classifier.
[0061] Description of the training objectives of each structure and the update method of its parameters included in each part of the adversarial transfer framework proposed by the present invention:
[0062] The optimization objective of the action classifier can be expressed as shown in formula (1).
[0063]
[0064] That is, the action classifier is implemented through formula (1), where λ·R(W, b) in formula (1) represents L2 regularization weighted by the hyperparameter λ (L2 regularization (Ridge Regression) is the square root of the sum of the squares of the elements of the vector). represents the action prediction loss of the i-th data sample. Here, the sample is the feature parameter extracted by the feature extractor, W and b are the feature extraction layer parameters in the migration process, and V and c are the action classification prediction layer parameters in the migration process.
[0065] The optimization objective of the domain classifier can be expressed as shown in formula (2).
[0066]
[0067] In the formula represents the domain label prediction loss of the i-th sample; u and z are the domain classification layer parameters in the migration process, that is, the domain recognition results; n represents the number of running data and action labels on the source domain, n′ represents the number of running data on the target domain, and N = n + n′ represents the total number of the training set.
[0068] Therefore, when performing adversarial migration on the target decision model, the optimization objective of the adversarial migration can be expressed as shown in formula (3).
[0069]
[0070] The parameter update processes of the three networks, namely the feature extractor, the action classifier, and the domain classifier, are shown in formulas (4) to (5):
[0071]
[0072]
[0073]
[0074] In formulas (4), (5), and (6), F represents the feature extractor, C represents the action classifier, d represents the domain classifier, and θ c , θd, and θF respectively represent the network parameters of C, F, and d; Lc and L d are the loss functions of C and d; μ represents the learning rate; -λ is the gradient reversal factor.
[0075] During the training process, to ensure the balance of the training data set. The difference in the ratio of the data from the source domain and the data from the target domain in the training set should be within the allowable difference range. For example, the values of n and n′ are equal.
[0076] Step S103: The learning adaptation of the migrated decision model to the new project.
[0077] After the above-mentioned adversarial transfer of the decision-making model, the decision-making model of the existing project can be transferred to the target project for use in the control decision-making in the initial stage. In order to enable the decision-making model to further learn and adapt during the operation of the actual project, it is necessary to perform further interactive training on the transferred decision-making model on the target project. Taking the reinforcement learning algorithm with the Actor-Critic structure as an example for illustration (only for illustration purposes and not as a basis for limiting the scope of application of this patent), after the adversarial transfer, a decision-making model applied to the target project is obtained, and this decision-making model is denoted as the final decision-making model. At this time, the decision-making control of the final decision-making model is reasonable. Subsequently, based on this final decision-making model for policy improvement, it is necessary to further train according to the interaction between the intelligent agent and the actual system. In order to utilize the advantages of the final decision-making model, the present invention proposes that when using the final decision-making model to learn and adapt to the target project, at the initial stage of continuous learning and adaptation of the final decision-making model, the Critic network is preferentially trained. The transferred final decision-making model interacts with the environment of the target project to generate changes in decision-making actions, states, and rewards of the target project, and these data are used to train the Critic network, that is, the value network. Through a period of training, the prediction accuracy of the value network is improved, and then the training mode of the normal reinforcement learning algorithm with the Actor-Critic structure is used to train and update both the action and value networks simultaneously. This can avoid, in the initial stage of the transferred training, the influence of the low value determination degree of the value network on the action on the correctness of the policy improvement direction of the final decision-making model. Make full use of the advantages of the final decision-making model obtained by adversarial transfer in the initial stage control. Avoid the degradation of the final decision-making model caused by the low value prediction accuracy of the Critic network in the initial stage. Thereby affecting the decision-making model to make relatively correct control instructions and affecting the decline of the overall end-user experience.
[0078] The present invention proposes to use the domain adversarial decision-making model transfer method to transfer the trained decision-making model in the existing project to a new project with similar systems and working conditions through adversarial learning after adjustment. Make full use of the decision-making model and data accumulated in the existing project, thereby alleviating the problem that the corresponding method is difficult to implement in practice due to the long training time of the reinforcement learning decision-making model in the initial debugging stage.
[0079] In this embodiment, a decision-making model adversarial transfer device is disclosed. For the specific working content of each unit in the device, please refer to the content of the above method embodiment.
[0080] Next, the decision-making model adversarial transfer device provided by the embodiment of the present invention will be described. The decision-making model adversarial transfer device described below can be mutually corresponding and referred to the decision-making model adversarial transfer method described above.
[0081] See Figure 4 , the device may include:
[0082] A model selection unit 10, configured to obtain the decision target setting conditions of the target project, denoted as the first decision condition; obtain the decision project setting conditions of the projects corresponding to each decision model in the decision model database, denoted as the second decision condition; and use the decision model corresponding to the second decision condition that matches the first decision condition as the target decision model that can be migrated.
[0083] A migration unit 20, configured to perform adversarial migration on the target decision model.
[0084] A learning adaptation unit 30, configured to perform learning adaptation on the target project by using the target decision model after adversarial migration.
[0085] The model selection unit 10 obtains the target decision model that can be applied to the target project by selecting from the decision models trained in the existing projects, and then the migration unit 20 performs adversarial migration on the target decision model and migrates it to the target project. Then, the learning adaptation unit 30 performs learning adaptation on the target project by using the migrated target decision model. Thus, the decision models and data accumulated in the existing projects are fully utilized, thereby alleviating the problem that the corresponding method is difficult to implement in practice due to the long training time of the reinforcement learning decision model in the initial debugging stage.
[0086] Corresponding to the above device, the present application also discloses a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it executes the method described in any one of the above.
[0087] In an embodiment of the present application, an electronic device is further provided. Refer to Figure 5 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 5 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0088] As Figure 5As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. The program in the random access memory (RAM) 603 is used to implement the various steps of the decision model adversarial migration method disclosed in any of the above embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0089] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0090] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0091] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0092] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0093] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. A decision model adversarial transfer method, characterized in that, Including: Obtain the decision target setting conditions of the target project, denoted as the first decision condition; Obtain the decision project setting conditions of the projects corresponding to each decision model in the decision model database, denoted as the second decision condition; Use the decision model corresponding to the second decision condition that matches the first decision condition as the target decision model to be transferred; Perform adversarial transfer on the target decision model; Use the target decision model after adversarial transfer to perform learning adaptation on the target project.
2. The method according to claim 1, characterized in that The decision target setting conditions include: Task target, system form, and environmental conditions; The step of using the decision model corresponding to the second decision condition that matches the first decision condition as the target decision model to be transferred includes: Obtain the similarity between the first decision condition and each second decision condition, and use the decision model corresponding to the second decision condition with the highest similarity to the first decision condition as the target decision model to be transferred.
3. The method according to claim 2, wherein Performing adversarial transfer on the target decision model includes: Use an adversarial transfer framework to perform adversarial transfer on the target decision model. The use of the adversarial transfer framework includes a feature extractor, an action classifier, and a domain classifier; The feature extractor is used for feature extraction and feature mapping of the environmental state data of the transferable project and the target project. The transferable project is the project corresponding to the target decision model; The action classifier is used to classify and make decisions on the actions to be controlled according to the feature data extracted by the feature extractor; The domain classifier is used to determine whether the data features extracted by the feature extractor come from the source domain or the target domain to which the decision model needs to be transferred.
4. The method according to claim 3, wherein When performing adversarial transfer on the target decision model, the optimization objective of the adversarial transfer is: Among them, W and b are the feature extraction layer parameters during the migration process, V and c are the action classification prediction layer parameters during the migration process, and u and z are the domain classification layer parameters during the migration process. represents the prediction loss of the action label for the i-th data sample. represents the prediction loss of the domain label for the i-th data sample, where n represents the number of running data and action labels on the source domain, n' represents the number of running data on the target domain, and N = n + n' represents the total number of the training set.
5. The method according to claim 4, characterized in that, Before obtaining the decision project setting conditions of the projects corresponding to each decision model in the decision model database, it further includes: Obtain the project type of the target project: Determine the decision model database that matches the project type.
6. The method according to claim 5, characterized in that, When the project type is a multi-connected air-conditioning system, the system form is the number of indoor and outdoor units of the multi-connected air-conditioning system, and the environmental conditions are the operating environments of the source multi-connected air-conditioning system and the target multi-connected air-conditioning system.
7. A decision model adversarial transfer device, characterized in that Including: A model selection unit, configured to obtain the decision target setting conditions of the target project, denoted as the first decision condition; Obtain the decision project setting conditions of the projects corresponding to each decision model in the decision model database, denoted as the second decision condition; use the decision model corresponding to the second decision condition that matches the first decision condition as the target decision model to be transferred; A migration unit, configured to perform adversarial transfer on the target decision model; A learning adaptation unit, configured to use the target decision model after adversarial transfer to perform learning adaptation on the target project.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 6.
9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 6 through the computer program.