Multi-agent collaborative intelligent equipment control method, device and equipment and medium
Through the multi-agent collaborative intelligent device control method, Bayesian network and particle filtering technology are used to dynamically adjust action priorities and optimize the coordination and response between devices, which solves the problem of insufficient coordination between devices in the intelligent device system and realizes efficient and personalized device control and response.
Patent Information
- Application Number
- CN202511060469.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing intelligent device systems suffer from insufficient coordination between devices, poor global optimization capabilities, and low response efficiency. This is especially true in the healthcare and financial technology fields, where there are problems such as poor coordination mechanisms between devices, poor global optimization capabilities, and low response efficiency.
A multi-agent collaborative intelligent device control method is adopted. The real-time scene data and user status data are analyzed through Bayesian networks to build a structured causal reasoning model. The scene activation and action instruction generation of particle filtering are combined to dynamically adjust the action priority, optimize the initial action instructions, perform conflict detection and resolution, generate collaborative decision-making reports, and achieve efficient coordination and personalized control between devices.
It improves the coordination mechanism, global optimization capability and response efficiency between devices, enhances the stability and response speed of the system, satisfies the user's personalized experience, and significantly improves the flexibility and user satisfaction of the smart device system.
Smart Images

Figure CN120802787A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a multi-agent cooperative intelligent device control method and device, equipment and medium. BACKGROUND
[0002] The intelligent device scene switching system currently generally adopts technical solutions such as centralized rule engine, single-agent decision system and simple distributed cooperative system. The centralized system realizes device linkage through fixed rules, the single-agent system relies on a central controller to collect sensor data and issue instructions, and handles conflicts between devices; and the simple distributed system communicates through Zigbee or Wi-Fi protocols, and performs local control based on an event triggering mechanism. However, the existing technology has some deficiencies.
[0003] In the medical health field, the existing intelligent device system faces the problem of insufficient coordination between devices. Especially when multiple devices need to respond to the health needs of patients at the same time, due to priority conflicts or system response delays, the environment settings may not be adjusted in time, thereby affecting the health status of patients. In addition, the system lacks deep learning and personalized analysis of patient health data, and fails to dynamically adjust the weight coefficients and settings of the devices, so that the special needs of individual patients cannot be fully met. Moreover, the real-time problem of the system, especially when the number of devices increases, may cause congestion of data transmission or delay of calculation, and cannot provide timely feedback in critical situations, affecting the accuracy and efficiency of medical decision-making.
[0004] In the field of financial technology business, the limitations of the intelligent device system mainly lie in data fusion and security. Due to the sensitivity of financial data, the intelligent device system is prone to data privacy and security risks when collecting and processing users' financial information. Especially in the cooperative work between devices, the flow of information may expose users' private data such as financial status and consumption habits. In addition, the system's prediction accuracy of user demand is insufficient, still relying on preset rules or single-dimensional behavior analysis, which is difficult to cope with complex and changing financial environment. With the increase in the number of devices, the lack of data transmission and computing resources may cause delay in information processing, affecting real-time financial product recommendation and decision-making execution, and reducing user experience.
[0005] In summary, when multiple scenarios trigger conflicting instructions, the existing system cannot evaluate and optimize the conflict resolution in real time, and usually relies on fixed priority or manual intervention. Local decision-making is disconnected from global optimization, centralized systems are prone to computational bottlenecks, distributed systems lack cross-device collaborative optimization models, and the linkage effects between devices are not fully considered. Preset rules are difficult to adapt to complex dynamic environments, and lack the ability to learn from historical user data to adjust device weights. Existing communication modes are prone to message congestion as the number of devices increases, and the performance of edge computing is not fully utilized, resulting in decision-making delays and affecting the real-time and robustness of the system.
[0006] Therefore, the prior art has the problems of poor coordination mechanism between devices, poor global optimization capability, and low response efficiency. SUMMARY
[0007] The present application provides a multi-agent collaborative intelligent device control method, device, equipment and medium, which mainly aims to solve the problems of poor coordination mechanism between devices, poor global optimization capability and low response efficiency.
[0008] In a first aspect, to achieve the above-mentioned purpose, the present application provides a multi-agent collaborative intelligent device control method, comprising: obtaining a target scene of a target user and a target parameter of the target scene, and collecting scene real-time data of the target parameter and user state data of the target user; using the scene real-time data and the user state data, analyzing the environment state of the target scene through a Bayesian network to obtain an environment current state; activating the scene according to the environment current state based on the scene real-time data, and generating an initial action instruction according to the activated scene; obtaining a state space, an action space and a reward function of the target scene, dynamically adjusting the action priority of the action space using the state space and the reward function, and optimizing the initial action instruction according to the adjusted action priority to obtain a target action instruction; executing the target action instruction on a target device in the target scene, and performing conflict detection and resolution on the target device during the execution process to generate a collaborative decision report; performing preference analysis on the target user using the user state data, and controlling the target device based on the user preference behavior obtained by the analysis, the target action instruction and the collaborative decision report.
[0009] In a second aspect, the present application further provides a multi-agent collaborative intelligent device control device, comprising: A data acquisition module is used to acquire a target scenario of a target user and target parameters of the target scenario, and collect real-time scenario data of the target parameters and user status data of the target user; A state analysis module is used to analyze the environmental state of the target scene using the real-time scene data and the user state data through a Bayesian network to obtain the current state of the environment; A scene activation module, configured to activate the scene real-time data according to the current state of the environment, and generate an initial action instruction according to the activated scene; an instruction adjustment module, configured to obtain the state space, action space, and reward function of the target scene, dynamically adjust the action priority of the action space using the state space and the reward function, and optimize the initial action instruction according to the adjusted action priority to obtain a target action instruction; A conflict detection module is used to execute the target action instruction on the target device in the target scene, and perform conflict detection and resolution on the target device during the execution process to generate a collaborative decision report; The device control module is used to use the user status data to perform preference analysis on the target user, and control the target device in combination with the user preference behavior obtained by the analysis, the target action instruction and the collaborative decision report.
[0010] In a third aspect, the present invention further provides an electronic device, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned multi-agent collaborative intelligent device control method.
[0011] In a fourth aspect, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned multi-agent collaborative intelligent device control method.
[0012] The application obtains a target scene of a target user and a target parameter of the target scene, collects scene real-time data of the target parameter and user state data of the target user, uses the scene real-time data and the user state data, analyzes an environment state of the target scene through a Bayesian network to obtain an environment current state, a Bayesian network analysis method fuses the scene real-time data and the user state data to construct a structured causal reasoning model, can not only realize dynamic perception and uncertainty modeling of the environment state, but also can maintain high reasoning accuracy in the case that multiple source data exist noise or are partially missing, activates the scene according to the environment current state and generates an initial action instruction according to the activated scene, a scene activation and action instruction generation method based on particle filtering can continuously optimize the recognition accuracy of the environment state in a complex scene where uncertainty and real-time coexist through multi-particle simulation and dynamic iteration of the environment current state, obtains a state space, an action space and a reward function of the target scene, dynamically adjusts an action priority of the action space using the state space and the reward function, and optimizes the initial action instruction according to the adjusted action priority to obtain a target action instruction, an action priority dynamic adjustment mechanism driven by the state space and the reward function can realize intelligent optimization of the initial action instruction, ensures that the system always executes the operation instruction with the highest value and the best effect under different states, executes the target action instruction on a target device in the target scene, and performs conflict detection and resolution on the target device in the execution process to generate a collaborative decision report, discovers and effectively solves the action conflict between devices in time through real-time monitoring of the execution state of the device, avoids device abnormities or function failures caused by operation conflicts, realizes optimized scheduling of device collaborative work through sequential adjustment based on the action priority, improves the stability and response efficiency of the overall operation of the system, uses the user state data to analyze the preferences of the target user, and controls the target device in combination with the user preference behavior obtained through the analysis, the target action instruction and the collaborative decision report, meets the individualized experience of the user, ensures efficient operation of the device, and improves the coordination mechanism between devices, the global optimization capability and the response efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the application. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0014] Figure 1An application environment schematic diagram of a multi-agent cooperative intelligent device control method in an embodiment of the present application; Figure 2 A flowchart of a multi-agent cooperative intelligent device control method in an embodiment of the present application; Figure 3 A flowchart of a device control process in a multi-agent cooperative intelligent device control method in an embodiment of the present application; Figure 4 A module schematic diagram of a multi-agent cooperative intelligent device control device in an embodiment of the present application; Figure 5 A structure schematic diagram of an electronic device for implementing a multi-agent cooperative intelligent device control method in an embodiment of the present application; Figure 6 Another structure schematic diagram of an electronic device for implementing a multi-agent cooperative intelligent device control method in an embodiment of the present application.
[0015] The object, function characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0016] In order to make the person in the art better understand the technical solutions of the present disclosure, and understand the implementation process of how to apply technical means to solve the technical problems and achieve the corresponding technical effects, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all the embodiments. The embodiments of the present disclosure and each feature in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present disclosure.
[0017] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present disclosure and above-described drawings are intended to distinguish similar objects and are not necessarily used to describe a particular sequential or chronological order. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product or apparatus including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or apparatus.
[0018] The embodiment of the present application provides a kind of multi-agent collaborative intelligent device control method, the execution subject of the multi-agent collaborative intelligent device control method includes but is not limited to at least one of the electronic device that can be configured to execute the device provided by the embodiment of the present application, such as server, terminal etc.It is said that the multi-agent collaborative intelligent device control method can be executed by the software or hardware installed in terminal device or server device.The server includes but is not limited to: single server, server cluster, cloud server or cloud server cluster etc.The server can be independent server, can also be cloud server that provides cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network (Content Delivery Network, CDN), and big data and artificial intelligence platform etc.Basic cloud computing services.
[0019] The embodiment of the present application provides a kind of multi-agent collaborative intelligent device control method, which can be applied to Figure 1In the application environment, the client communicates with the server through the network. The server can obtain the target scene of the target user and the target parameters of the target scene through the client, collect scene real-time data of the target parameters and user state data of the target user, analyze the environment state of the target scene through the Bayesian network by using the scene real-time data and the user state data, obtain the current environment state, and construct a structured causal reasoning model by fusing the scene real-time data and the user state data. The Bayesian network analysis method can not only realize dynamic perception and uncertainty modeling of the environment state, but also maintain high reasoning accuracy in the case of noise or partial missing of multi-source data. The scene real-time data is activated according to the current environment state, and an initial action instruction is generated according to the activated scene. The scene activation and action instruction generation method based on particle filtering can continuously optimize the recognition accuracy of the environment state in the complex scene with uncertainty and real-time coexistence through multi-particle simulation and dynamic iteration of the current environment state, obtain the state space, action space and reward function of the target scene, dynamically adjust the action priority of the action space by using the state space and the reward function, optimize the initial action instruction according to the adjusted action priority, and obtain a target action instruction. The action priority dynamic adjustment mechanism driven by the state space and the reward function can realize intelligent optimization of the initial action instruction, ensure that the system always executes the operation instruction with the highest value and the best effect under different states, execute the target action instruction on the target device in the target scene, and perform conflict detection and resolution on the target device during the execution process to generate a collaborative decision report. Through real-time monitoring of the execution state of the device, the action conflict between devices can be found and effectively solved in time, and the device abnormity or function failure caused by operation conflict can be avoided. Through the sequence adjustment based on the action priority, the optimization scheduling of device collaborative work is realized, the stability and response efficiency of the overall operation of the system are improved, the user state data is used for preference analysis of the target user, and the target device is controlled in combination with the user preference behavior obtained by analysis, the target action instruction and the collaborative decision report, so that the user's individual experience is met, and the efficient operation of the device is ensured. The coordination mechanism between devices, the global optimization capability and the response efficiency are improved. Finally, the state of the target device is output and fed back to the user client. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be realized by an independent server or a server cluster composed of multiple servers. The application will be described in detail through specific embodiments.
[0020] The following is explained in the description of the present application. The present application fuses scene real-time data and user state data through a Bayesian network analysis method. Not only can the dynamic perception of the environmental state and the uncertainty modeling be realized, but also the high inference accuracy can be maintained in the case of noise or partial missing of multi-source data. The scene activation and action instruction generation method based on particle filtering can simulate and dynamically iterate the current state of the environment through multiple particles. In the complex scene where uncertainty and real-time coexist, the recognition accuracy of the environmental state can be continuously optimized. The action priority dynamic adjustment mechanism based on the state space and reward function can realize the intelligent optimization of the initial action instruction, ensuring that the system always executes the operation instruction with the highest value and best effect in different states. Combined with the target action instruction and the collaborative decision report, a highly adaptive control strategy can be generated, and the device state can be monitored and adjusted in real time, significantly improving the flexibility, response speed and user satisfaction of the intelligent device system.
[0021] Referring to Figure 2 Fig. 1 is a flowchart of a multi-agent collaborative intelligent device control method according to an embodiment of the present application. In this embodiment, the multi-agent collaborative intelligent device control method comprises the following steps: S1, obtaining a target scene of a target user and a target parameter of the target scene, and collecting scene real-time data of the target parameter and user state data of the target user.
[0022] In this embodiment of the present application, the system obtains the target scene currently selected by the target user through a user interface or a voice assistant. For example, the user selects a scene such as a “movie watching mode”, a “leaving home mode” or a “sleeping mode” through a touch screen, a voice instruction or a mobile device. This step automatically obtains the user’s intention through user input or system prediction, and determines the current scene type.
[0023] Once the target scene is determined, the corresponding target parameters are automatically extracted according to the scene type. For example, in the “movie watching mode”, the target parameters may include the light color temperature (such as 2700K), the air conditioner temperature (such as 22℃), the sound volume (such as medium volume) and the like; in the “leaving home mode”, the target parameters may include air conditioner off, curtain closing, light off and the like. The system defines these target parameters through a preset scene parameter library or a dynamic scene engine.
[0024] Real-time data related to the target scene is collected. The real-time data comes from sensors or device state feedback. For example, the real-time data that may need to be obtained in the scene includes environmental temperature and humidity, light intensity, air quality (such as PM2.5 concentration), device energy consumption and the like. These data will be transmitted to the central control system in real time to support subsequent decision-making.
[0025] Obtain the state data of the target user, usually collected by wearable devices (such as smart bracelets, heart rate monitors) or environmental sensors (such as UWB positioning systems, smart mattresses) and the like. The user state data may include heart rate, body temperature, fatigue, user location, etc., reflecting the current physical state or activity state of the user. For example, when the user is in a "fatigue state", the system can automatically adjust the light brightness or air conditioner temperature to provide a more comfortable environment.
[0026] S2, analyze the environment state of the target scene by Bayesian network based on the scene real-time data and the user state data, and obtain the current environment state.
[0027] In the embodiment of the application, the scene real-time data and the user state data are normalized to obtain standardized scene data and user state data respectively. The environment state parameters and device state parameters are taken as state nodes, and the standardized data is taken as observation nodes. The node dependency relationship is constructed to form a Bayesian network. The conditional probability table of each node is generated based on the transition probability matrix of the network. The probability distribution of the current environment state of the target scene is inferred using the conditional probability table, and the current state of the environment is determined accordingly.
[0028] In the medical health specific scene, the Bayesian network analysis method can be used in the intelligent ward environment monitoring system. By collecting real-time scene data such as temperature, humidity, air quality, and device running state in the ward, as well as patient vital sign monitoring data (such as heart rate, blood pressure, activity level, etc.), the current ward environment state is probabilistically modeled, and the ward comfort level or infection risk level is dynamically evaluated, thereby assisting medical staff to optimize the environment control strategy and improving the patient rehabilitation efficiency.
[0029] In the financial technology specific scene, the Bayesian network analysis method can be applied to the intelligent risk control system. By performing Bayesian network modeling on user behavior data (such as transaction frequency, login location, device fingerprint) and real-time scene data (such as market fluctuations, system access pressure, etc.), the risk state of the current account is inferred, real-time early warning and intervention of potential fraudulent transactions are realized, and the security and response efficiency of the financial transaction system are improved.
[0030] In the embodiment of the application, the environment state of the target scene is analyzed by Bayesian network based on the scene real-time data and the user state data, and the current environment state is obtained, including: The scene real-time data and the user state data are normalized to obtain scene standard data and user state standard data; Extract the environment state parameters and device state parameters of the scene standard data, and the user state parameters of the user state standard data; The environment state parameter, the device state parameter and the user state parameter are taken as state nodes; The scene standard data and the user state standard data are taken as observation nodes; A node dependency relationship is acquired, the state nodes and the observation nodes are connected by using the node dependency relationship, and a Bayesian network is constructed; A transition probability matrix of the Bayesian network is acquired, and a conditional probability table of each node is generated by using the transition probability matrix; The conditional probability table is used to determine a probability distribution of a current environment state of the target scene; According to the probability distribution, a current environment state of the target scene is determined.
[0031] In detail, the collected scene real-time data (such as temperature, humidity, device power and the like) and user state data (such as physiological parameters, behavior characteristics and the like) are converted into a unified format and cleaned, and abnormal values and missing values are removed. According to the value range of each data, a minimum-maximum normalization or Z-score standardization method is used to map data of different dimensions to a unified standardized interval, and scene standard data and user state standard data available for subsequent analysis are obtained respectively, so as to eliminate the influence of data scale difference.
[0032] After completing the standardization processing, parameters related to the environment (such as temperature, humidity, light intensity, air quality and the like) are identified and extracted from the scene standard data as environment state parameters, and data reflecting the running conditions of various intelligent devices (such as switch state, running mode, power consumption and the like) are extracted as device state parameters. Through parameter classification and labeling, the structural distinction of environmental factors and device behaviors in the scene is realized, and clear state variable input is provided for subsequent Bayesian network modeling.
[0033] In the process of constructing the Bayesian network, the extracted environment state parameters (such as temperature and humidity, air quality and the like), device state parameters (such as device switch state, running mode and the like) and user state parameters (user position, heart rate, fatigue degree) are set as state nodes, representing the core environment and device state to be inferred in the target scene. The normalized scene standard data and user state standard data are taken as observation nodes, representing the external input information that can be directly acquired by the system, so as to establish the network structure through the dependency relationship between the observation nodes and the state nodes, and provide a basis support for subsequent probability inference.
[0034] By analyzing the logical and statistical dependencies between state nodes and observation nodes, combined with historical data or domain knowledge, the causal or conditional correlations between nodes are identified, and the directed connections between nodes are constructed; for example, the temperature and humidity state nodes may affect the user's sense of comfort, an observation node, and the device operating state node may be affected by user behavior data. Based on these dependencies, the state nodes and observation nodes are connected in order to form a structured Bayesian network, providing graph model support for probabilistic reasoning of environmental state.
[0035] On the basis of the constructed Bayesian network structure, the joint probability of state transition between nodes is calculated through historical observation data statistics or expert knowledge inference, and the transition probability matrix of the network is generated , describing the time series correlation of environmental parameters, based on the matrix to establish the corresponding conditional probability table for each node, and clearly define the probability distribution of the node value under different parent node states. The joint probability model is:
[0036] Among them, represents all states from time 1 to , represents all observation nodes from time 1 to ; represents the joint probability of state sequence and observation sequence in the time range from 1 to ; represents the probability of state at time ; represents the conditional probability of state transition, the probability of state at time , given that the previous state is ; represents the conditional probability of observation, the probability of observation value at time , given that the current state is ; represents the conditional probability of each state transition from time 2 to ; represents the observation probability given the current state at each time from time 1 to .
[0037] By inputting the real-time data of the current observation node, the state node is updated using the Bayesian inference mechanism, and the probability distribution of the target scene current environmental state under each possible state is finally obtained, providing a basis for state judgment and control decision.
[0038] Based on the environment state probability distribution obtained by Bayesian network inference, each possible environment state is evaluated, and the state with the maximum posterior probability is selected as the current environment state of the target scene; at the same time, if there are multiple states with close high probability values, a threshold or strategy can be set in combination with business requirements to make a fuzzy judgment or joint decision on the environment state, so as to realize accurate identification and dynamic feedback of the current scene environment state.
[0039] The Bayesian network analysis method can not only realize dynamic perception and uncertainty modeling of the environment state by fusing scene real-time data and user state data, but also maintain high inference accuracy in the case of noise or partial missing of multi-source data; compared with the traditional static rule or threshold judgment method, it has stronger adaptability and interpretability, and is helpful for the system to realize intelligent and personalized environment management and active decision response.
[0040] S3, according to the current environment state of the scene, the scene is activated, and the initial action instruction is generated according to the activated scene.
[0041] In the embodiment of the application, an initial particle set is generated based on the current environment state of the target scene, which is used to simulate the possible state evolution path in the scene, and the state of each particle is iteratively updated in combination with the scene real-time data and the preset state transition function, forming an updated particle state. By matching with the real-time data collected by the sensor, the observation probability of each particle is calculated, and the initial weight of the particle is updated. Then, based on the updated weight, the particles are screened to form an updated particle set, and it is judged whether the particle set meets the preset condition of the activated scene. If it meets the condition, the current target scene is confirmed as the activated scene, and on this basis, the action instruction matching is performed to generate the corresponding initial action instruction, which provides a basis for the subsequent control execution of the system.
[0042] In the medical health specific scene, the particle filtering driven scene activation and action instruction generation method can be applied in an intelligent ward or home health monitoring system. For example, by analyzing the current environment state and physiological data of the patient, simulating state particles are generated, and the particle state and weight are updated in real time by comparing the sensor data, and it is dynamically identified whether there is a risk state (such as abnormal air quality, abnormal patient activity, etc.) in the ward. Once a high-confidence risk state is matched, the corresponding scene is activated and the initial instruction is automatically triggered, such as adjusting the air conditioning system, notifying the nursing staff or starting the alarm mechanism, to realize intelligent response and intervention to the patient environment.
[0043] In the specific scene of financial technology, the particle filtering driven scene activation and action instruction generation method can be used in abnormal transaction identification or intelligent service triggering system. By analyzing the current transaction behavior state of the user and the market environment, the user behavior particle set is generated and its state is updated in real time. The particle weight is adjusted according to the matching degree of user operation and market feedback, and it is judged whether the preset risk triggering condition is reached. Once the activation condition is reached, it can be identified as a potential risk transaction scene, and the corresponding initial action instruction is automatically triggered, such as temporarily freezing the account, pushing the verification request or adjusting the risk control strategy, so as to realize the active prevention and rapid response to abnormal behavior.
[0044] In the embodiment of the application, the scene activation is performed on the scene real-time data according to the current state of the environment, and the initial action instruction is generated according to the activated scene, which comprises: generating an initial particle set according to the current state of the environment; obtaining a state transition function, updating the state of each particle in the initial particle set according to the scene real-time data and the state transition function, and obtaining an updated particle state; obtaining sensor real-time data, sequentially selecting a particle in the initial particle set as a target particle, and determining the data state matching degree between the updated particle state of the target particle and the sensor real-time data; determining the observation probability of the target particle according to the data state matching degree; obtaining the initial weight of each particle in the initial particle set, updating the initial weight by using the observation probability, and obtaining an updated particle weight; performing particle screening on the initial particle set according to the updated particle weight, and collecting the screened particles into an updated particle set; judging whether the updated particle set meets the preset activation condition; if the updated particle set does not meet the preset activation condition, the updated particle set is updated and screened again; if the updated particle set meets the preset activation condition, the target scene is taken as an activated scene; performing action instruction matching on the activated scene to obtain an initial action instruction.
[0045] In detail, after obtaining the current state of the environment of the target scene, a set of state samples with small perturbations or random distribution is constructed by taking the current state of the environment as the center, through the set state variable dimension and value range, forming an initial particle set. Each particle represents a possible environment state hypothesis, which is used to simulate the state evolution trend of the current scene under different conditions, so as to provide diversity basis for subsequent particle state updating and weight adjustment, and improve the modeling ability and identification accuracy of complex dynamic environment.
[0046] The state transition function is established according to historical data analysis or expert rules, and is used to depict the evolution law of the environment state under different conditions. The state transition function is combined with the real-time data of the current scene to update the state of each particle in the initial particle set, that is, the possible state of the next moment is calculated according to the current state and the environment input, so as to obtain a set of updated particle states reflecting the dynamic change trend of the current environment, thereby laying a foundation for subsequent observation matching and weight adjustment.
[0047] Real-time data of sensors (such as temperature, humidity, device state, user behavior, etc.) are collected, and particles in the initial particle set are selected one by one as target particles. The updated state of each target particle is evaluated, and the matching degree between the updated particle state and the sensor observation data is quantified by calculating the difference, such as Euclidean distance, Mahalanobis distance or other similarity measurement indicators, so as to judge the credibility of the particle under the current observation, and provide a basis for observation probability calculation.
[0048] According to the matching degree between the target particle and the real-time data of the sensor, the matching degree is converted into the observation probability by using a preset probability mapping function (such as a Gaussian distribution function). The higher the matching degree, the closer the particle state is to the actual observation, and the larger the observation probability. This observation probability is used to measure the credibility of the particle under the current environment, and provides a quantitative basis for the update of the particle weight, so as to ensure that the system can more accurately identify the real state trend of the environment.
[0049] The initial weight of each initial particle is obtained, which is usually a uniformly distributed or weight value set based on prior knowledge. Combined with the observation probability corresponding to each particle, the initial weight is corrected by the Bayesian update rule, that is, the observation probability is multiplied by the original weight as a likelihood item and is normalized to obtain the updated weight of each particle, which reflects the explanation ability of the particle to the environment state under the current observation condition, and provides a basis for particle screening and scene judgment.
[0050] According to the updated particle weight, a resampling mechanism is used to screen the initial particle set, and particles with higher weight and better matching with the current observation are preferentially retained, and the number of particles with lower weight is discarded or reduced, thereby forming a new updated particle set. By copying high-weight particles and suppressing low-weight particles, the focus and optimization of the particle set are realized, so that the newly generated particle set more accurately reflects the real distribution of the current environment state, and the reliability of subsequent judgment and decision is improved.
[0051] After the particle screening is completed and the updated particle set is generated, it is judged whether the updated particle set meets preset activation conditions, such as particle distribution concentration, a state proportion threshold or a confidence upper limit; if the conditions are not met, it indicates that the current state is not stable or does not have sufficient judgment basis, and the state update and resampling of the updated particle set will be continued; if the activation conditions are met, it is considered that the environment state is clear enough, the current target scene is marked as an activated scene, and the corresponding initial action instruction is matched based on the feature state of the activated scene and the predefined action mapping rule, to provide an instruction basis for intelligent control system execution. Specifically, when a user triggers a mode instruction (such as “reading mode”), evidence reasoning is performed through a particle filtering algorithm: input: current observation data , historical state sequence , mode activation threshold (such as heart rate > 100bpm), output: optimal state, and initial action instruction is generated according to the optimal state.
[0052] The scene activation and action instruction generation method based on particle filtering can continuously optimize the recognition accuracy of the environment state in a complex scene with uncertainty and real-time coexistence, and judge whether to trigger scene activation through a high-confidence particle set, to effectively avoid false activation and delayed response; at the same time, accurate action instruction matching is performed based on the activated scene, to ensure that the system response has context adaptability and intelligent execution capability, and significantly improves the decision efficiency and intelligence level of the environment control or service system.
[0053] S4, obtain the state space, action space and reward function of the target scene, dynamically adjust the action priority of the action space by using the state space and the reward function, and optimize the initial action instruction according to the adjusted action priority to obtain a target action instruction.
[0054] In the embodiment of the application, the reward value of each action in the action space is calculated based on the current state space and the reward function, all actions are sorted in descending order of the reward value, and the corresponding action priority is allocated accordingly, the matching degree between the initial action instruction and the action priority is analyzed, the target action instruction that meets the priority order and has better effect is generated according to the matching degree, and the dynamic optimization and accurate control of the action strategy are realized.
[0055] In the specific scenario of medical health, the state space can represent the current health status of the patient (such as body temperature, heart rate, blood pressure), environmental parameters (such as ward temperature and humidity), etc., and the action space includes adjusting the ward equipment, pushing health suggestions or notifying medical staff, etc. The reward function is used to evaluate the degree of improvement of patient comfort or rehabilitation effect of each action. By dynamically evaluating the priority of the action, the system can intelligently optimize the intervention strategy to achieve safer and personalized health management.
[0056] In the specific scenario of financial technology, the state space can describe the user account behavior state, market fluctuation trend and risk indicator, etc., and the action space includes transaction interception, risk control warning, identity verification, etc. The reward function measures the positive influence of the action on risk reduction, user experience maintenance or compliance. By real-time evaluation of these factors, the priority of the risk control measures is dynamically adjusted, effectively improving the timeliness and accuracy of abnormal transaction identification, and enhancing the safety and intelligence of the financial system.
[0057] In the embodiment of the present application, the dynamic adjustment of the action priority of the action space by using the state space and the reward function, and the optimization of the initial action instruction according to the adjusted action priority to obtain the target action instruction, comprises: According to the state space and the reward function, the reward value of each action in the action space is determined; According to the reward value, all actions in the action space are sorted, and the action priority is allocated to the sorted action space from high to low, to obtain the action priority of each action in the action space; The matching degree of the initial action instruction and the action priority is analyzed to obtain the instruction action matching degree; It is judged whether the instruction action matching degree is greater than a preset matching degree threshold; If the instruction action matching degree is greater than the matching degree threshold, the initial action instruction is taken as the target action instruction; If the instruction action matching degree is less than or equal to the matching degree threshold, the initial action instruction is adjusted according to the action priority to obtain the target action instruction.
[0058] In detail, the state space refers to the set of all possible environmental states in the target scenario, covering various variables and parameters that affect system decision-making, such as temperature, humidity, device status or user behavior, etc., which is used to comprehensively describe the specific situation of the current scenario. Specifically, the state space contains 12-dimensional parameters, and the calculation formula is as follows:
[0059] wherein, environmental parameters (temperature, humidity, light, user location), device real-time energy consumption (collecting device power data through Zigbee), user preferences (historical operation records converted into a weighted vector, such as [lighting preferences, temperature preferences]), device status (working mode, fault status, remaining life).
[0060] Action space: refers to the set of all possible actions that the system can perform in the target scenario, including 10 types of device control actions such as turning on devices, adjusting temperature, sending reminders, etc., representing all behavior choices that the system can take.
[0061] Reward function: a mapping function that evaluates the feedback value obtained by the system after performing an action, usually based on the impact of the action on the target achievement, guiding the system to optimize action selection to achieve optimal control effect. Specifically, it includes three elements: comprehensive comfort, energy consumption, and user satisfaction. The reward function formula is as follows:
[0062]
[0063] where, represents the comprehensive comfort, represents the real-time energy consumption, represents the user satisfaction, represents the weight of comprehensive comfort, represents the weight of real-time energy consumption, represents the weight of user satisfaction.
[0064] Based on the information contained in the current state space, such as environmental state, user behavior or system parameters, combined with the pre-set reward function, the execution effect of each selectable action in the action space under the current state is evaluated. The reward function will quantify the positive or negative impact of each action on the system target (such as patient comfort improvement, risk reduction, energy consumption reduction, etc.), calculate the corresponding reward value, and provide a basis for subsequent action sorting and priority allocation.
[0065] According to the reward value corresponding to each action, all selectable actions in the action space are sorted, and the higher the reward value, the more beneficial it is to achieve the system target under the current state. After sorting, the action priority is allocated from high to low according to the reward value, ensuring that the optimal action is given the highest priority, and the sub-optimal action is lowered in turn, thereby building a dynamic priority system reflecting the value of action under the current state, providing a clear basis for action instruction optimization and strategy decision-making.
[0066] The operations in the initial action instruction are compared and analyzed with the current sorted action priority, the distribution of the actions contained in the instruction in the priority sorting is evaluated, whether the initial action contains high-priority actions, low-priority actions, or the overall priority weighted score is calculated, and the consistency of the initial instruction and the optimal action sorting is quantified, and a matching degree index indicating the rationality and adaptability is generated, which provides a basis for judging whether optimization is needed.
[0067] The calculated matching degree is compared with the preset matching degree threshold: if the matching degree is higher than the threshold, it means that the initial action instruction is basically consistent with the optimal action sorting under the current state, and the instruction is directly adopted as the target action instruction; if the matching degree is lower than or equal to the threshold, it means that there are many low-priority or non-optimal actions in the initial instruction, and the sorted action priority is referred to to adjust and replace the initial instruction, retain high-priority actions, and eliminate or modify low-priority actions, so as to optimize the target action instruction that meets the current scene requirements.
[0068] Based on the state space and reward function driven action priority dynamic adjustment mechanism, the intelligent optimization of the initial action instruction can be realized, and the system can always preferentially execute the operation instruction with the highest value and the best effect under different states. Through the introduction of matching degree evaluation and threshold judgment, not only the rationality and response efficiency of action selection are improved, but also the system can dynamically adapt to environmental changes and user needs, avoid the execution of inefficient or redundant instructions, and enhance the decision accuracy, flexibility and overall intelligence level of the system.
[0069] S5, performing the target action instruction on the target device in the target scene, and performing conflict detection and resolution on the target device in the execution process to generate a collaborative decision report.
[0070] In the embodiment of the application, the target action instruction is pushed to the corresponding target device end, and the behavior data returned by the device end is monitored in real time to obtain device behavior monitoring data. According to a preset conflict type model, the data is subjected to conflict detection to determine whether there are instruction conflicts, resource preemption, state inconsistency and the like. The conflict actions are sequentially optimized and adjusted according to the action priority to form a new execution order, and the conflict behaviors are coordinated and processed according to the order to complete conflict resolution. A collaborative decision report is generated based on the conflict detection, adjustment and execution process.
[0071] In the specific scenario of medical health, the device execution and conflict resolution mechanism can be applied to intelligent ward management. For example, when the system issues multiple device control instructions (such as adjusting the bed angle, starting the air purifier, and dimming the lighting), the device response process is monitored in real time. If it is found that there is an operation conflict between devices (such as the bed adjustment and the patient moving device occupying the electric track at the same time), the system can adjust the execution order according to the medical operation priority, coordinate the device operation, ensure medical safety and patient comfort, and generate a collaborative decision report for medical staff to review and trace.
[0072] In the specific scenario of financial technology, the device execution and conflict resolution mechanism can be used in automated operations or intelligent risk control systems. For example, when the system executes multiple automatic transactions, risk control freezing, customer notification, and other instructions, strategy conflicts or resource competition may occur (such as the same account triggering risk control and marketing strategies at the same time). By monitoring the execution behavior in real time, identifying abnormal or conflict situations, and dynamically sorting and adjusting conflicting instructions according to strategy priority (such as risk control over marketing), process decoupling and strategy unification are achieved, and a collaborative report is generated for audit tracking and strategy optimization.
[0073] In an embodiment of the present application, the target action instruction is executed on the target device in the target scene, and conflict detection and resolution are performed on the target device during execution to generate a collaborative decision report, comprising: pushing the target action instruction to a preset target device end; monitoring the device behavior data returned by the preset target device end in real time to obtain device behavior monitoring data; obtaining a conflict type, detecting conflicts in the device behavior monitoring data according to the conflict type, and obtaining a conflict detection result; determining whether there is a conflict in the device behavior monitoring data according to the conflict detection result; if there is no conflict in the device behavior monitoring data, continue to monitor the device behavior data returned by the preset target device end in real time; if there is a conflict in the device behavior monitoring data, extract the conflict behavior data in the device behavior monitoring data, and take the device corresponding to the conflict behavior data as a target conflict device group; extracting the execution action of the target conflict device group from the conflict behavior data; adjusting the execution order of the execution action according to the action priority to obtain the execution order of the execution action; resolving the conflict of the target conflict device group using the execution order to obtain updated behavior data; generating a collaborative decision report according to the conflict resolution process and the updated behavior data.
[0074] In detail, the optimized target action instructions are pushed to the corresponding preset target device end, such as air conditioners, lights, curtains, air purifiers, etc. through a communication interface. After receiving the instructions, the device starts to execute and returns its execution status, response results and running parameters in real time. The system continuously monitors the running state and generates device behavior monitoring data to provide data support for subsequent conflict detection and control optimization.
[0075] The system compares and analyzes the device behavior monitoring data collected in real time according to predefined conflict types (such as resource preemption conflict, state logic conflict, timing execution conflict, etc.). By identifying the behaviors of multiple devices in the same time period, such as executing mutually exclusive operations, violating the dependency order or occupying the same resource, the system determines whether there is a potential conflict and outputs the corresponding conflict detection results to provide clear basis for subsequent conflict resolution.
[0076] If no conflict is found, the system continues to monitor the behavior data returned by the preset target device end in real time to ensure normal operation of the device. If a conflict is detected, the system extracts specific behavior information related to the conflict from the monitoring data and identifies the devices that cause the conflict as the target conflict device group to prepare for subsequent conflict resolution and coordination.
[0077] From the conflict behavior data, the system extracts the specific actions currently being executed or about to be executed by the target conflict device group, reorders these actions according to the predetermined action priority, ensures that high-priority actions are executed first and low-priority actions are executed later, and forms a reasonable execution order to avoid conflicts between devices and improve overall coordination efficiency.
[0078] According to the adjusted execution order, the system coordinates each action of the target conflict device group in turn to avoid resource contention or state conflict caused by simultaneous execution, resolves the conflict through strategies such as delaying, replacing or merging actions, and finally generates new updated behavior data reflecting the device behavior after the conflict is resolved to ensure the stability and efficiency of device coordination.
[0079] Based on the action adjustment and coordination during the conflict resolution process, the system combines the updated device behavior data, comprehensively summarizes the types, processing methods and results of conflicts, generates a detailed coordination decision report, records the specific performance and resolution measures of conflicts between devices, reflects the effect and optimization suggestions of the system on conflict processing, and provides important reference for subsequent system tuning and operation and maintenance management.
[0080] By monitoring the execution state of the device in real time, discovering and effectively solving the action conflict between devices in time, avoiding device abnormalities or functional failures caused by operation conflicts; through the order adjustment based on the action priority, the optimal scheduling of the collaborative work of the device is realized, and the stability and response efficiency of the overall operation of the system are improved; at the same time, a detailed collaborative decision report is generated, which is helpful for subsequent system optimization and operation and maintenance management, and significantly enhances the reliability and user experience of the intelligent device system.
[0081] S6, preference analysis is performed on the target user by using the user state data, and the target device is controlled in combination with the user preference behavior obtained by analysis, the target action instruction and the collaborative decision report.
[0082] In the embodiment of the application, the historical behavior data of the target user is acquired, the historical preference features thereof are extracted, the physiological state and the work and rest regularity of the user are analyzed, the dynamic preference analysis is performed on the real-time data corresponding to the target time stamp, and finally the user preference behavior is summarized. In combination with the user preference behavior vector and the target action instruction, an adjustment control strategy is generated to preliminarily control the device, an initial device state is obtained, the device state is monitored and dynamically adjusted according to the collaborative decision report, the operation of the device is ensured to meet the user's individual preference and to be efficient, and finally the intelligent and personalized environment control is realized.
[0083] In the medical health specific scene, by analyzing the historical health behavior data, physiological state and work and rest regularity of the patient, in combination with the real-time monitoring of the ward environment data, the comfort preference and activity habit of the patient in different time periods are dynamically identified, and then the control strategy of the ward device (such as air conditioner, lighting, bed angle, etc.) is adjusted, the personalized environment adjustment is realized, and the treatment experience and rehabilitation effect of the patient are improved.
[0084] In the financial technology specific scene, based on the user's past transaction behavior, account state and active time regularity, in combination with the real-time market quotation and risk control index, the user's risk preference and operation habit are dynamically analyzed, and then the risk control strategy and transaction permission control are adjusted, the potential risk operation is accurately identified and personalized intervention is realized, and the account safety is ensured while the user service experience is optimized.
[0085] In the embodiment of the application, the user state data is used to analyze the preference of the target user, comprising: Acquiring user historical behavior data, extracting historical preference features of the user historical behavior data; extracting the user physiological state and work and rest regularity of the user state data; acquire a target timestamp of the user physiological state and scene real-time data of the target timestamp, perform preference dynamic analysis on the user physiological state and the scene real-time data of the target timestamp to obtain a user preference environment state; generate a user activity preference time according to the work-rest rule; integrate the historical preference features, the user preference environment state and the user activity preference time into a user preference behavior.
[0086] In detail, by collecting historical behavior data of the user in a target scene, such as operation records, device usage frequency and time period distribution, and using data mining and feature extraction techniques to identify long-term stable behavior patterns and preference features of the user, a quantitative description of the historical preferences of the user is formed, providing basic support for subsequent personalized analysis and strategy formulation.
[0087] From the user state data, key indicators related to physiological health are extracted, such as heart rate, blood pressure, sleep duration and quality, and the daily work-rest time rule of the user is analyzed, including wake-up time, activity peak period and rest period, and by fusing these information, the physiological state and life rhythm of the user are accurately described, providing real-time basis for dynamic preference analysis.
[0088] Determine the target timestamp corresponding to the user physiological state, and synchronously acquire scene real-time data (such as environmental temperature, humidity, illumination, etc.) at the time point, by combining the physiological indicator changes of the user with the environmental conditions, use a dynamic analysis model to evaluate the preference changes of the user to the environment in this time period, and then generate a preference environment state reflecting the current needs and comfort of the user, to realize accurate response to the dynamic needs of the user.
[0089] Based on the work-rest rule data of the user, analyze the time distribution and high-frequency activity period of the daily activities of the user, determine the activity preference of the user in different time periods through statistical and pattern recognition techniques, generate time windows reflecting the active period and rest period of the user as the user activity preference time, to guide the regulation and service strategy time arrangement of intelligent devices.
[0090] Integrate the historical preference features of the user, the dynamically generated preference environment state and the activity preference time extracted based on the work-rest rule, adopt multi-dimensional data fusion and weighting method to form a unified user preference behavior model, comprehensively reflect the individualized needs and behavior habits of the user under different time and environmental conditions, and provide accurate basis for subsequent intelligent control and personalized service.
[0091] Figure 3 A flowchart of a device control process in a multi-agent collaborative intelligent device control method provided by an embodiment of the present application.
[0092] In the embodiments of the present application, the user preference behavior obtained through the binding analysis, the target action instruction and the collaborative decision report control the target device, comprising: The user preference behavior is converted into a user preference behavior vector; An adjustment control strategy is generated according to the user preference behavior vector and the target action instruction; The target device is controlled according to the adjustment control strategy, and an initial device state is obtained; The initial device state of the target device is monitored and adjusted in real time according to the collaborative decision report, and a target device state is obtained.
[0093] In detail, through feature coding technology, various data (such as historical preferences, environmental status, activity time, etc.) in the summarized user preference behavior are converted into unified numerical representation, and a user preference behavior vector in the form of a multi-dimensional vector is constructed, which can quantify and express the preference intensity and feature distribution of the user in different dimensions, facilitating subsequent model calculation and accurate matching of control strategies.
[0094] The user preference behavior vector and the target action instruction are fused and analyzed, the consistency and adaptability of the action instruction and the user preference are evaluated in combination with a preset control rule or a machine learning model, the parameters or priority of the action instruction are dynamically adjusted, an individualized and user-preference-compliant adjustment control strategy is generated, so that more accurate and user-demand-compliant device control is realized.
[0095] According to the generated adjustment control strategy, corresponding control instructions are issued to the device to adjust the running parameters and state settings of the device; after receiving the instructions, the device starts to perform operations, synchronously collects device feedback information, and finally forms an initial device state reflecting the current device response, providing basic data for subsequent state monitoring and optimization.
[0096] Based on the feedback of the device running status and conflict processing results in the collaborative decision report, the initial state of the device is monitored in real time, and possible existing abnormalities or uncoordinated behaviors are identified; by dynamically adjusting the control parameters or action sequence, the conflicts and performance bottlenecks among devices are solved, and finally the device state is optimized and updated to obtain a stable and user-demand-compliant target device state.
[0097] By comprehensively analyzing the user's historical behavior, real-time physiological state and work-rest rules, a personalized user preference behavior model is dynamically constructed to accurately grasp the user's demand; in combination with the target action instruction and the collaborative decision report, a highly adaptive control strategy can be generated, and the device state can be monitored and adjusted in real time, which not only meets the user's individualized experience, but also guarantees the collaborative and efficient operation of the device, significantly improving the flexibility, response speed and user satisfaction of the intelligent device system.
[0098] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0099] like Figure 4 1 is a functional module diagram of a multi-agent collaborative intelligent device control device provided by one embodiment of the present invention.
[0100] In the embodiment of the present disclosure, a multi-agent collaborative intelligent device control device is provided, and the multi-agent collaborative intelligent device control device corresponds one-to-one with the multi-agent collaborative intelligent device control method of the above embodiment. Figure 4 As shown, the multi-agent collaborative intelligent device control device 100 can be installed in an electronic device. According to the functions to be implemented, the multi-agent collaborative intelligent device control device 100 includes a data acquisition module 101, a state analysis module 102, a scene activation module 103, a command adjustment module 104, a conflict detection module 105, and a device control module 106. The functional modules are described in detail as follows: The data acquisition module 101 is used to acquire the target scene of the target user and the target parameters of the target scene, and collect the real-time scene data of the target parameters and the user status data of the target user; A state analysis module 102 is configured to analyze the environment state of the target scene using the scene real-time data and the user state data through a Bayesian network to obtain the current state of the environment; A scene activation module 103 is configured to activate the scene real-time data according to the current state of the environment and generate an initial action instruction according to the activated scene; An instruction adjustment module 104 is configured to obtain a state space, an action space, and a reward function of the target scene, dynamically adjust the action priority of the action space using the state space and the reward function, and optimize the initial action instruction according to the adjusted action priority to obtain a target action instruction; The conflict detection module 105 is used to execute the target action instruction on the target device in the target scene, perform conflict detection and resolution on the target device during the execution process, and generate a collaborative decision report; The device control module 106 is used to perform preference analysis on the target user using the user status data, and control the target device in combination with the user preference behavior obtained by the analysis, the target action instruction and the collaborative decision report.
[0101] In an embodiment, the state analysis module 102, when performing analysis on the environment state of the target scene by using the scene real-time data and the user state data through a Bayesian network, comprises: normalizing the scene real-time data and the user state data respectively to obtain scene standard data and user state standard data; extracting environment state parameters and device state parameters of the scene standard data, and user state parameters of the user state standard data; taking the environment state parameters, the device state parameters and the user state parameters as state nodes; taking the scene standard data and the user state standard data as observation nodes; obtaining node dependency relationships, connecting the state nodes and the observation nodes by using the node dependency relationships to construct a Bayesian network; obtaining a transition probability matrix of the Bayesian network, and generating a conditional probability table of each node by using the transition probability matrix; determining the probability distribution of the current environment state of the target scene by using the conditional probability table; determining the current environment state of the target scene according to the probability distribution.
[0102] In an embodiment, the scene activation module 103, when performing scene activation on the scene real-time data according to the current environment state, and generating an initial action instruction according to the activated scene, comprises: generating an initial particle set according to the current environment state; obtaining a state transition function, updating the state of each particle in the initial particle set according to the scene real-time data and the state transition function to obtain an updated particle state; obtaining sensor real-time data, sequentially selecting a particle in the initial particle set as a target particle, and determining the data state matching degree between the updated particle state of the target particle and the sensor real-time data; determining the observation probability of the target particle according to the data state matching degree; obtaining the initial weight of each particle in the initial particle set, updating the initial weight by using the observation probability to obtain an updated particle weight; performing particle screening on the initial particle set according to the updated particle weight, and collecting the screened particles into an updated particle set; determining whether the updated particle set meets a preset activation condition; if the updated particle set does not meet the preset activation condition, updating and screening the updated particle set again; If the update particle set meets a preset activation condition, the target scene is taken as an activation scene. An action instruction matching is performed on the activation scene to obtain an initial action instruction.
[0103] In an embodiment, when performing dynamic adjustment of action priorities of the action space by using the state space and the reward function, and optimizing the initial action instruction according to the adjusted action priorities to obtain a target action instruction, the instruction adjustment module 104 comprises: According to the state space and the reward function, a reward value of each action in the action space is determined. According to the reward value, all actions in the action space are sorted, and the sorted action space is assigned with action priorities from high to low to obtain an action priority of each action in the action space. An instruction action matching degree is analyzed according to a matching degree of the initial action instruction and the action priority to obtain an instruction action matching degree. It is judged whether the instruction action matching degree is greater than a preset matching degree threshold. If the instruction action matching degree is greater than the matching degree threshold, the initial action instruction is taken as a target action instruction. If the instruction action matching degree is less than or equal to the matching degree threshold, the initial action instruction is adjusted according to the action priority to obtain a target action instruction.
[0104] In an embodiment, when performing the target action instruction on the target device in the target scene, and performing conflict detection and resolution on the target device in the execution process to generate a collaborative decision report, the conflict detection module 105 comprises: The target action instruction is pushed to a preset target device end. Device behavior monitoring data is obtained by monitoring device behavior data returned by the preset target device end in real time. A conflict type is obtained, and the device behavior monitoring data is detected according to the conflict type to obtain a conflict detection result. It is judged according to the conflict detection result whether there is a conflict in the device behavior monitoring data. If there is no conflict in the device behavior monitoring data, real-time monitoring of the device behavior data returned by the preset target device end is continued. If there is a conflict in the device behavior monitoring data, a conflict behavior data in the device behavior monitoring data is extracted, and a device corresponding to the conflict behavior data is taken as a target conflict device group. An execution action of the target conflict device group is extracted from the conflict behavior data. sequentially adjust the execution actions according to the action priorities, to obtain an execution order of the execution actions; resolve the conflict behavior data of the target conflict device group by using the execution order, to obtain updated behavior data; generate a collaborative decision report according to the conflict resolution process and the updated behavior data.
[0105] In an embodiment, the device control module 106, when performing preference analysis on the target user by using the user state data, includes: obtain user historical behavior data, and extract historical preference features of the user historical behavior data; extract a user physiological state and a work-rest schedule of the user state data; obtain a target timestamp of the user physiological state and scene real-time data of the target timestamp, and perform preference dynamic analysis on the user physiological state and the scene real-time data of the target timestamp, to obtain a user preference environment state; generate a user activity preference time according to the work-rest schedule; summarize the historical preference features, the user preference environment state, and the user activity preference time into a user preference behavior.
[0106] In an embodiment, the device control module 106, when performing control on the target device by combining the user preference behavior obtained through analysis, the target action instruction, and the collaborative decision report, includes: convert the user preference behavior into a user preference behavior vector; generate an adjustment control strategy according to the user preference behavior vector and the target action instruction; control the target device according to the adjustment control strategy, to obtain an initial device state; monitor and adjust the initial device state of the target device in real time according to the collaborative decision report, to obtain a target device state.
[0107] In the present application, for a multi-agent collaborative intelligent device control device, first, the present application obtains the target scene of the target user and the target parameters of the target scene, collects the scene real-time data of the target parameters and the user state data of the target user, analyzes the environment state of the target scene by using the scene real-time data and the user state data through the Bayesian network, obtains the current environment state, the Bayesian network analysis method fuses the scene real-time data and the user state data, constructs a structured causal reasoning model, which can not only realize dynamic perception and uncertainty modeling of the environment state, but also maintain high reasoning accuracy in the case of noise or partial missing of multi-source data, activates the scene according to the scene real-time data of the environment current state, and generates an initial action instruction according to the activated scene. The scene activation and action instruction generation method based on particle filtering can continuously optimize the recognition accuracy of the environment state in the complex scene where uncertainty and real-time coexist through multi-particle simulation and dynamic iteration of the current environment state. Then, the state space, action space and reward function of the target scene are obtained, the action priority of the action space is dynamically adjusted by using the state space and the reward function, and the initial action instruction is optimized according to the adjusted action priority to obtain a target action instruction. The action priority dynamic adjustment mechanism driven by the state space and the reward function can realize intelligent optimization of the initial action instruction, ensure that the system always executes the operation instruction with the highest value and the best effect under different states, execute the target action instruction on the target device in the target scene, and detect and eliminate the target device in the execution process. Conflict detection and resolution generate a collaborative decision report. Through real-time monitoring of the execution state of the device, the action conflict between devices is discovered and effectively solved in time, avoiding device abnormalities or function failures caused by operation conflicts. Through the sequence adjustment based on the action priority, the optimization scheduling of device collaborative work is realized, and the stability and response efficiency of the overall operation of the system are improved. Finally, the user state data is used to analyze the preferences of the target user, and the target device is controlled in combination with the user preference behavior obtained by the analysis, the target action instruction and the collaborative decision report, which meets the user's individual experience and ensures the efficient operation of the device collaboration, improves the coordination mechanism between devices, the global optimization capability and the response efficiency. The specific limitations of the multi-agent collaborative intelligent device control device can be referred to the limitations of the multi-agent collaborative intelligent device control method in the foregoing, which will not be repeated here. Each module in the multi-agent collaborative intelligent device control device can be realized by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations of the above-mentioned modules by the processor.
[0108] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multi-agent collaborative intelligent device control method.
[0109] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a multi-agent collaborative intelligent device control method.
[0110] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed: Acquire a target scenario of a target user and target parameters of the target scenario, and collect real-time scenario data of the target parameters and user status data of the target user; Analyzing the environmental state of the target scene using the real-time scene data and the user state data through a Bayesian network to obtain the current state of the environment; Activating the scene real-time data according to the current state of the environment, and generating initial action instructions according to the activated scene; Obtaining a state space, an action space, and a reward function of the target scene, dynamically adjusting the action priority of the action space using the state space and the reward function, and optimizing the initial action instruction according to the adjusted action priority to obtain a target action instruction; Perform the target action instruction on the target device in the target scene, and perform conflict detection and resolution on the target device during the execution process to generate a collaborative decision report; Perform preference analysis on the target user by using the user state data, and control the target device in combination with the user preference behavior obtained through the analysis, the target action instruction, and the collaborative decision report.
[0111] In several embodiments provided by the present application, it should be understood that the disclosed devices and apparatuses can be implemented in other manners. For example, the above described system embodiments are merely illustrative. For example, the division of the modules is merely logical function division. In actual implementation, another division manner can be adopted.
[0112] In addition, the various function modules in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of hardware plus software function modules.
[0113] Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and therefore all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0114] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.
[0115] In some embodiments of the present embodiment, a computer readable storage medium is provided, and a computer program is stored on the computer readable storage medium. The computer program is characterized in that when executed by a processor, the computer program implements the steps of the method described in the above embodiments.
[0116] The readable storage medium of the present application stores a computer program, and the computer program can achieve the following when executed by a processor of an electronic device: Obtain a target scene of a target user and a target parameter of the target scene, and collect scene real-time data of the target parameter and user state data of the target user; Analyze the environment state of the target scene by using the scene real-time data and the user state data through a Bayesian network to obtain a current environment state; According to the current state of the environment, the scene activation is performed on the scene real-time data, and initial action instructions are generated according to the activated scene; The state space, the action space and the reward function of the target scene are acquired, the action priority of the action space is dynamically adjusted by using the state space and the reward function, the initial action instructions are optimized according to the adjusted action priority, and target action instructions are obtained; The target action instructions are executed on the target device in the target scene, and the target device in the execution process is subjected to conflict detection and resolution, and a collaborative decision report is generated; The target user is subjected to preference analysis by using the user state data, and the target device is controlled in combination with the user preference behavior obtained through the analysis, the target action instructions and the collaborative decision report.
[0117] It should be noted that the functions or steps described above with respect to the computer-readable storage medium or the computer device can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0118] The computer-readable storage medium can also store at least one computer executable program / instruction, such as computer readable instructions. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may, for example, include read-only memory (ROM), hard disk, flash memory, etc. For example, the non-transitory computer-readable storage medium can be connected to a computing device such as a computer, and then when the computing device runs the computer readable instructions stored on the computer readable storage medium, the various methods described above can be performed.
[0119] In addition, the computer device can also include (but not limited to) a data bus, an input / output (I / O) bus, a display, and an input / output device (such as a keyboard, a mouse, a speaker, etc.), etc.
[0120] The processor can communicate with external devices through the I / O bus via wired or wireless networks.
[0121] In one embodiment, the at least one computer executable instruction can also be compiled into or constitute a software product / computer program product, wherein one or more computer executable instructions are executed by the processor to perform the steps of the various functions and / or methods described in the embodiments of the present technology.
[0122] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0123] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0124] In the embodiments provided by the present disclosure, it should be understood that the disclosed apparatus and method can also be implemented in other manners. The embodiments described above are merely exemplary for describing the present disclosure. For example, the flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operation of the apparatus, method and computer program product according to the embodiments of the present disclosure. In this regard, each block in the flowcharts and block diagrams can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logic function. It should also be noted that, in some alternative implementations, the functions noted in the blocks can occur in different orders from those noted in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a special-purpose hardware-based system for implementing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0125] It should be noted that, in the present disclosure, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element limited by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0126] The above-described embodiments are merely used to illustrate the technical solutions of the present disclosure, rather than limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent replacements; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure.
[0127] It should be noted that, in the embodiments of the present disclosure, if non-company software tools or components appear, they are only used for example introduction, and do not represent actual use.
Claims
1. A multi-agent collaborative intelligent device control method, characterized in that: The method comprises: Acquire a target scenario of a target user and target parameters of the target scenario, and collect real-time scenario data of the target parameters and user status data of the target user; Analyzing the environmental state of the target scene using the real-time scene data and the user state data through a Bayesian network to obtain the current state of the environment; Activating the scene real-time data according to the current state of the environment, and generating initial action instructions according to the activated scene; Obtaining a state space, an action space, and a reward function of the target scene, dynamically adjusting the action priority of the action space using the state space and the reward function, and optimizing the initial action instruction according to the adjusted action priority to obtain a target action instruction; Executing the target action instruction on the target device in the target scene, performing conflict detection and resolution on the target device during the execution process, and generating a collaborative decision report; The user status data is used to perform a preference analysis on the target user, and the target device is controlled in combination with the user preference behavior obtained from the analysis, the target action instruction and the collaborative decision report.
2. The multi-agent collaborative intelligent device control method according to claim 1, characterized in that: The method of analyzing the environment state of the target scene using the real-time scene data and the user state data through a Bayesian network to obtain the current state of the environment includes: Normalizing the scene real-time data and the user status data to obtain scene standard data and user status standard data; Extracting the environment state parameters and the device state parameters of the scene standard data, and the user state parameters of the user state standard data; Taking the environment state parameter, the device state parameter and the user state parameter as state nodes; Using the scene standard data and the user state standard data as observation nodes; Obtaining node dependencies, connecting the state nodes and the observation nodes using the node dependencies, and constructing a Bayesian network; Obtaining a transition probability matrix of the Bayesian network, and generating a conditional probability table for each node using the transition probability matrix; Determining the probability distribution of the current environmental state of the target scene using the conditional probability table; According to the probability distribution, a current state of the environment of the target scene is determined.
3. The multi-agent collaborative intelligent device control method according to claim 1, characterized in that: The step of activating the scene real-time data according to the current state of the environment and generating an initial action instruction according to the activated scene includes: generating an initial set of particles according to the current state of the environment; Obtaining a state transfer function, and updating the state of each particle in the initial particle set according to the scene real-time data and the state transfer function to obtain an updated particle state; Acquire real-time sensor data, select one particle from the initial particle set as a target particle in turn, and determine a data state matching degree between an updated particle state of the target particle and the real-time sensor data; Determining the observation probability of the target particle according to the data state matching degree; Obtaining an initial weight of each particle in the initial particle set, and updating the initial weight using the observation probability to obtain an updated particle weight; Perform particle screening on the initial particle set according to the updated particle weight, and aggregate the screened particles into an updated particle set; Determining whether the updated particle set meets a preset activation condition; If the updated particle set does not meet the preset activation condition, the updated particle set is updated and screened again; If the updated particle set meets the preset activation condition, the target scene is used as the activation scene; Matching action instructions to the activation scene to obtain initial action instructions.
4. The multi-agent collaborative intelligent device control method according to claim 1, characterized in that: The dynamically adjusting the action priority of the action space by using the state space and the reward function, and optimizing the initial action instruction according to the adjusted action priority to obtain a target action instruction, includes: Determining a reward value for each action in the action space according to the state space and the reward function; sorting all actions in the action space according to the reward value, assigning action priorities to the sorted action space from high to low, and obtaining the action priority of each action in the action space; Analyzing the matching degree between the initial action instruction and the action priority to obtain the instruction-action matching degree; Determining whether the command action matching degree is greater than a preset matching degree threshold; If the command action matching degree is greater than the matching degree threshold, the initial action command is used as the target action command; If the instruction-action matching degree is less than or equal to the matching degree threshold, the initial action instruction is adjusted according to the action priority to obtain a target action instruction.
5. The multi-agent collaborative intelligent device control method according to claim 1, characterized in that: The executing the target action instruction on the target device in the target scene, performing conflict detection and resolution on the target device during the execution process, and generating a collaborative decision report includes: Pushing the target action instruction to the preset target device end; Monitor the device behavior data returned by the preset target device in real time to obtain device behavior monitoring data; Obtaining a conflict type, performing conflict detection on the device behavior monitoring data according to the conflict type, and obtaining a conflict detection result; Determining whether there is a conflict in the device behavior monitoring data according to the conflict detection result; If there is no conflict in the device behavior monitoring data, continue to monitor the device behavior data returned by the preset target device end in real time; If there is a conflict in the device behavior monitoring data, extracting the conflicting behavior data in the device behavior monitoring data, and taking the devices corresponding to the conflicting behavior data as the target conflicting device group; extracting the execution action of the target conflicting device group from the conflicting behavior data; Adjusting the order of the execution actions according to the action priorities to obtain the execution order of the execution actions; Resolving conflicts on the conflicting behavior data of the target conflicting device group using the execution order to obtain updated behavior data; A collaborative decision report is generated based on the conflict resolution process and the updated behavior data.
6. The multi-agent collaborative intelligent device control method according to claim 1, characterized in that: The performing preference analysis on the target user by utilizing the user status data includes: Obtaining user historical behavior data and extracting historical preference features of the user historical behavior data; Extracting the user's physiological state and work and rest patterns from the user state data; Obtaining a target timestamp of the user's physiological state and real-time scene data of the target timestamp, performing a preference dynamic analysis on the user's physiological state and the real-time scene data of the target timestamp to obtain a user preferred environment state; Generating user activity preference time according to the work and rest pattern; The historical preference features, the user preference environment state and the user activity preference time are aggregated into user preference behavior.
7. The multi-agent collaborative intelligent device control method according to claim 1, characterized in that: The control of the target device by combining the analyzed user preference behavior, the target action instruction, and the collaborative decision report includes: Converting the user preference behavior into a user preference behavior vector; Generate an adjustment control strategy according to the user preference behavior vector and the target action instruction; Controlling the target device according to the adjustment control strategy to obtain an initial device state; The initial device state of the target device is monitored and adjusted in real time according to the collaborative decision report to obtain the target device state.
8. A multi-agent collaborative intelligent device control device, characterized in that: The device comprises: A data acquisition module is used to acquire a target scenario of a target user and target parameters of the target scenario, and collect real-time scenario data of the target parameters and user status data of the target user; A state analysis module is used to analyze the environmental state of the target scene using the real-time scene data and the user state data through a Bayesian network to obtain the current state of the environment; A scene activation module, configured to activate the scene real-time data according to the current state of the environment, and generate an initial action instruction according to the activated scene; an instruction adjustment module, configured to obtain the state space, action space, and reward function of the target scene, dynamically adjust the action priority of the action space using the state space and the reward function, and optimize the initial action instruction according to the adjusted action priority to obtain a target action instruction; a conflict detection module, configured to execute the target action instruction on the target device in the target scene, perform conflict detection and resolution on the target device during the execution process, and generate a collaborative decision report; The device control module is used to use the user status data to perform preference analysis on the target user, and control the target device in combination with the user preference behavior obtained by the analysis, the target action instruction and the collaborative decision report.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a multi-agent collaborative intelligent device control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements a multi-agent collaborative intelligent device control method as described in any one of claims 1 to 7.
Citation Information
Cited By
Intelligent park operation center control management system based on artificial intelligence
CN121209270A
Multi-AI agent collaborative arbitration guest obtaining and putting method and device, medium and equipment
CN122199071A