Water-saving intelligent agent interaction control method and system based on multi-modal fusion
By employing a multimodal fusion-based water-saving intelligent agent interactive control method, and utilizing Granger causality testing and graph neural network updates, the system dynamically adapts to changes in building water usage characteristics, achieving accurate identification of water usage anomalies and generation of water-saving suggestions, thus solving the problem of static model adaptation failure.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI JICHENSHUI DIGITAL TECH CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-01
AI Technical Summary
Existing building water management systems struggle to accurately identify water usage anomalies and implement water-saving interventions when faced with varying water usage characteristics across different time periods, resulting in static model adaptation failure.
By employing a multimodal fusion-based interactive control method for water-saving intelligent agents, Granger causality tests are used to construct the temporal dependencies of multimodal data, dynamically update the graph neural network map, and generate water-saving suggestions adapted to the current time period's feature distribution.
It achieves accurate identification and adaptation of water usage characteristics at different times, ensuring the global adaptability and real-time nature of water-saving recommendations, and solving the problem of static model adaptation failure.
Smart Images

Figure CN121598326B_ABST
Abstract
Description
A Multimodal Fusion-Based Interactive Control Method and System for Water-Saving Intelligent Agents Technical Field
[0001] This invention relates to the field of water-based intelligent agent interaction technology, and more specifically to a water-saving intelligent agent interaction control method and system based on multimodal fusion. Background Technology
[0002] In water management of various buildings and public facilities such as commercial buildings, hospitals, and schools, water use behavior exhibits significant temporal dynamics. For example, people concentrate on water use during morning peak hours and activity periods, while water load drops sharply during off-peak hours. The statistical distribution of multimodal characteristics such as the number of people using water, flow rate, and business activities varies greatly at different times.
[0003] With the popularization of IoT and multimodal sensing technologies, buildings have achieved the collection of multi-dimensional data such as video (number of people using water), audio (pipeline status), flow (water meter), and business data (OA activities). However, existing solutions mostly rely on static models to analyze water usage patterns in a single time period, which is difficult to adapt to the changes in characteristic distribution in different time periods. This results in insufficient accuracy in identifying abnormal water usage and intervening in water conservation. Therefore, existing technologies have shortcomings. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention aims to provide a water-saving intelligent agent interactive control method and system based on multimodal fusion. By constructing the temporal dependency relationship of multimodal data through Granger causality test, the method enables real-time tracing of the root causes of anomalies. Furthermore, the method updates the graph neural network spectrum based on the drift results obtained from the structural causal graph, making the model adapt to the feature distribution of the current time period and completely solving the adaptation failure problem of static models.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] This invention provides a water-saving intelligent agent interaction control method based on multimodal fusion, wherein the water-saving intelligent agent includes multiple regional intelligent agents and a fusion intelligent agent, and the method includes:
[0007] Obtain a multimodal dataset, which includes video modal data, audio modal data, structural modal data, and activity modal data;
[0008] Based on the multimodal dataset and Granger causality test, a structural causality map is obtained;
[0009] The graph neural network graph is updated based on the causal graph of the structure to obtain the graph embedding vector corresponding to each regional agent, so that the fused agent can generate water-saving suggestions based on the graph embedding vector corresponding to each regional agent.
[0010] As a further improvement of the present invention, a structural causal graph is obtained based on the multimodal dataset and the Granger causality test, including:
[0011] The edge set is obtained based on historical modal data and Granger causality test;
[0012] The node set and reconstructed edge weights are obtained based on the multimodal dataset;
[0013] Based on the reconstructed edge weights and the structural causal model, the structural causal graph and graph vector are obtained.
[0014] As a further improvement of the present invention, the reconstructed edge weights are obtained based on the multimodal dataset, including:
[0015] Based on the multimodal dataset and preset thresholds, the modality to be replaced is obtained;
[0016] Based on the historical correlation matrix and the mode to be replaced, the alternative mode is obtained;
[0017] Based on the correlation coefficients corresponding to the alternative modes in the historical correlation matrix, the original graph edge weights are updated to obtain the reconstructed edge weights.
[0018] As a further improvement of the present invention, the graph neural network graph is updated based on the structural causal graph to obtain the graph embedding vector corresponding to each regional agent, including:
[0019] Based on the graph vector, the embedding dimension corresponding to the modality to be replaced is obtained, and the key dimension vector is obtained.
[0020] The drift result is obtained based on the key dimension vector and the baseline distribution parameters corresponding to the current time.
[0021] Based on the structural causal graph and the drift result, the graph neural network graph is updated to obtain the graph embedding vector corresponding to each regional agent.
[0022] As a further improvement of the present invention, the drift result is obtained based on the key dimension vector and the baseline distribution parameters corresponding to the current time, including:
[0023] The KL divergence is obtained based on the baseline distribution parameters corresponding to the current time and the current distribution parameters corresponding to the key dimension vector.
[0024] If the KL divergence is greater than a preset value, the drift node is located based on mutual information analysis to obtain the drift result.
[0025] As a further improvement of the present invention, based on the structural causal graph and the drift result, the graph neural network graph is updated to obtain the graph embedding vector corresponding to each regional agent, including:
[0026] The initial weights corresponding to the graph neural network graph are obtained based on the structural causal graph.
[0027] The update object in the graph neural network graph is determined based on the drift result;
[0028] The initial weights corresponding to the updated objects in the graph neural network graph are updated based on the historical correlation matrix to obtain the graph embedding vector corresponding to each regional agent.
[0029] As a further improvement of the present invention, water-saving suggestions are generated based on the graph embedding vector corresponding to each regional agent, including:
[0030] Based on the graph embedding vector corresponding to each regional agent, the aggregate embedding vector is obtained;
[0031] Based on the aggregated embedding vector and the preset model, a global embedding vector is obtained;
[0032] Based on the global embedding vector, a global fusion graph is obtained;
[0033] Based on the drift results and the global fusion map, water-saving suggestions are generated.
[0034] As a further improvement of the present invention, the step of obtaining the aggregated embedding vector based on the graph embedding vector corresponding to each regional agent includes:
[0035] Calculate the collaboration coefficient based on the graph embedding vector corresponding to each regional agent;
[0036] Based on the aforementioned coordination coefficient, the attention weights are obtained;
[0037] The aggregated embedding vector is obtained based on the attention weights and the graph embedding vector corresponding to each region agent.
[0038] As a further improvement of the present invention, a global fusion graph is obtained based on the global embedding vector, including:
[0039] Calculate the similarity between the global embedding vector and the positive and negative samples;
[0040] The loss function is obtained based on the confidence level and similarity level corresponding to the multimodal dataset;
[0041] The global embedding vector is updated according to the loss function to obtain the global fusion graph.
[0042] This invention provides a water-saving intelligent agent interactive control system based on multimodal fusion, comprising a regional intelligent agent and a fusion intelligent agent, wherein the regional intelligent agent includes:
[0043] The acquisition module is used to acquire a multimodal dataset, which includes video modal data, audio modal data, structural modal data, and activity modal data.
[0044] The verification module is used to obtain a structural causal graph based on the multimodal dataset and the Granger causality test;
[0045] The update module is used to update the graph neural network graph based on the structural causal graph to obtain the graph embedding vector corresponding to each regional agent;
[0046] The fusion agent is used to generate water-saving suggestions based on the graph embedding vector corresponding to each regional agent.
[0047] This invention first integrates multimodal data including video, audio, structure, and activity. Then, it constructs a structural causal graph by combining Granger causality tests, clarifying the temporal dependencies between the various modalities. Based on the structural causal graph, it dynamically updates the graph neural network graph and generates graph embedding vectors for regional agents. Finally, the fusion agent generates water-saving suggestions by aggregating the graph embedding vectors of each region, solving the adaptation failure problem of static models and ensuring the global adaptability of water-saving suggestions. Attached Figure Description
[0048] Figure 1 is a schematic diagram of the method steps of the present invention;
[0049] Figure 2 is a schematic diagram of the steps to obtain the map embedding vector;
[0050] Figure 3 is a schematic diagram of the steps to obtain the aggregated embedding vector;
[0051] Figure 4 is a schematic diagram of the steps to obtain water-saving suggestions. Detailed Implementation
[0052] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof.
[0053] The term "and / or" in the following text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0054] This application provides a water-saving intelligent agent interaction control method based on multimodal fusion, including:
[0055] Obtain a multimodal dataset, which includes video modal data, audio modal data, structural modal data, and activity modal data;
[0056] Based on the multimodal dataset and Granger causality test, a structural causality map is obtained;
[0057] The graph neural network graph is updated based on the structural causal graph, and the graph embedding vector corresponding to each regional agent is obtained, so that the fusion agent can generate water-saving suggestions based on the graph embedding vector corresponding to each regional agent.
[0058] In this embodiment, the water-saving intelligent agent includes multiple regional intelligent agents and a fusion intelligent agent. Each regional intelligent agent corresponds to a region, such as a tea room. Each regional intelligent agent needs to obtain the corresponding graph embedding vector based on its current multimodal data and send the graph embedding vector to the fusion intelligent agent to obtain water-saving suggestions. The method provided in this embodiment is executed once every preset time interval. Therefore, the current multimodal data is obtained based on the data collected in the adjacent previous preset time interval. This embodiment does not limit the specific length of the preset time interval, for example, it can be set to 20 minutes.
[0059] For example, video modal data may include faucet switch status frames, audio modal data may include the spectrum of abnormal noises from pipe leaks, structural modal data may include flow rate and water temperature, and activity modal data may include the number of people active in the area. The multimodal data is preprocessed to obtain a multimodal dataset.
[0060] This embodiment first integrates multimodal data including video, audio, structure, and activity. Then, it constructs a structural causal graph using Granger causality tests to clarify the temporal dependencies between the various modalities. Based on the structural causal graph, it dynamically updates the graph neural network graph and generates graph embedding vectors for regional agents. Finally, the fusion agent generates water-saving suggestions by aggregating the graph embedding vectors of each region, thus solving the adaptation failure problem of static models and ensuring the global adaptability of water-saving suggestions.
[0061] Furthermore, this embodiment provides a step for obtaining a structural causal graph based on a multimodal dataset and Granger causality test, including:
[0062] The edge set is obtained based on historical modal data and Granger causality test;
[0063] The node set and reconstructed edge weights are obtained from the multimodal dataset;
[0064] Based on the reconstructed edge weights and the structural causal model, the structural causal graph and graph vector are obtained.
[0065] Specifically, the node set is first obtained based on the data types included in the multimodal dataset. For example, the node set N can be represented as N={V, A, S, OA}.
[0066] In the node set, V corresponds to the faucet status, A corresponds to the pipe leakage status, S corresponds to the flow rate, and OA indicates whether there is human activity in the current area. The node set also carries a label to indicate the current time period. For example, based on the faucet's on / off status within a preset time period captured by the camera, the number of times water is drawn can be identified, the maximum number of times water is drawn in history can be obtained, and the ratio of the current number of times water is drawn to the maximum number of times water is drawn is used as the value corresponding to V. Based on the pipe spectrum within a preset time period captured by the sensor, the corresponding spectral energy is calculated using Fourier transform, and the spectral energy corresponding to when the pipe leaks is obtained is obtained. The ratio of the current spectral energy to the spectral energy corresponding to when the leak occurs is used as the value corresponding to A. Based on the water meter data within the preset time period, the flow rate of the area within the preset time interval can be determined. The average flow rate is used as the ratio of the average flow rate to the maximum flow rate in historical data, and this ratio is used as the value corresponding to S. Based on the data in the OA system, it is determined whether there is any activity near the current area, and the value corresponding to OA is obtained. For example, if the tea room corresponding to the current area is on the 3rd floor, and the OA system shows that there is a meeting on the 3rd floor, then OA=1; otherwise, OA=0. It is determined whether the current time is during the peak water usage period. If so, the label is marked as value 1; otherwise, the label is marked as value 0. This embodiment does not limit the criteria for determining the peak water usage period. For example, the peak water usage period may include the lunchtime peak (11:30-12:30) caused by employees collecting water and washing cups before / after lunch break, and the temporary peak caused by attendees collecting water after large meetings or events (such as the meeting ending at 10:30).
[0067] Next, Granger causality tests are performed based on the historical modal data corresponding to the region. These tests identify whether a causal relationship exists between any two nodes in the node set. Based on the results of the Granger causality tests, edges are created between the two causally related nodes, and the direction of the edges is determined based on the causal relationship. Finally, the set of each edge is used as the edge set. For example, for V and S, the values corresponding to each preset time interval within a preset time period (e.g., 7 days) can be obtained. These values are arranged in chronological order to obtain the time-series data corresponding to these two nodes. Then, a Granger causality test is performed on the two time-series data. Since the causality test is bidirectional, two tests are required for any two nodes. For example, for V and S, the first test determines whether V is a Granger cause of S, and the second test determines whether S is a Granger cause of V. Afterward, edges are created between the two causally related nodes, and the direction of the edges is determined based on the causal relationship. For example, if V is a Granger cause of S, the direction of the edge is V→S. Furthermore, to reduce latency, edge sets do not need to be generated every time. If the scene corresponding to the current region has not changed, the edge set can be reused directly without repeated generation. If the scene corresponding to the current region changes, such as adjusting employees' working hours or changing the pipeline structure, the edge set needs to be regenerated.
[0068] Then, based on the data of each modality in the multimodal dataset, the corresponding modal confidence is calculated. The confidence is used to quantify the reliability of each modal data. For example, for video modal data, since the interference in faucet status recognition mainly comes from area occlusion, the video modal confidence needs to be determined in conjunction with the severity of occlusion. For example, for multiple frames of images collected within a preset time interval, the total area of the faucet area and the area where the person is getting water in the multiple frames of images is first determined as the first area. Then, the total area of the occluded part (such as the person's body or debris occluding the faucet area) in the area is identified as the second area. Then, the ratio of the second area to the first area is calculated. Finally, the ratio is subtracted from 1 to obtain the video modal confidence. The image recognition method is a conventional technique, which will not be described in detail in this embodiment. Audio modal analysis is used to identify pipe leaks, but ambient noise can mask the effective sound, causing the acoustic features of the leak to be covered by noise and making accurate identification impossible. Therefore, this embodiment obtains the audio modal confidence level based on the interference of noise on the audio data. Specifically, it can be expressed as 1 minus the ratio of real-time noise intensity to noise threshold. The real-time noise intensity is the ambient noise decibels read directly by the decibel detection function of the audio sensor, and the noise threshold is the maximum noise threshold at which the leak features can be identified. That is, when the real-time noise intensity exceeds this threshold, the acoustic features of the leak will be covered by noise. The noise threshold can be calibrated experimentally, and this embodiment does not limit its value or acquisition method. The reliability of structural modes (such as water meter flow) is determined by the error in the data. The smaller the error, the more reliable the data. Therefore, this embodiment uses the trace of the error variance matrix corresponding to the flow to measure the magnitude of the error. For example, the flow value corresponding to each preset time interval within a preset time can be obtained, normalized, and the covariance matrix can be calculated. The trace of the covariance matrix is added to 1, summed, and the reciprocal is taken to obtain the structural mode confidence score. The smaller the fluctuation of the flow data, the smaller the trace of the covariance matrix, and the higher the confidence score. The activity mode confidence score is used to measure the time deviation between the current time and the peak water consumption period. Specifically, it can be expressed as 1 minus the ratio of the time deviation to the time threshold. For example, assuming the current time is 8:35, the meeting time can be obtained from the OA system as 8:30-9:30. At this time, the time deviation is 5 minutes. The time threshold is the preset maximum effective deviation of the activity impact. This embodiment does not limit its value. For example, after statistics, it is found that when the time deviation exceeds 60 minutes, the impact of the OA activity on water consumption can be ignored. At this time, the time threshold can be set to 60 minutes. Then, a modality confidence list is constructed based on the confidence level corresponding to each modality data, where each modality corresponds to a node in the node set.
[0069] Next, the historical correlation matrix is obtained. The number of rows and columns in the historical correlation matrix is equal to the number of modes. The elements in the matrix represent the correlation coefficient between the two corresponding modes. The correlation coefficient can be calculated based on the historical mode data. For example, for V and S, the values corresponding to each preset time interval within a preset time (such as 7 days) can be obtained. Each value is arranged in chronological order to obtain the time series data corresponding to these two nodes. The correlation coefficient between the two time series data is calculated to obtain the corresponding element values in the historical correlation matrix. Calculating the correlation coefficient is a conventional technique, which will not be elaborated on in this embodiment. It should be noted that the historical correlation matrix needs to be updated based on the data collected in real time. That is, for the current moment, the corresponding historical correlation matrix is not calculated based on the currently collected data. The historical correlation matrix updated based on the currently collected data is used for the next execution of the method provided in this embodiment.
[0070] Furthermore, this embodiment provides a step for obtaining reconstructed edge weights based on multimodal data, including:
[0071] Based on the multimodal dataset and preset thresholds, the modality to be replaced is obtained;
[0072] Based on the historical correlation matrix and the mode to be replaced, the alternative mode is obtained;
[0073] Based on the correlation coefficients corresponding to the alternative modes in the historical correlation matrix, the edge weights of the original graph are updated to obtain the reconstructed edge weights.
[0074] Specifically, after obtaining the list of modal confidence levels, it is necessary to compare the confidence level corresponding to each modality with the preset threshold corresponding to that modality. The preset threshold represents the minimum confidence threshold value for the usable data of that modality. When the confidence level of a certain modality is lower than this value, it indicates that the data is unreliable, and at this time, the edge weights of that modality in the original graph need to be updated.
[0075] Then, modes with confidence scores below a preset threshold are selected as modes to be replaced. Next, the historical correlation matrix is searched to determine the correlation coefficients between other modes and the modes to be replaced, and the confidence scores of each other mode are obtained. For each other mode, the correlation coefficient between it and the mode to be replaced is weighted by its own confidence score, and the mode with the highest weighted value is selected as the replacement mode. The edge weights of the original graph are updated based on the replacement mode. In this embodiment, no restrictions are placed on the weights.
[0076] The original graph is determined based on the node set and edge set mentioned above. Since this is the first update of edge weights after the graph is determined, the current edge weights are equal to the correlation coefficients between the corresponding modes in the historical correlation matrix. However, at this time, they are only numerically the same. The correlation coefficient is used to measure the degree of association between modes. The edge weights in the graph are parameters of the model. With the dynamic update of real-time data, that is, with the continuous update of edge weights, the values of the two will be different. Furthermore, each subsequent update of the graph edge weights is performed based on the adjacent previous update. Specifically, after obtaining the mode to be replaced, each edge connected to this mode is determined based on the original graph, and the weight corresponding to each edge is obtained. These edge weights are the edge weights that need to be reconstructed. The other mode corresponding to these edges is also determined, i.e., each mode directly connected to the mode to be replaced by an edge is identified, denoted as the relative mode. Taking the weight of reconstructing one edge as an example, firstly, the correlation coefficient between the two modes corresponding to the edge is determined based on the historical association matrix or the original graph, and the correlation coefficient between the mode to be replaced and the relative mode corresponding to the edge is also determined. Then, the correlation coefficient between the replacement mode and the relative mode is determined, and the ratio of the two correlation coefficients is calculated. The weight corresponding to the edge is then multiplied by this ratio to obtain the reconstructed edge weight corresponding to this edge. Similarly, the reconstructed edge weights corresponding to each edge connected to the mode to be replaced can be obtained. For example, suppose the mode to be replaced corresponds to node A, the replacement mode corresponds to node V, and there is an edge A→S between A and S, i.e., S is the corresponding relative mode. The weight of the edge between A and S needs to be reconstructed. At this time, the correlation coefficient between A and S needs to be determined. And determine the correlation coefficient between V and S. The updated edge weights are , This represents the weight of the edge between A and S before reconstruction.
[0077] In this embodiment, when single-modal data is unreliable, a replacement modality is selected and the edge weights are updated using the correlation coefficient. Specifically, the purpose of updating the edge weights in this embodiment is to address the unreliability of data without altering the causal contribution logic between the two variables. The causal contribution logic can be determined by the causal contribution amount, which can be determined by the edge weights and the correlation coefficient between the two modalities. The edge weights are model parameters in the graph, representing the weight percentage of the edge's influence in the graph. The correlation coefficient represents the degree of correlation between the modalities. Multiplying the two yields the causal contribution amount, which represents the actual influence strength of a certain modality on the relative modality through the corresponding edge. Therefore, the principle behind the formula for updating the edge weights in this embodiment is to keep the causal contribution amount unchanged.
[0078] After obtaining the reconstructed edge weights, for each node, its corresponding value is used as the initial feature of that node. Then, for each node, its corresponding neighboring nodes are first determined. Neighboring nodes are nodes directly adjacent to it through an edge. Then, the initial features corresponding to the neighboring nodes are weighted according to their corresponding edge weights, and the sum of these corresponding edge weights is calculated. Finally, the ratio of the weighted value to the sum is used as the neighbor aggregated feature corresponding to that node, thus obtaining the neighbor aggregated feature corresponding to each node. Next, the initial features corresponding to the node itself and the neighbor aggregated features are concatenated to obtain the preliminary feature vector corresponding to each node. Then, the preliminary feature vector corresponding to each node is mapped through a fully connected layer to obtain the graph vector. The role of the fully connected layer is to increase the dimension through linear transformation and activation function to mine the hidden information in the preliminary feature vector. This embodiment does not limit the specific dimension, such as 62 or 128 dimensions, etc., and this embodiment does not limit the structure of the fully connected layer. For example, two fully connected layers can be set and the activation function can be set to ReLU. The graph vector integrates the features of all nodes in the graph and the causal relationships between nodes. Furthermore, based on the reconstructed edge weights, the structural causal graph can be represented as G={N (node set), E (edge set), W (weight set)}.
[0079] This embodiment first targets the low-confidence mode and repairs the invalid associated edges in the graph by using the edge replacement logic of the high-confidence mode. This avoids the interference of low-confidence data on the analysis and ensures the integrity of the graph structure. At the same time, based on the reconstructed edge weights and the corresponding values of each node, the initial features corresponding to each node are obtained. Finally, the graph vector is generated based on the initial features.
[0080] Furthermore, as shown in Figure 2, this embodiment provides a step for updating the graph neural network graph based on the structural causal graph to obtain the graph embedding vector corresponding to each region agent, including:
[0081] Based on the graph vector, the embedding dimension corresponding to the modality to be replaced is obtained, and the key dimension vector is obtained.
[0082] The drift result is obtained based on the key dimension vector and the baseline distribution parameters corresponding to the current time.
[0083] Based on the structural causal graph and the drift results, the graph neural network graph is updated to obtain the graph embedding vector corresponding to each regional agent.
[0084] Specifically, after obtaining the spectral vector, since the ultimate goal of this embodiment is to output water-saving suggestions, and flow rate is the core analysis object in the water-saving scenario, it is necessary to select multiple dimensions with the highest correlation to flow rate from the spectral vector and recombine them to obtain a key dimension vector. This is to eliminate interference from irrelevant information and reduce the computational complexity of the subsequent detection process. For example, the correlation can be determined based on the correlation coefficient. The flow rate value corresponding to each preset time interval within a preset time period can be obtained to obtain the flow rate sequence. The spectral vector corresponding to each preset time interval within the preset time period can also be obtained to obtain the sequence corresponding to each dimension in each spectral vector. The correlation coefficient between each dimension and the flow rate sequence can be calculated to obtain multiple dimensions with the highest correlation coefficient. Finally, these dimensions are selected from the current spectral vector and combined to obtain the key dimension vector. This embodiment does not limit the specific number of dimensions selected.
[0085] Furthermore, this embodiment provides a step for obtaining the drift result based on the key dimension vector and the baseline distribution parameters corresponding to the current time, including:
[0086] The KL divergence is obtained based on the baseline distribution parameters and the current distribution parameters corresponding to the key dimension vectors at the current time.
[0087] If the KL divergence is greater than a preset value, the drift node is located based on mutual information analysis to obtain the drift result.
[0088] The baseline distribution is obtained from historical modal data samples. Different time periods correspond to different baseline distributions. Taking the baseline distribution corresponding to the midday peak (11:30-12:30) as an example, firstly, multimodal data within this time period is collected to generate corresponding key dimension vectors. For each sample's corresponding key dimension vector, its mean and covariance matrix are calculated as the baseline mean and baseline covariance for the morning peak period. Finally, the normal distribution defined by the baseline mean and baseline covariance is the baseline distribution corresponding to the morning peak period. The baseline mean and baseline covariance are the baseline distribution parameters.
[0089] The current distribution parameters corresponding to the key dimension vectors are obtained based on the currently collected multimodal data. Since only one key dimension vector is currently obtained, multiple key dimension vectors corresponding to each preset time interval within a preset time period can be obtained, and their mean and covariance can be calculated as the current distribution parameters. This embodiment does not limit the specific number of key dimension vectors or the length of the preset time period.
[0090] Then, based on the baseline distribution parameters at the current time and the current distribution parameters corresponding to the key dimension vectors, the KL divergence is obtained. For example, assume the current distribution... benchmark distribution The KL divergence is:
[0091]
[0092] in, Indicates a normal distribution. This represents the mean of the current distribution. This represents the covariance corresponding to the current distribution. This represents the mean of the baseline distribution. This represents the covariance corresponding to the baseline distribution. Represents the trace of a matrix. This represents the dimension corresponding to the key dimension vector. for The inverse matrix, This indicates transpose.
[0093] KL divergence is used to measure the difference between two distributions. If the KL divergence is greater than the preset value, it means that the overall distribution of the multimodal association features corresponding to the current key dimension vector has shifted compared with the baseline distribution, and water usage anomalies have occurred. However, at this time, the specific node causing the shift cannot be determined. This embodiment does not limit the value of the preset value.
[0094] Therefore, this embodiment further calculates the mutual information between the numerical value and the key dimension vector corresponding to each node. When calculating the mutual information, it is necessary to determine the joint distribution and marginal distribution of the numerical value and the key dimension vector corresponding to each node. The larger the mutual information value, the stronger the correlation between the node and the key dimension. If the mutual information of a certain node is higher than the preset benchmark, it means that the node is the core node that causes the distribution drift, and it is regarded as the drift node. Finally, the KL divergence results and the drift nodes are integrated to obtain the drift results. The joint distribution is determined according to the frequency of the simultaneous occurrence of node features and key dimension vectors in historical situations. The marginal distribution is the probability distribution of a single variable (the numerical value / key dimension vector corresponding to the node) itself, which can be obtained by summing all cases of the other variable in the joint distribution. This is a technical means that can be implemented by those skilled in the art, and this embodiment will not elaborate on it. The preset benchmark is the upper limit of the mutual information between the numerical value and the key dimension vector corresponding to the node under normal scenarios. It can be obtained through historical normal data or set by those skilled in the art. This embodiment does not limit it.
[0095] This embodiment uses KL divergence to quantify the overall difference between the current distribution parameters and the baseline distribution parameters at the current time, solving the problem that traditional detection methods cannot detect potential anomalies in advance. Then, drift nodes are located by mutual information calculation, making up for the shortcomings of the previous steps that can only detect correlations but cannot identify anomalies. This provides accurate anomaly anchor points for subsequent map updates and the generation of water-saving suggestions, ensuring that the system can identify changes in water use patterns in real time and avoid analysis failure or misjudgment due to feature drift, which would affect the accuracy of the final water-saving suggestion generation.
[0096] Furthermore, this embodiment provides a step of updating the graph neural network graph based on the structural causal graph and drift results to obtain the graph embedding vector corresponding to each regional agent, including:
[0097] The initial weights corresponding to the graph neural network graph are obtained from the structural causal graph.
[0098] Determine the update object in the graph neural network graph based on the drift results;
[0099] The initial weights corresponding to the updated objects in the graph neural network graph are updated based on the historical correlation matrix, and the graph embedding vector corresponding to each regional agent is obtained.
[0100] Specifically, the edge weights in the structural causal graph correspond to the initial weights of each edge in the graph neural network graph. Each node and edge in the graph neural network graph corresponds one-to-one with each node and edge in the structural causal graph. Then, the edges between drifting nodes are used as update objects in the graph neural network graph. Next, the correlation coefficients corresponding to the update objects are determined based on the historical correlation matrix, and the historical correlation matrix is updated based on the currently collected multimodal data to obtain the current correlation coefficients corresponding to the update objects. The absolute value of the difference between the two correlation coefficients is used as the loss function. Then, the initial weights are updated based on a preset learning rate and number of iterations. For example, the gradient descent method can be used to continuously update the weights according to the gradient of the loss function until the preset number of iterations is reached, and the updated weight values are output. If there are multiple update objects, the above update steps need to be repeated for each update object. This embodiment does not limit the learning rate and number of iterations. Gradient descent is a conventional technique, and this embodiment will not elaborate on it.
[0101] Furthermore, since the weights of each edge in the graph neural network graph need to maintain the overall reasonable correlation ratio to avoid abrupt changes in the weight of one edge leading to an imbalance in the graph correlation, this embodiment further updates the weights corresponding to the associated edges based on the adjusted weights of the updated object. The associated edges are all edges with the drifting node as the endpoint, excluding the updated object. Specifically, firstly, the weight change rate of the updated object before and after the update is calculated. Then, for each associated edge, its corresponding weight is multiplied by the weight change rate and an empirical ratio to obtain the weight adjustment amount. The weight adjustment amount is added to its corresponding weight to obtain the updated weight of the associated edge, thus obtaining the updated graph neural network graph. If there are multiple weight change rates, the average weight change rate can be calculated, and then its weight is multiplied by the empirical ratio. The empirical ratio needs to make the associated edges respond to the changes of the updated object without destroying the original correlation ratio of the graph. Those skilled in the art can determine the empirical ratio by simulating different ratios, such as 10%. Then, based on the above steps of obtaining the graph vector, each node calculates the graph vector according to the updated edge weights, that is, updates the graph vector, and obtains the graph embedding vector.
[0102] This embodiment, based on the drift nodes located in the previous steps and the adjusted edge weights, uses incremental graph neural network logic to locally update only the edges corresponding to the drift node set. This avoids the high computational cost of full graph reconstruction and integrates the real-time adjusted weights into the graph structure. At the same time, the final generated graph embedding vector compresses the updated associations, solving the problem that static graphs cannot adapt to real-time water usage pattern changes. Furthermore, the lightweight incremental update method ensures the real-time performance and efficiency of the system in high-concurrency scenarios.
[0103] Furthermore, this embodiment provides a step for generating water-saving suggestions based on the graph embedding vector corresponding to each regional agent, including:
[0104] Based on the graph embedding vector corresponding to each regional agent, the aggregate embedding vector is obtained;
[0105] Based on the aggregated embedding vector and the preset model, the global embedding vector is obtained;
[0106] Based on the global embedding vector, the global fusion graph is obtained;
[0107] Based on the drift results and the global fusion map, water-saving suggestions are generated.
[0108] Furthermore, this embodiment provides a step for obtaining an aggregated embedding vector based on the graph embedding vector corresponding to each regional agent, including:
[0109] Calculate the collaboration coefficient based on the graph embedding vector corresponding to each regional agent;
[0110] The attention weight is obtained based on the synergy coefficient;
[0111] The aggregated embedding vector is obtained based on the attention weights and the graph embedding vector corresponding to each region agent.
[0112] Specifically, as shown in Figure 3, after each regional agent (such as regional agent A and regional agent B) generates its corresponding graph embedding vector, it needs to send its own corresponding graph embedding vector (such as graph embedding vector A and graph embedding vector B), drift result, and the associated dimension vector corresponding to the drift node to the neighboring agents. The neighboring agents can be determined according to physical location and function type. This embodiment does not limit this. For example, if the current regional agent corresponds to the tea room, its corresponding neighboring agents are the tea room on the adjacent floor and the restroom on the same floor. The tea room on the adjacent floor is the neighboring agent determined according to physical location, and the restroom on the same floor is the neighboring agent determined according to function type.
[0113] First, upon receiving the drift result from a neighboring agent, the drift node in the drift result is determined in its own graph neural network graph, and its corresponding association dimension vector is obtained. Then, the KL divergence is calculated based on its own association dimension vector and the received association dimension vector to measure the degree of feature difference between the two agents at the same node. The method of calculating the KL divergence is similar to the above steps, that is, it is necessary to calculate the current distribution corresponding to its own association dimension vector and the received association dimension vector respectively, and obtain the KL divergence based on the current distribution parameters. This embodiment will not be elaborated here.
[0114] After obtaining the KL divergence, it is first multiplied by the difference sensitivity coefficient to obtain the difference compensation term. Since the feature differences between adjacent agents are usually small, in order to avoid the difference having too much impact on the weights, this embodiment sets a small amplification factor, such as the difference sensitivity coefficient, of 0.1. This allows the weight adjustment range to be sufficient to respond to the difference without deviating excessively from the original association pattern. This embodiment does not limit the value of the amplification factor; for example, those skilled in the art can obtain it through testing. Then, the difference compensation term is added to 1 to obtain the coordination coefficient, where 1 is the baseline coefficient, representing that the original weights are maintained when there is no difference. For example, assuming the current agent is... ,correspond A number of neighboring agents, which receive The first of the neighboring agents After receiving drift results from neighboring agents, the calculated KL divergence is used to measure the degree of feature difference between the two agents at the same node. Then, the KL divergence is multiplied by the difference sensitivity coefficient to obtain a difference compensation term. Based on this difference compensation term, a coordination coefficient is derived, which represents the current agent's performance. and the The cooperative coefficients corresponding to each neighboring agent .
[0115] Next, the collaboration coefficient is multiplied by the edge weight to update the weight of the edge corresponding to the drifting node. Then, the corresponding graph embedding vector is concatenated with the graph embedding vectors sent by neighboring agents (as shown in Figure 3, both regional agents A and B need to concatenate graph embedding vectors A and B). Then, the inner product of the attention vector and the concatenated vector is performed to obtain the association matching degree between agents. Then, the difference in association matching degree is amplified by exponential operation. Then, the attention weight is obtained based on the result of the exponential operation and the collaboration coefficient (as shown in Figure 3, regional agents A and B respectively calculate attention weight A and attention weight B). Finally, based on the attention weight, the corresponding graph embedding vector is weighted with the graph embedding vectors sent by neighboring agents to obtain the aggregated embedding vector (as shown in Figure 3, regional agents A and B respectively obtain aggregated embedding vector A and aggregated embedding vector B). For example, assume the current agent is... Its corresponding A number of neighboring agents, and their... The first of the neighboring agents The attention weights between neighboring agents are:
[0116]
[0117] in, Represents the attention vector. express The transpose of || indicates concatenation. Represents the characteristic transformation matrix. Indicates the current intelligent agent The corresponding graph embedding vector, and They represent the first The neighboring agents and the first The graph embedding vectors corresponding to each neighboring agent. , Indicates the current intelligent agent and the The cooperative coefficients corresponding to each neighboring agent. Indicates the current intelligent agent and the The cooperative coefficients corresponding to each neighboring agent. Indicates exponentiation. Indicates the current intelligent agent and the The association matching degree corresponding to each adjacent agent. The feature transformation matrix is used to adjust the dimension of the graph embedding vector. This embodiment does not limit this dimension; those skilled in the art can determine it based on computational cost, for example, setting the dimension to 128×64. The attention vector can be obtained through training, a common technique used by those skilled in the art, which will not be elaborated upon in this embodiment. As a simple example, the attention vector can be randomly initialized first, and the attention weights can be calculated based on the initialized attention vector and training data to obtain the aggregated embedding vector. Then, based on the aggregated embedding vector, it can be predicted whether the current water flow is abnormal, and the actual abnormal state can be obtained. The prediction error is obtained through cross-entropy loss, and finally, the prediction error is used as the loss function to iteratively update the initialized attention vector based on backpropagation.
[0118] After each regional agent obtains its corresponding aggregated embedding vector, it needs to send it to the fusion agent. As shown in Figure 4, regional agents A and B send aggregated embedding vector A and aggregated embedding vector B to the fusion agent respectively. The fusion agent generates a global embedding vector based on a preset model and finally outputs water-saving suggestions. The preset model includes a graph mapper and a domain discriminator. The graph mapper and domain discriminator are trained iteratively, eliminating time-varying differences between agents in different regions and time periods to obtain a globally usable embedding vector. For example, during training, multiple aggregated embedding vectors are first mapped to latent vectors. Then, the domain discriminator calculates the probability that a latent vector belongs to a peak water usage period. Based on this probability and the actual time period, a cross-entropy loss is calculated, and the parameters of the domain discriminator are updated using backpropagation. The graph mapper then regenerates latent vectors based on the updated parameters of the domain discriminator and calculates the inverse cross-entropy loss. The corresponding parameters of the graph mapper are then updated using the inverse cross-entropy loss and backpropagation. This process of updating the parameters of the graph mapper and domain discriminator is repeated until the cross-entropy loss and the inverse cross-entropy loss converge. Finally, the latent vector obtained by mapping multiple aggregated embedding vectors using the last updated graph mapper is used as the global embedding vector.
[0119] This embodiment first determines neighboring agents based on physical location and functional type, and uses information from neighboring agents to assist in generating water-saving suggestions, ensuring the accuracy of the suggestions. Specifically, after receiving drift information from neighboring nodes, the agent calculates the KL divergence to measure the feature differences of the same node, and generates a coordination coefficient based on the difference sensitivity coefficient to adjust the edge weights. Subsequently, through attention vectors and feature transformation matrices, the agent calculates the association matching degree of its own and neighboring agents' graph embeddings, exponentially amplifies the differences, obtains attention weights, and weights them to obtain an aggregated embedding vector. Finally, the aggregated embedding vector is sent to the fusion agent, and adversarial training between the graph mapper and the domain discriminator eliminates time-period differences to obtain a global embedding vector. This solves the problem of incomparable features across regions and time periods, ensuring the consistency and accuracy of water use pattern analysis in different scenarios.
[0120] Furthermore, this embodiment provides a step for obtaining a global fusion graph based on a global embedding vector, including:
[0121] Calculate the similarity between the global embedding vector and the positive and negative samples;
[0122] The loss function is derived based on the confidence and similarity scores corresponding to the multimodal datasets.
[0123] The global embedding vector is updated based on the loss function to obtain the global fused graph.
[0124] In this model, positive samples are global embedding vectors corresponding to peak water usage scenarios, and negative samples are global embedding vectors corresponding to scenarios outside of peak water usage scenarios. There are multiple positive and negative samples, which can be obtained from historical data. The similarity between the global embedding vectors and the positive and negative samples is then calculated. The confidence level for each modality is obtained from the confidence level list. Finally, a loss function is derived based on the confidence level weights and similarity. For example, the loss function... The formula is:
[0125]
[0126] in, This represents the mean cosine similarity between the current global embedding vector and the positive samples. Based on the cosine similarity, the positive sample with the highest similarity to the current global embedding vector can be determined. A confidence list corresponding to this positive sample is obtained, and the confidence scores for each modality in the confidence list are multiplied together. The result of the multiplication is... ,in Indicates multiplication. Indicates the first in the confidence list One confidence level, Indicates the relationship between the current global embedding vector and the first... Cosine similarity of negative samples Indicates the first The confidence list corresponding to the nth negative sample is the nth... One confidence level, Indicates the first The result is the product of the confidence scores for each modality in the confidence score list corresponding to each negative sample. These are manually set hyperparameters used to control the similarity difference between positive and negative samples. The smaller the value, the more the similarity difference between positive and negative samples will be amplified. This embodiment... There is no limit to the specific value; for example, if the similarity difference between positive and negative samples is small, a smaller value is usually set. For example, a value of 0.05-0.1 is used to avoid eliminating similarity differences. If the similarity difference between positive and negative samples is already large, a larger value is usually set. For example, 0.1-0.2, to avoid numerical overflow. The similarity difference between positive and negative samples can be determined by the difference between the average cosine similarity between the current global embedding vector and the positive sample and the average cosine similarity between the current global embedding vector and the negative sample. This embodiment does not restrict the specific criteria for dividing the similarity difference. For example, when the difference is less than 0.2, it can be considered that the similarity difference between positive and negative samples is small.
[0127] Specifically, in this embodiment, the numerator of the loss function This represents the effective association strength between the current global embedding vector and positive samples, obtained from the positive sample with the highest similarity. This indicates the credibility of the sample, and the denominator represents the total association strength between the current global embedding vector and the positive sample and all negative samples. In this embodiment, the purpose of selecting only the positive sample with the highest similarity for calculation is to make the current global embedding vector closely match the target pattern best suited to the current scene, facilitating accurate generation of water-saving suggestions. The purpose of using all negative samples for calculation is to keep the current global embedding vector away from all possible interfering target patterns. Furthermore, the logarithmic function has the property that the larger the value, the higher the correlation coefficient. The smaller the result, the more likely the effective association strength is to be much greater than the association strength corresponding to the negative sample. When the value of the loss function approaches 0, it indicates that the current global embedding vector is close to the positive sample. Conversely, when the correlation strength of the negative sample is too high, the value of the loss function is larger.
[0128] Finally, the global embedding vector is adjusted based on the value of the loss function. Specifically, when the value of the loss function is less than the preset loss value, the current global embedding vector is multiplied by the fine-tuning coefficient to obtain the updated global embedding vector. The fine-tuning coefficient is used to further enhance the correlation between the current global embedding vector and positive samples. Its value can be determined based on the value of the loss function. The smaller the value of the loss function, the smaller the fine-tuning coefficient. For example, it can be set to 1.05, but this embodiment does not limit its specific value. Finally, the updated global embedding vector is recorded as the global fusion map. If the value of the loss function is greater than or equal to the preset loss value, positive and negative samples can be reselected and the above steps can be repeated. This embodiment does not limit the specific value of the preset loss value.
[0129] Finally, after obtaining the global fusion graph, water-saving suggestions adapted to the current scenario are generated based on the VAE. The VAE is an existing model, which can be exemplified by an encoder, a reparameterization layer, and a decoder. After obtaining the global fusion graph, the specific values corresponding to each drift node are obtained according to the drift results. After normalizing and preprocessing the global fusion graph and the specific values corresponding to the drift nodes, they are concatenated to obtain the input vector. The encoder is a two-layer fully connected neural network used to map the input vector to the latent space, obtaining the latent space mean and the log-variance of the latent space, where the log is used to ensure that the values are non-negative. Then, the reparameterization layer generates latent vectors using reparameterization techniques based on the latent space mean and the log-variance of the latent space. The decoder can also be a two-layer fully connected neural network used to map the latent vectors to the specific parameter vectors corresponding to the water-saving suggestions. This embodiment... The dimensions of the specific parameter vector are not limited and can be set according to the actual situation. For example, assuming that the area in this embodiment includes tea rooms on floors 1-5, the specific parameter vector can be [0,0,1,0,0,2,520,5,525,5,27.5,15.4]. The first to fifth dimensions correspond to the tea rooms on floors 1-5. The third dimension of 1 indicates that the current area agent is the area agent corresponding to the tea room on floor 3, that is, the current water-saving suggestion is applied to the tea room on floor 3. 2 indicates that water is collected in two batches at off-peak times. By collecting water at off-peak times, empty flow waste is avoided and pipeline pressure is reduced, thereby achieving the purpose of water saving. 520 and 525 indicate the start time of the two batches of water collection. The unit is minutes, which need to be converted to hours. For example, 520 corresponds to 8:40. 5 indicates the duration. 27.5 is the expected water flow rate and 15.4 is the expected water saving rate.
[0130] Furthermore, VAEs are obtained through training. Model training is a technical means that can be achieved by those skilled in the art, and this embodiment will not elaborate on it. For example, training data can be set and the parameters in the model can be initialized. The parameters in the model include the weight matrix and the bias. The loss function is set by the difference between the water-saving suggestions output by the model and the water-saving suggestions labeled by humans. The Adam optimizer is selected through the loss function to iteratively optimize the model parameters according to the preset learning rate, and finally the water-saving suggestions after training are obtained.
[0131] This embodiment first strengthens the core associated features through contrastive learning in the preset model, and then uses VAE to generate off-peak water-saving suggestions adapted to the scenario, transforming abstract feature data into an executable water-saving plan, solving the problem of unavailable multimodal and cross-regional features, and ensuring the accuracy of water-saving suggestions.
[0132] This embodiment provides a water-saving intelligent agent interactive control system based on multimodal fusion, including a regional intelligent agent and a fusion intelligent agent. The regional intelligent agent includes:
[0133] The acquisition module is used to acquire a multimodal dataset, which includes video modal data, audio modal data, structural modal data, and activity modal data.
[0134] The testing module is used to obtain a structural causal graph based on the multimodal dataset and Granger causality test;
[0135] The update module is used to update the graph neural network graph based on the structural causal graph to obtain the graph embedding vector corresponding to each region agent;
[0136] The fusion agent is used to generate water-saving suggestions based on the graph embedding vector corresponding to each regional agent.
[0137] This application provides a water-saving intelligent agent interactive control method and system based on multimodal fusion. First, it integrates multimodal data such as video, audio, structure, and activity. Then, it constructs a structural causal graph by combining Granger causality test, clarifying the temporal dependencies between the data of each modality. Based on the structural causal graph, it dynamically updates the graph neural network graph and generates the graph embedding vector of the regional intelligent agent. Finally, the fused intelligent agent generates water-saving suggestions by aggregating the graph embedding vectors of each region. This solves the adaptation failure problem of static models in traditional solutions and ensures the global adaptability of water-saving suggestions.
[0138] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0140] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0141] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A water-saving intelligent agent interactive control method based on multimodal fusion, characterized in that, The water-saving agent comprises multiple regional agents and a fusion agent, with each regional agent corresponding to a region. The method includes: acquiring a multimodal dataset, which includes video modal data, audio modal data, structural modal data, and activity modal data. The video modal data includes faucet on / off status frames, the audio modal data includes the spectrum of abnormal noises from pipe leaks, the structural modal data includes flow rate and water temperature, and the activity modal data includes the number of people active within the region. Based on the multimodal dataset and Granger causality test, a structural causal graph is obtained. The graph neural network graph is updated based on the structural causal graph to obtain a graph embedding vector corresponding to each regional agent, so that the fusion agent generates water-saving suggestions based on the graph embedding vector corresponding to each regional agent. Generating water-saving suggestions based on the graph embedding vector corresponding to each regional agent includes: obtaining an aggregated embedding vector based on the graph embedding vector corresponding to each regional agent; and then, based on the aggregated embedding vector... The process involves embedding vectors and a preset model to obtain a global embedding vector; obtaining a global fusion map based on the global embedding vector; and generating water-saving suggestions based on the drift results and the global fusion map. Specifically, this includes obtaining the specific values corresponding to each drift node based on the drift results; normalizing and preprocessing the global fusion map and the specific values corresponding to the drift nodes, then concatenating them to obtain an input vector; mapping the input vector to the latent space using an encoder to obtain the latent space mean and log-variance; generating latent vectors using a reparameterization layer based on the latent space mean and log-variance; mapping the latent vectors to specific parameter vectors corresponding to the water-saving suggestions using a decoder; and parsing the specific parameter vectors to obtain the water-saving suggestions. The encoder, reparameterization layer, and decoder constitute a VAE (Variable Energy Array), which is obtained through training. Each dimension in the specific parameter vector corresponds to the current regional agent, off-peak water collection batch, water collection start time, duration, expected water flow, and expected water saving rate in the water-saving suggestions.
2. The water-saving intelligent agent interactive control method based on multimodal fusion according to claim 1, characterized in that, Based on the multimodal dataset and Granger causality test, a structural causal graph is obtained, including: obtaining an edge set based on historical modal data and Granger causality test; obtaining a node set and reconstructed edge weights based on the multimodal dataset; and obtaining the structural causal graph and graph vector based on the reconstructed edge weights and structural causal model.
3. The water-saving intelligent agent interactive control method based on multimodal fusion according to claim 2, characterized in that, Obtaining reconstructed edge weights based on the multimodal dataset includes: obtaining the modality to be replaced based on the multimodal dataset and a preset threshold; obtaining the replacement modality based on the historical correlation matrix and the replacement modality; and updating the original graph edge weights based on the correlation coefficients corresponding to the replacement modality in the historical correlation matrix to obtain the reconstructed edge weights.
4. The water-saving intelligent agent interactive control method based on multimodal fusion according to claim 3, characterized in that, The graph neural network graph is updated based on the structural causal graph to obtain the graph embedding vector corresponding to each regional agent, including: obtaining the embedding dimension corresponding to the modality to be replaced based on the graph vector, and obtaining the key dimension vector; obtaining the drift result based on the key dimension vector and the baseline distribution parameter corresponding to the current time; updating the graph neural network graph based on the structural causal graph and the drift result to obtain the graph embedding vector corresponding to each regional agent.
5. The water-saving intelligent agent interactive control method based on multimodal fusion according to claim 4, characterized in that, The drift result is obtained based on the key dimension vector and the baseline distribution parameters corresponding to the current time, including: obtaining the KL divergence based on the baseline distribution parameters corresponding to the current time and the current distribution parameters corresponding to the key dimension vector; if the KL divergence is greater than a preset value, the drift node is located based on mutual information analysis to obtain the drift result.
6. The water-saving intelligent agent interactive control method based on multimodal fusion according to claim 4, characterized in that, Based on the structural causal graph and the drift result, the graph neural network graph is updated to obtain the graph embedding vector corresponding to each regional agent. This includes: obtaining the initial weights corresponding to the graph neural network graph based on the structural causal graph; determining the update object in the graph neural network graph based on the drift result; and updating the initial weights corresponding to the update object in the graph neural network graph based on the historical association matrix to obtain the graph embedding vector corresponding to each regional agent.
7. The water-saving intelligent agent interactive control method based on multimodal fusion according to claim 1, characterized in that, The step of obtaining the aggregated embedding vector based on the graph embedding vector corresponding to each regional agent includes: calculating a coordination coefficient based on the graph embedding vector corresponding to each regional agent; obtaining an attention weight based on the coordination coefficient; and obtaining the aggregated embedding vector based on the attention weight and the graph embedding vector corresponding to each regional agent.
8. The water-saving intelligent agent interactive control method based on multimodal fusion according to claim 1, characterized in that, Obtaining a global fusion graph based on the global embedding vector includes: calculating the similarity between the global embedding vector and positive and negative samples; obtaining a loss function based on the confidence level corresponding to the multimodal dataset and the similarity; and updating the global embedding vector based on the loss function to obtain the global fusion graph.
9. A water-saving intelligent agent interactive control system based on multimodal fusion, characterized in that, The system includes regional agents and fusion agents, with each regional agent corresponding to a region. Each regional agent comprises: a data acquisition module for acquiring a multimodal dataset, including video modal data, audio modal data, structural modal data, and activity modal data. The video modal data includes faucet on / off status frames, the audio modal data includes the spectrum of abnormal noises from pipe leaks, the structural modal data includes flow rate and water temperature, and the activity modal data includes the number of people active within the region; a verification module for obtaining a structural causal graph based on the multimodal dataset and Granger causality tests; and an update module for updating the graph neural network graph based on the structural causal graph to obtain a graph embedding vector corresponding to each regional agent. The fusion agent generates water-saving suggestions based on the graph embedding vectors corresponding to each regional agent. Generating water-saving suggestions based on the graph embedding vectors corresponding to each regional agent includes: obtaining an aggregated embedding vector based on the graph embedding vectors corresponding to each regional agent. The process involves: obtaining a global embedding vector based on the aggregated embedding vector and a preset model; obtaining a global fusion graph based on the global embedding vector; and generating water-saving suggestions based on the drift results and the global fusion graph. Specifically, this includes obtaining the specific values corresponding to each drift node based on the drift results; performing normalization preprocessing on the global fusion graph and the specific values corresponding to the drift nodes, then concatenating them to obtain an input vector; mapping the input vector to the latent space using an encoder to obtain the latent space mean and log-variance; generating latent vectors using a reparameterization layer based on the latent space mean and log-variance; mapping the latent vectors to specific parameter vectors corresponding to the water-saving suggestions using a decoder; and parsing the specific parameter vectors to obtain the water-saving suggestions. The encoder, reparameterization layer, and decoder constitute a VAE (Variable Energy Array), which is obtained through training. Each dimension in the specific parameter vector corresponds to the current regional agent, off-peak water collection batch, water collection start time, duration, expected water flow, and expected water saving rate in the water-saving suggestions.
Citation Information
Patent Citations
Multi-modal equipment integrated management system and method based on intelligent AI
CN120217271A
Decision scheme generation method and device based on complex multi-modal management data, equipment and medium
CN120974219A
Intelligent irrigation method and device based on large language model
CN120982399A
Digital human interaction system and method based on multi-modal emotion recognition
CN121116129A
Office process optimization method and device based on artificial intelligence, equipment and medium
CN121119947A