Fine-grained home electricity monitoring system and method combining smart speaker and electricity meter

By combining smart speakers and electricity meters, and utilizing cross-modal learning and self-supervised methods, the correlation between appliance power and sound is automatically learned, solving the accuracy and cost issues of household electricity monitoring and realizing low-cost, fine-grained electricity monitoring and user behavior analysis.

CN117131462BActive Publication Date: 2025-12-12BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311075482.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-24
Publication Date
2025-12-12
Estimated Expiration
2043-08-24

AI Technical Summary

Technical Problem

Existing household electricity monitoring methods struggle to achieve accurate, fine-grained decomposition, especially when multiple devices are operating simultaneously or have similar power outputs. Furthermore, existing technologies often require additional equipment or extensive data annotation, resulting in high costs and low accuracy.

Method used

By combining smart speakers and electricity meters, a cross-modal learning module is used to automatically learn the correlation between appliance power and sound. A self-supervised method is used to extract consistency and complementarity information, enabling fine-grained monitoring of appliance energy consumption and reducing equipment and labeling costs.

Benefits of technology

It enables low-cost, low-labeling-cost fine-grained monitoring of household electricity consumption, improves the accuracy of appliance energy consumption breakdown, helps users identify electricity consumption behavior, and assists in smart home and health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117131462B_ABST
    Figure CN117131462B_ABST
Patent Text Reader

Abstract

The application discloses a fine-grained household power monitoring system and method combined with an intelligent sound box and an electric meter, and belongs to the related field of household electrical appliance energy consumption monitoring.The method comprises the following steps: the system automatically learns the correlation between electrical appliance power and sound, i.e., consistency information and complementary information; power events are divided into high power changes and low power changes, and scene discovery is iteratively performed; sound features and power features with correlation, i.e., key feature pairs, are found, and it is understood through clustering that which key feature pairs belong to the same electrical appliance state; a noise-robust sound-based electrical appliance state recognizer is trained; and the learned consistency information and complementary information, i.e., the recognition result of the electrical appliance state recognizer, is used to realize electrical appliance energy consumption decomposition in a cross-modal correlation fusion manner, and then the power consumption of each type of electrical appliance is inferred.The application helps users understand fine-grained household power consumption in a low device cost and low labeling cost manner, helps users cultivate low-carbon power consumption habits, and can also assist in user activity recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of household appliance energy consumption monitoring, and particularly relates to a fine-grained household electricity monitoring system and method combined with an intelligent sound box and an electric meter. BACKGROUND

[0002] Building energy consumption accounts for about 46% of total human energy consumption, and energy saving and emission reduction issues make building energy digitalization research a hot topic. Among building energy, household electricity generates a large proportion of carbon emissions, and the Ministry of Housing and Urban-Rural Development estimates that the proportion of electricity consumption in building energy consumption will exceed 55% by 2025. Coarse-grained total electric meter data cannot reflect the direction of energy consumption, and cannot develop effective energy saving plans. Fine-grained understanding of energy consumption helps to improve the living habits of households and reduce up to 20% of carbon emissions. Therefore, it is necessary to understand the fine-grained household electricity behavior of users, which helps to identify the energy consumption distribution of households and thus help to make building energy low-carbon.

[0003] To solve this problem, the existing mainstream power monitoring method is achieved by energy decomposition, that is, the energy consumption of a single appliance is estimated according to the total power data in the house. However, the existing method is difficult to balance accuracy and use burden. Invasive load monitoring (ILM) methods can achieve high decomposition accuracy, but they require the deployment of additional devices, such as smart sockets and circuit breakers. Non-intrusive load monitoring (NILM) methods have lower device costs (only one electric meter is needed). However, supervised NILM methods require the collection and labeling of a large amount of training data, and unsupervised NILM methods require a large amount of prior knowledge. In addition, when multiple devices are running simultaneously or when multiple devices have similar power, the accuracy of unsupervised NILM methods is low.

[0004] Sound as a complementary feature of appliance power has obvious advantages in the following aspects: (1) distinguishing different appliances with similar power, and (2) decomposing the power of parallel running appliances. Some multi-modal methods use sound to improve the performance of energy consumption decomposition, but they still require users to place multiple sensors or carry portable devices to place the sound sensor close to the appliance, and require a large amount of data labeling or prior knowledge. SUMMARY

[0005] In order to solve the problems existing in the prior art, the present application provides a fine-grained household power monitoring system and method combined with a smart speaker and a power meter. Considering the advantages of sound in distinguishing similar power appliances and parallel appliances, and the high popularity of smart speakers in homes, the present application combines the power data and sound data of smart power meters and smart speakers, mines the correlation between them, helps users understand the fine-grained power consumption in the home at low device cost and low labeling cost, enables low-carbon life, and assists in user activity recognition, which is used in the fields of smart home, health monitoring, etc.

[0006] The technical solution adopted by the present application to solve the technical problems is as follows:

[0007] The fine-grained household power monitoring system combined with a smart speaker and a power meter of the present application comprises a cross-modal learning module and a cross-modal inference module, the cross-modal learning module is used to automatically learn the correlation between appliance power and sound, i.e. consistency information and complementary information, the cross-modal inference module is used to utilize the learned consistency information and complementary information, i.e. the recognition result of the appliance state recognizer, to realize appliance energy consumption decomposition in a cross-modal correlation fusion manner, and further infer the power consumption of each type of appliance.

[0008] Further, the cross-modal learning module comprises a two-stage scene discovery component, a self-supervised consistency information extraction component and a noise-robust sound-based appliance state recognizer training component; the two-stage scene discovery component and the self-supervised consistency information extraction component are used to extract consistency information from appliance power and sound data; the noise-robust sound-based appliance state recognizer training component first performs data augmentation on appliance power and sound data with consistency, and then performs iterative training, and obtains the running state of the appliance, i.e. complementary information, from the appliance sound data through the noise-robust sound-based appliance state recognizer.

[0009] Further, the cross-modal inference module comprises a cross-modal data correlation component, a sound-assisted scene discovery component and an appliance label correlation component; the cross-modal data correlation component is used to verify the recognized sound-related appliance running state; the sound-assisted scene discovery component is used to identify the sound-related appliance cycle, and after discovering the sound-related scene, the remaining candidate scenes are discovered by using a two-stage scene discovery operation; the appliance label correlation component is used to cluster the sound-related appliance cycle, and identifies through one-time interaction to obtain the final label result of appliance energy consumption decomposition.

[0010] The fine-grained household power monitoring method combined with a smart speaker and a power meter of the present application is realized by using a fine-grained household power monitoring system combined with a smart speaker and a power meter, and the method comprises the following steps:

[0011] Step S1: The system automatically learns the correlation between the power and sound of the electrical appliance, i.e., consistency information and complementary information;

[0012] Step S1.1: Divide the power events into high power changes and low power changes, and iteratively perform scenario discovery;

[0013] Step S1.2: Find sound features and power features that have a correlation, i.e., key feature pairs, and understand which key feature pairs belong to the same electrical appliance state through clustering;

[0014] Step S1.3: Noise-robust sound-based electrical appliance state recognizer training;

[0015] Step S2: Use the learned consistency information and complementary information, i.e., the recognition results of the electrical appliance state recognizer, to realize electrical energy consumption decomposition in a cross-modal correlation fusion manner, and then infer the power consumption of each type of electrical appliance;

[0016] Step S2.1: Use the learned consistency information to verify the recognition results of the noise-robust sound-based electrical appliance state recognizer;

[0017] Step S2.2: Identify the electrical appliance cycle related to sound, and use a two-stage scenario discovery operation to discover the remaining candidate scenarios after discovering the scenario related to sound;

[0018] Step S2.3: Cluster the electrical appliance cycle related to sound, and identify it through one-time interaction to obtain the final label result of electrical energy consumption decomposition.

[0019] Further, the specific operation process of step S1.1 is as follows:

[0020] The DPGMMS model is used to divide the power events into high power changes and low power changes, and then iteratively perform scenario discovery; a scenario E n =(e n1 ,e n2 ,…,e nL ) is a subsequence of an ordered power event sequence; if the power change P n =(P n1 ,P n2 ,…,P nL ) corresponding to the subsequence satisfies the axiom constraint, it can be regarded as an effective scenario set E valid , i.e., an electrical appliance cycle; scenarios that have not been verified by the axiom constraint are referred to as candidate scenario set E cand .

[0021] Further, the specific operation process of step S1.2 is as follows:

[0022] Step S1.2.1: Sound feature extraction

[0023] According to a given sound data stream, it is divided into sound segments with a length of 0.96 seconds. For the sound segment around time t, if its decibel value (DB) is greater than -45 DB, it is considered as a sound event a t ; convert each 0.96s sound segment into sound features using VGG-ish as a sound feature extractor

[0024] Step S1.2.2: Encoding and decoding

[0025] Use the encoder to decouple the sound features into a noise feature vector Z t,Noise and an effective feature vector Z t,Clean , and reconstruct the sound features The calculation formula of the reconstruction loss is shown in equation (1):

[0026]

[0027] Step S1.2.3: Power prediction

[0028] According to the decoupled effective feature vector Z t,Clean , predict the power value; the calculation formula of the loss function of the power predictor is shown in equation (2):

[0029]

[0030] where p t is the actual power value, is the predicted power value.

[0031] Step S1.2.4: Construct the overall loss function

[0032] The calculation formula of the overall loss function is shown in equation (3):

[0033]

[0034] where λ is a hyperparameter that balances the prediction error and the reconstruction error.

[0035] Further, the specific operation process of step S1.3 is as follows:

[0036] Move and overlap the target sound segment with other sound segments to generate rich training data with noise through data enhancement; according to the identification result of the appliance state recognizer, iteratively train new negative samples to enhance the performance of the appliance state recognizer.

[0037] Further, the specific operation process of step S2.1 is as follows:

[0038] The recognition result of the appliance state recognizer is directly associated with the power event at the moment, and the recognition result of the noise-robust sound-based appliance state recognizer is verified by the cross-modal learning module using the consistency information learned, and if the power value does not match, it is discarded.

[0039] Further, in step S2.2, different scenario discovery rules are used according to the state combination of different appliances: (i) if the start, running and end states are all recognized, the final appliance cycle is directly selected from the candidate scenarios according to the power value of each state; (ii) if only part of the state is recognized, the candidate space of the candidate scenario is reduced, and the final scenario is screened out using the axiomatic constraint.

[0040] Further, in step S2.3, a single appliance cycle contains multiple power events, if a single appliance cycle contains a sound-related power event, it is marked as a pseudo-label of the sound class; if all power events in the appliance cycle do not contain sound pseudo-labels, use density peak clustering to cluster similar appliance cycles and mark pseudo-labels for identification; at the same time, ask different types of interactive questions according to the scenario.

[0041] The beneficial effects of the present application are:

[0042] The prior art mainly uses invasive methods or supervised non-invasive methods, and rarely uses unsupervised or self-supervised non-invasive methods, even if it is used, only a single mode is considered, and the accuracy is limited. The present application uses non-invasive technology to fuse sound and power data in a self-supervised manner to realize fine-grained monitoring of household appliance energy consumption. The present application only uses the existing sensing devices in the home, i.e. smart speakers and smart meters, to reduce the cost of equipment required for monitoring. In addition, the present application ingeniously uses the consistency and complementary information of appliance sound and power to improve system performance while reducing the data labeling cost required for training. The present application is helpful for analyzing user electricity consumption behavior, thereby helping users to develop low-carbon electricity consumption habits, and can also assist in user activity recognition, which is used in the fields of smart home, health monitoring, etc. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 A schematic diagram of a fine-grained household electricity monitoring method combining a smart speaker and a meter according to the present application.

[0044] Figure 2 A schematic diagram of self-supervised consistency information extraction according to the present application. DETAILED DESCRIPTION

[0045] With reference to the drawings, the specific embodiments of the present application will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0046] With reference to the drawings, the specific embodiments of the present application will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application. Figure 1 To illustrate, the fine-grained household power monitoring system combined with an intelligent sound box and an electric meter specifically includes: a cross-modal learning module and a cross-modal inference module; wherein the cross-modal learning module is used to automatically learn the correlation between the power of electrical appliances and the sound (consistency information and complementary information); the cross-modal inference module is used to realize electrical appliance energy consumption decomposition in a cross-modal correlation fusion manner by using the consistency information and the complementary information (the identification result of the electrical appliance state identifier) learned by the cross-modal learning module, and then infer the power consumption of each type of electrical appliance.

[0047] In addition, the cross-modal learning module includes three components, which are a two-stage scene discovery component, a self-supervised consistency information extraction component, and a noise-robust sound-based electrical appliance state identifier training component; the two-stage scene discovery component and the self-supervised consistency information extraction component are used to extract consistency information from electrical appliance power and sound data; the noise-robust sound-based electrical appliance state identifier training component first performs data enhancement on electrical appliance power and sound data with consistency, and then performs iterative training, and obtains the running state of the electrical appliance from the electrical appliance sound data through the noise-robust sound-based electrical appliance state identifier, that is, complementary information. The self-supervised consistency information extraction component includes: a sound feature extractor, an encoder, a decoder, a power predictor, and an error function calculation module; wherein the functions of each component are as follows:

[0048] (1) Sound feature extractor

[0049] According to a given sound data stream, it is divided into sound segments with a length of 0.96 seconds. For the sound segment around time t, if its decibel value (DB) is greater than -45DB, it is considered as a sound event a t ; using VGG-ish as a sound feature extractor to convert each 0.96s sound segment into a sound feature

[0050] (2) Encoder and decoder

[0051] The sound feature is decoupled into a noise feature vector Z t,Noise and an effective feature vector Z t,Clean by the encoder, and the sound feature is reconstructed by the decoder The calculation formula of the reconstruction loss is shown as formula (1):

[0052]

[0053] (3) power predictor;

[0054] According to the decoupled effective feature vector Z t,Clean The calculation formula of the loss function of the power predictor is shown as formula (2):

[0055]

[0056] Wherein, p t is the actual power value, is the predicted power value.

[0057] (4) error function calculation module;

[0058] The overall loss function is constructed, and the calculation formula of the overall loss function is shown as formula (3):

[0059]

[0060] Wherein, λ is a hyperparameter for balancing the prediction error and the reconstruction error.

[0061] In addition, the cross-modal inference module utilizes the consistency information and the complementary information (the identification result of the appliance state identifier) learned by the cross-modal learning module to realize more accurate appliance energy consumption decomposition in a cross-modal correlation fusion manner, and then accurately infer the power consumption of each type of appliance. The cross-modal inference module comprises three components, namely a cross-modal data correlation component, a sound-assisted scene discovery component and an appliance label correlation component; the cross-modal data correlation component is used to verify the identified sound-related appliance operating state; the sound-assisted scene discovery component is used to identify a sound-related appliance cycle, and after discovering a sound-related scene, a two-stage scene discovery operation is used to discover the remaining candidate scenes; and the appliance label correlation component is used to cluster the sound-related appliance cycle, and identify through one-time interaction to obtain a final appliance energy consumption decomposition label result.

[0062] Referring to Figure 1 and Figure 2 It is explained that the fine-grained household power monitoring method of the joint intelligent sound box and the electric meter specifically comprises the following steps:

[0063] Step S1: cross-modal learning;

[0064] The cross-modal learning module is used to automatically learn the correlation between the power and the sound of the electrical appliance, i.e. consistency information and complementary information. The specific operation process is as follows:

[0065] Step S1.1: two-stage scene discovery;

[0066] Firstly, the DPGMMS model, i.e. Dirichlet Process Gaussian Mixture Model, is used to divide the power events into high power changes and low power changes, and then iterative scene discovery is performed. Taking high power as an example, firstly, the power events e1, e2, …, e m A candidate scene set E cand is constructed. valid .

[0067] Among them, a scene E n =(e n1 ,e n2 ,…,e nL ) is a subsequence of an ordered power event sequence. If the power change P n =(P n1 ,P n2 ,…,P nL ) corresponding to the subsequence satisfies the axiom constraint, i.e. can be regarded as an effective scene set E valid , i.e. the period of the electrical appliance; the scene which has not been verified by the axiom constraint is called a candidate scene set E cand .

[0068] Step S1.2: consistency information extraction based on self-supervision;

[0069] The consistency information extraction component based on self-supervision is adopted to find the associated sound features and power features, and the associated sound features and power features are called key feature pairs. Through clustering, it can be known which key feature pairs belong to the same electrical appliance state. As shown in FIG. 2, the specific operation process is as follows: Figure 2

[0070] Step S1.2.1: sound feature extraction;

[0071] According to a given sound data stream, it is divided into sound segments with a length of 0.96 seconds. For the sound segment around time t, if the decibel value (DB) is greater than -45DB, it is considered to be a sound event a t .

[0072] The VGG-ish is used as a sound feature extractor in the application to convert each 0.96s sound segment into a sound feature The sound feature is a 128-dimensional feature vector for subsequent encoder and decoder processing. ​

[0073] Step S1.2.2: Encoding and decoding;

[0074] Since the sound feature extractor generates sound features which can be overlapped with noise, the present application designs an encoder to decouple them into a noise feature vector Z t,Noise and an effective feature vector Z t,Clean Then the sound feature is reconstructed by the decoder The calculation formula of the reconstruction loss is shown in equation (1):

[0075]

[0076] Step S1.2.3: Power prediction;

[0077] The power value is predicted by the power predictor according to the decoupled effective feature vector Z t,Clean If the sound feature and the power feature have consistency, the predicted power value will have a smaller error with the actual power value p t ; Wherein, the calculation formula of the loss function of the power predictor is shown in equation (2):

[0078]

[0079] Wherein, p t is the actual power value, is the predicted power value.

[0080] Step S1.2.4: Constructing the overall loss function;

[0081] The encoder, the decoder and the power predictor are jointly trained to obtain better consistency information extraction performance. The calculation formula of the overall loss function of the model is shown in equation (3):

[0082]

[0083] Wherein, λ is a hyperparameter balancing the prediction error and the reconstruction error.

[0084] Step S1.3: Noise-robust sound-based appliance state recognizer training;

[0085] The consistency information (key feature pairs) is limited in quantity, and lacks sufficient negative samples, making it difficult to train a noise-robust sound-based appliance state recognizer, and further difficult to obtain effective complementary information. To solve this problem, the present application generates rich training data with noise by moving and overlapping target sound segments with other sound segments through data augmentation. In addition, the present application also designs an iterative training module to enhance the performance of the appliance state recognizer by iteratively training new negative samples based on the recognition results of the appliance state recognizer.

[0086] Step S2: Cross-modal inference;

[0087] The consistency information and complementary information (recognition results of the appliance state recognizer) learned by the cross-modal learning module are used to realize more accurate appliance energy consumption decomposition in a cross-modal correlation fusion manner, and further accurately infer the power consumption of each type of appliance. The specific operation process is as follows:

[0088] Step S2.1: Cross-modal data correlation;

[0089] The recognition results of the appliance state recognizer can be directly correlated with the power event at that moment. However, due to the complex relationship between sound features and power features or recognition errors of the appliance state recognizer itself, incorrect state correlations often occur. To solve this problem, the present application verifies the recognition results of the noise-robust sound-based appliance state recognizer using the learned consistency information through the cross-modal learning module. In other words, if the power value does not match, it is discarded.

[0090] Step S2.2: Sound-assisted context discovery;

[0091] The two-stage context discovery operation cannot guarantee that the discovered power events in the appliance period belong to the correct appliance, which can lead to error accumulation and cause other power events to be incorrectly discovered as the same appliance period. To solve this problem, the present application first identifies the sound-related appliance period to reduce the number of candidate contexts to facilitate the discovery of silent appliances. Since sound appliances do not produce sound in all states, when discovering sound-related appliance periods, the present application uses different context discovery rules based on different state combinations of appliances: (i) If the start, run, and end states are all recognized, the final appliance period can be selected from the candidate contexts based on the power values of each state; (ii) If only part of the state is recognized, the candidate space of the candidate context can be reduced, and the final context can be selected using the axiomatic constraint.

[0092] After discovering the sound-related context, the present application uses the two-stage context discovery operation to discover the remaining candidate contexts.

[0093] Step S2.3: Appliance label correlation;

[0094] The appliance energy decomposition cannot reveal the relationship between the decomposed appliance category and power consumption. Therefore, the appliance cycle related to the sound is clustered through the appliance tag association component, and identified through a one-time interaction to obtain the final appliance energy decomposition tag result. The specific operation process is as follows:

[0095] An appliance cycle contains multiple power events. The appliance cycle mentioned refers to a group of power events meeting the axiom constraints. The axiom constraints describe the physical characteristics of an appliance power: (i) the running power cannot be negative; (ii) the sum of power change values should be less than a threshold V p ; (iii) the sum of power change values and adjacent power change differences should also be less than a threshold V p . In the present application, the power change value sum threshold V p is set to 30W. In addition, the power event e i can be regarded as the timestamp of the ith power change, and the power change value is p i . Given the total power data stream, a series of ordered power events e1, e2, …, e m and related power changes p1, p2, …, p m can be obtained.

[0096] If an appliance cycle contains a power event related to the sound, it can be marked as a pseudo tag of the sound category. If all power events in the appliance cycle do not contain the sound pseudo tag, the present application uses density peak clustering to cluster similar appliance cycles and identify the pseudo tag. Specifically, the time length, maximum power value of the appliance cycle, and the time interval of the appliance cycle closest to the appliance cycle and having similar power can be used as the features of density clustering. If the time interval is greater than 30 minutes, it is set to a maximum value.

[0097] The present application will ask different types of interactive questions according to the situation: (i) open questions, such as “what is that sound / appliance?”. When a new pseudo tag appears, the system will ask the user about the appliance category, and replace the pseudo tag with the queried appliance category; (ii) confirmatory questions, such as “is that a microwave?”. Ask when the boundary of the known appliance category is too large (power difference exceeds 10%) and seems to contain other devices (contains another high local density area in the cluster).

[0098] The above only describes the preferred embodiments of the present application, and it should be noted that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements should also be considered within the protection scope of the present application.

Claims

1. A fine-grained household electricity monitoring method that combines a smart speaker and an electricity meter, characterized in that, The fine-grained household power monitoring system is realized by using a joint smart speaker and an electric meter, and the fine-grained household power monitoring system of the joint smart speaker and the electric meter comprises a cross-modal learning module and a cross-modal inference module, the cross-modal learning module is used for automatically learning the correlation between the power of the electrical appliance and the sound, i.e. consistency information and complementary information; the cross-modal inference module is used for utilizing the learned consistency information and complementary information, i.e. the identification result of the electrical appliance state recognizer, to realize electrical appliance energy consumption decomposition in a cross-modal correlation fusion manner, and further infer the power consumption of each type of electrical appliance; The method comprises the following steps: Step S1: the system automatically learns the correlation between the power of the electrical appliance and the sound, i.e. consistency information and complementary information; Step S1.1: the power event is divided into high power change and low power change, and scene discovery is iteratively performed; The DPGMMS model is used to divide power events into high power changes and low power changes, and then iteratively perform scenario discovery; one scenario E n = (e n1 ,e n2 ,…,e nL ) is a subsequence of an ordered power event sequence; if the power change P n = (P n1 ,P n2 ,…,P nL ) corresponding to the subsequence satisfies the axiom constraint, it is considered as a valid scenario set E valid , that is, the electrical appliance cycle; scenarios that have not been verified by the axiom constraint are referred to as candidate scenario set E cand ; Step S1.2: find the sound features and power features, i.e. key feature pairs, that have a correlation, and understand which key feature pairs belong to the same electrical appliance state through clustering; Step S1.2.1: sound feature extraction; According to a given sound data stream, it is divided into sound segments with length of 0.96 seconds, for the sound segment around time t, if its decibel value (DB) is greater than -45 DB, it is considered as a sound event a t ; using VGG-ish as a sound feature extractor to convert each 0.96s sound segment into a sound feature Step S1.2.2: encoding and decoding; using an encoder to encode sound features decoupled into noise feature vectors Z t,Noise and effective feature vectors Z t,Clean reconstruct sound features through a decoder The calculation formula of reconstruction loss is shown in equation (1): Step S1.2.3: power prediction; According to the decoupled effective feature vector Z t,Clean The predicted power value; the loss function of the power predictor is calculated as shown in equation (2): wherein p t is the actual power value, is the predicted power value; Step S1.2.4: constructing an overall loss function; The calculation formula of the overall loss function is shown in formula (3): Wherein, λ is a hyperparameter for balancing the prediction error and the reconstruction error; Step S1.3: noise-robust sound-based electrical appliance state recognizer training; Rich training data with noise is generated by moving and overlapping the target sound segment with other sound segments through data enhancement; new negative samples are generated by iteratively training the electrical appliance state recognizer to enhance the performance of the electrical appliance state recognizer; Step S2: utilize the learned consistency information and complementary information, i.e. the identification result of the electrical appliance state recognizer, to realize electrical appliance energy consumption decomposition in a cross-modal correlation fusion manner, and further infer the power consumption of each type of electrical appliance; Step S2.1: utilize the learned consistency information to verify the identification result of the noise-robust sound-based electrical appliance state recognizer; The identification result of the electrical appliance state recognizer is directly related to the power event at this moment, and the learned consistency information is used by the cross-modal learning module to verify the identification result of the noise-robust sound-based electrical appliance state recognizer, and if the power value does not match, it is discarded; Step S2.2: identify the electrical appliance cycle related to the sound, and utilize a two-stage scene discovery operation to discover the remaining candidate scenes after discovering the scene related to the sound; According to the state combination of different electrical appliances, different scene discovery rules are adopted: (i) if the start, running and end states are all identified, then according to the power value of each state, the final electrical appliance cycle is directly selected from the candidate scenes; (ii) if only part of the states are identified, then the candidate space of the candidate scenes is reduced, and the final scene is selected by using the axiomatic constraint; Step S2.3: cluster the electrical appliance cycle related to the sound, and identify through one-time interaction to obtain the label result of the final electrical appliance energy consumption decomposition; If an appliance cycle contains a power event related to the sound, it is marked as a pseudo label of the sound category; if all power events in an appliance cycle do not contain sound pseudo labels, similar appliance cycles are clustered and pseudo labels are identified using density peak clustering; at the same time, different types of interactive questions are asked according to the scene.

2. The method for joint smart speaker and electricity meter fine-grained home electricity monitoring of claim 1, wherein, The cross-modal learning module comprises a two-stage scene discovery component, a self-supervised consistency information extraction component and a noise-robust sound-based appliance state identifier training component; the two-stage scene discovery component and the self-supervised consistency information extraction component are used to extract consistency information in appliance power and sound data; the noise-robust sound-based appliance state identifier training component first performs data enhancement on the appliance power and sound data with consistency, and then performs iterative training, and obtains the running state of the appliance from the appliance sound data through the noise-robust sound-based appliance state identifier, that is, complementary information.

3. The method of claim 2, wherein the method further comprises: The cross-modal inference module comprises a cross-modal data association component, a sound-assisted scene discovery component and an appliance label association component; the cross-modal data association component is used to verify the identified sound-related appliance running state; the sound-assisted scene discovery component is used to identify the appliance cycle related to the sound, and after discovering the scene related to the sound, the remaining candidate scenes are discovered by using a two-stage scene discovery operation; the appliance label association component is used to cluster the appliance cycle related to the sound, and identify through one-time interaction to obtain the final label result of appliance energy consumption decomposition.

Citation Information

Patent Citations

  • Sound event detection method and device, equipment and storage medium

    CN114882911A