Compressing Information in End Nodes Using an Autoencoder Neural Network

By using pre-trained automatic encoder to compress and transmit spectrum diagrams in end-node devices, combined with cloud decoding and classifier processing, the problems of end-node devices' computing and power limiting are solved, and efficient data processing and low-power voice and image recognition are achieved.

CN114282644BActive Publication Date: 2025-07-18SILICON LABORATORIES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111092879.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-17
Filing Date
2021-09-17
Publication Date
2025-07-18
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

Due to computing and power limitations, end-node devices are difficult to effectively process sensed data such as voice or images, resulting in the transmission of raw data that requires a large amount of bandwidth and power consumption to the cloud for processing.

Method used

The sensed data is compressed into a spectrogram using a pre-trained automatic encoder in the end-node device and sent to a remote destination through a wireless circuit for further processing, combining cloud-based decoding and classifiers for data analysis.

Benefits of technology

It reduces the data transmission requirements of end-node devices, reduces power consumption and improves data processing efficiency, and supports local execution of complex tasks such as voice control and image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114282644B_ABST
    Figure CN114282644B_ABST
Patent Text Reader

Abstract

In one embodiment, a device includes: a sensor for sensing real-world information; a digitizer coupled to the sensor to digitize the real-world information into digitized information; a signal processor coupled to the digitizer to process the digitized information into a spectrogram; a neural engine coupled to the signal processor, the neural engine including an autoencoder for compressing the spectrogram into a compressed spectrogram; and a wireless circuit coupled to the neural engine to send the compressed spectrogram to a remote destination to enable the remote destination to process the compressed spectrogram.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] End-node devices typically sense and collect data about their environment. For example, such data can be audio or images. It is usually not allowed for end-nodes to perform meaningful computations on such data, such as performing speech recognition to detect keywords or detecting the presence of people in an image. Therefore, end-nodes typically send the raw data to the cloud, where powerful machine learning algorithms are used to process the raw data. However, the transmission of raw data requires bandwidth and power consumption, which can be prohibitively large for battery-powered devices or other power-constrained devices. Summary of the Invention

[0002] In one aspect, a device includes: a sensor for sensing real-world information; a digitizer coupled to the sensor to digitize the real-world information into digitized information; a signal processor coupled to the digitizer to process the digitized information into a spectrogram; a neural engine coupled to the signal processor, the neural engine including an autoencoder for compressing the spectrogram into a compressed spectrogram; and a wireless circuit coupled to the neural engine to send the compressed spectrogram to a remote destination so that the remote destination can process the compressed spectrogram.

[0003] In one example, the neural engine is used to store a model for the autoencoder. The model can include the structure of a neural network and a plurality of coefficients, and can be a pre-trained model generated using at least one correlation function according to real-world type information. The device can receive an updated model from a remote destination and update at least some of the coefficient weights of the model based on the updated model. The decoder of the autoencoder can decompress a first compressed spectrogram into a first reconstructed spectrogram, which is compressed from a first spectrogram by the encoder of the autoencoder.

[0004] In one example, the device can compare the first spectrogram with the first reconstructed spectrogram and send the first spectrogram rather than the first compressed spectrogram to the remote destination at least partially based on the comparison. The device can compare the first spectrogram with the first reconstructed spectrogram in response to a request from the remote destination.

[0005] In one example, the real-world information includes voice information, and the device includes a voice-controlled end-node device. The voice-controlled end-node device can receive at least one command from a remote destination at least partially based on the compressed spectrogram. The real-world information can be image information, and the device can take an action in response to a command from a remote destination based at least partially on the image information and detecting a person in the image information at the remote destination.

[0006] In another aspect, a method includes: generating an autoencoder including an encoder and a decoder, and generating a classifier, where the encoder is used to encode a spectrogram into a compressed spectrogram, the decoder is used to decode the compressed spectrogram into a restored spectrogram, and the classifier is used to determine at least one keyword from the decoded compressed spectrogram; calculating a first loss of the autoencoder and calculating a second loss of the classifier; jointly training the autoencoder and the classifier at least partially based on the first loss and the second loss; and storing the trained autoencoder and the trained classifier in a non-transitory storage medium.

[0007] In one example, the method further includes jointly training the autoencoder and the classifier based on a weighted sum of the first loss and the second loss. The method may further include: calculating the first loss according to a correlation coefficient; calculating the second loss according to binary cross-entropy. The method may further include sending the trained encoder part of the autoencoder to one or more end-node devices so that the one or more end-node devices can use the trained encoder part to compress the spectrogram. The method may further include generating the encoder asymmetrically from the decoder.

[0008] In yet another aspect, a voice-controlled device may include: a microphone for receiving a voice input; a digitizer coupled to the microphone to digitize the voice input into digitized information; a signal processor coupled to the digitizer to generate a spectrogram from the digitized information; a controller coupled to the signal processor, the controller including an encoder of an autoencoder for compressing the spectrogram into a compressed spectrogram corresponding to the voice input; and a wireless circuit coupled to the controller to send the compressed spectrogram to a remote server so that the remote server can process the compressed spectrogram. In response to a command from the remote server, the controller may cause the wireless circuit to send an uncompressed spectrogram corresponding to another voice input to the remote server.

[0009] In one example, the controller further includes a decoder of the autoencoder for decompressing one or more compressed spectrograms compressed by the encoder of the autoencoder into one or more reconstructed spectrograms. The controller may receive at least one second command from the remote server and implement an operation requested by the user in response to the at least one second command, where the voice input includes the operation requested by the user. The voice-controlled device may further include a cache for storing a model of the encoder. The model may include a structure of a neural network and a plurality of coefficient weights, where the model includes a pre-trained model generated using at least one correlation function for the spectrogram.

[0010] In one example, the voice-controlled device may further include a second cache for storing a second model of an encoder, where the second model includes the structure of a second neural network and a plurality of second coefficient weights. The second model may be a pre-trained model generated using at least one correlation function for image input. The encoder may be asymmetric with respect to the decoder of the autoencoder, which exists in a remote server to decompress the compressed spectrogram, and the decoder may have a larger kernel and more filters or a different architecture than the encoder. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a block diagram of an end-node device according to one embodiment.

[0012] Figure 2 is a block diagram of a cloud server according to one embodiment.

[0013] Figure 3 is a flowchart of a method according to one embodiment.

[0014] Figure 4 is a flowchart of a method according to another embodiment.

[0015] Figure 5 is a flowchart of a method according to yet another embodiment.

[0016] Figure 6 is a block diagram of a system for performing combined training of an autoencoder and a classifier according to one embodiment.

[0017] Figure 7 is a block diagram of a representative device that can be incorporated into a wireless network as an end node.

[0018] Figure 8 is a block diagram of a representative wireless automation network according to one embodiment. DETAILED DESCRIPTION

[0019] In various embodiments, the end-node device may be configured to sense real-world information, such as voice data, image data, etc. Further, the end-node device may minimally process the sensed information before sending the sensed information to a remote entity, such as a cloud-based destination for further processing the data.

[0020] Since end node devices typically have limited processing and power resources (e.g., end node devices can be battery-powered devices), more complex processing can be performed at a cloud-based destination. Although embodiments can vary, the implementations described herein can be in the context of voice-controlled end node devices, among other device types such as remote controls, home automation components such as lighting, security, HVAC components. For such voice-controlled devices, the sensed input can be voice, where the device can be configured to be awakened in response to a keyword. After awakening, the device can sense a voice string and encode the voice into a spectrogram or a sequence of spectrograms. Thereafter, the device can use a trained autoencoder to compress the (multiple) spectrograms into a compressed spectrogram (or a sequence of compressed spectrograms) format, which is then sent to a cloud-based destination.

[0021] In many cases, the end node device can be a wireless-enabled device such that communication to a remote source such as a cloud-based destination can be via a wireless local area network (WLAN), e.g., according to a given IEEE 802.11 standard, which is in turn coupled to the cloud-based destination via a network such as the Internet, e.g., through a gateway.

[0022] Although in some cases the device itself may use machine learning techniques to train the autoencoder, more typically due to the limited computational and power capabilities of the device, the autoencoder can be pre-trained and loaded into the end node device. Given the expected use of the end node device, the autoencoder can be pre-loaded into the end node device during manufacturing, e.g., loaded into the non-transitory storage medium of the device. Note that in such cases, the autoencoder can also be updated in the field in response to an update to the autoencoder model received from a cloud-based destination.

[0023] Furthermore, since the autoencoder is pre-trained for a specific type of input information, in some cases, the end node device can be configured with multiple autoencoders, each for processing a different type of information. For example, a first coding scheme can be used for voice input, and a second coding scheme can be used for another input type, e.g., image information. Thus, a suitable algorithm can be designed for the autoencoder according to the nature of the incoming data. Note that a given autoencoder implementation designed for a specific information type may not be optimally used for data of a different nature. In one embodiment having one or more autoencoders, an efficient compression of each suitable dataset can be automatically learned. In some embodiments, one or more relevant functions can be used to train the autoencoder depending on the characteristics of the type of input information. In contrast, mean squared error techniques are typically used to train autoencoders.

[0024] Now refer toFigure 1 , which shows a block diagram of an end - node device according to an embodiment. As Figure 1 shown, the device 100 can be any type of end - node device, such as an Internet of Things (IoT) device. In some cases, the device 100 can be a voice - controlled device such as those described above. For a small set of certain wake - up keywords and possibly other local - processing keywords, the device 100 can be configured to fully process and respond by taking a certain action. However, in many cases, the incoming voice can be minimally processed within the device 100 and then sent to a remote destination such as a cloud - based destination for further processing and disposition.

[0025] In Figure 1 the advanced configuration shown, the device 100 includes an input sensor 110. It should be understood that in different embodiments, there can be multiple sensors, each sensing a different type of real - world information, such as environmental conditions, image information, voice or other audio information, etc. Thus, the input sensor 110 can include one or more microphones and possibly other sensors, such as image sensors, video sensors, thermal sensors, accelerometers, passive infrared sensors, etc.

[0026] The input sensor 110 is then coupled to a digitizer 120, such as an analog - to - digital converter, which digitizes the information and provides it to a signal processor 130 that can perform various signal processing, such as filtering, amplification, etc. Note that in some embodiments, at least some of this signal processing can be performed in the analog domain before the digitizer 120. In any case, the digitized and processed information can be provided to a processor 140, which can be the main processor of the system of the device 100, such as a microcontroller, a processing core, or other general - purpose processing circuitry or other such controller.

[0027] Relevant to the discussion herein, the processor 140 can perform additional signal processing on the information, including converting the digitized information into a spectrogram. The processor 140 is then coupled to a neural engine 150, which in some cases can be implemented as a dedicated hardware circuit. In different embodiments, the neural engine 150 can be implemented with fully - connected layers or convolutional neural network layers. Note that in some cases, at least two of the signal processor 130, the processor 140, and the neural engine 150 can be implemented together.

[0028] In any case, the neural engine 150 can compress the spectrogram according to the trained parameters, which can be stored in the flat cache 155. In one embodiment, such parameters can include the structure of the neural network and a plurality of weight coefficients. Note that in some implementations, only the encoder part of the autoencoder exists. In other implementations, both the encoder and decoder parts of the autoencoder can exist. The processor 140 can also store the uncompressed spectrogram in the memory 160 at least temporarily.

[0029] To enable remote processing of the sensed information, the processor 140 can receive the compressed spectrogram from the neural engine 150 and provide it to the wireless circuitry 170, which can wirelessly transmit the compressed information via the antenna 180. It should be understood that the antenna 180 can also be configured to receive incoming information, such as command information from a cloud-based destination, to enable the device 100 to perform an action in response to the received command information. Although shown in this high-level configuration in Figure 1 the embodiment, many variations and substitutions are possible.

[0030] Now refer to Figure 2 , which shows a block diagram of a cloud server 200, which can be, for example, a remote server present in a data center. The cloud server 200 can be configured to receive incoming compressed spectrograms from a large number of remotely located end-node devices, decompress and process the decompressed information as described herein.

[0031] In addition, Figure 2 shows the components of the cloud server 200 involved in the initial training of an autoencoder model for a particular type of sensed information. Although all these components are shown in a single cloud server 200 for ease of illustration, it should be understood that in some cases, these components can also exist in different servers. And, in fact, the initial training (and training updates) as described herein can be performed by different entities on different hardware (such as different (e.g., cloud) servers). For example, the pre-training can be performed by an IC or IoT device manufacturer, who performs the training to load the autoencoder into the device before it is sold. In other cases, the pre-training can be performed by a system distributor, who implements the device into a network architecture such as a home automation network. And, in some cases, the processing of the compressed information received from the end-node device can also be performed by the same entity.

[0032] As shown in the figure, the cloud server 200 includes an autoencoder 210, which includes an encoder 212 and a decoder 214. In one embodiment, the autoencoder 210 can be a symmetric autoencoder with a 3×3 kernel and 16 filters. However, in other cases, the encoder 212 can be asymmetric with respect to the decoder 214, thus providing a simpler and less computationally intensive encoder, since encoding is typically performed in the end-node device. In an exemplary embodiment, the asymmetric autoencoder can include 2x filters, a 5x5 internal kernel, and a 7x7 output kernel. As another example, the asymmetric autoencoder can include 4x filters, a 7x7 internal kernel, and an 11x11 output kernel. In some embodiments, the autoencoder 210 can perform compression with a compression ratio of up to approximately 16x and a low reconstruction loss (although higher compression ratios can be achieved if a certain amount of reconstruction loss is allowed). In this way, a better compression ratio can be achieved by using a relatively simple (e.g., low loop count) encoder in the end-node device and a more complex (e.g., higher loop count) decoder in the cloud server 200.

[0033] As shown in the figure, the cloud server 200 can send a flat cache to an end-node device of a model that includes at least the encoder 212. The flat cache can include a neural network structure and a plurality of coefficient weights to implement the encoder 212 of the autoencoder 210. In some cases, in instances where it is desired to provide the decoder to the end-node device for use in performing local quality checks as described herein, the flat cache (or another flat cache) can also include the structure and coefficient weights for the decoder 214.

[0034] As further shown, the autoencoder 210 also receives compressed information from the connected end-node device, for example, in the form of a compressed spectrogram. This compressed spectrogram is provided to the decoder 214, which can decompress them. As shown in the figure, the decompressed spectrogram is provided to a spectrogram processor 240, which can be implemented as a classifier. This machine learning classifier can process the spectrogram to identify attributes of the originally sensed real-world information (such as voice information, image information, etc.).

[0035] The classification result can be provided from the classifier to a command interpreter 250, which can interpret these results into one or more commands to send back to the end-node device. For example, for voice information, the keywords identified by the classifier can be used to generate commands for actions to be taken by the end-node device. These actions can include playing a specific media file, performing an operation in an automated network, etc. In the case of image information, the classified result can indicate the presence of a person in the local environment, which can trigger the command interpreter 250 to send a specific command, such as turning on the light, triggering an alarm, etc.

[0036] Still referring to Figure 2 , when the spectrogram processor 240 has difficulty classifying the received information, for example due to poor quality, it can send a request for the uncompressed information to the end node device. The end node device can send the uncompressed information in response to the request, such as various uncompressed spectrograms that can be directly provided to the spectrogram processor 240 (since decoding does not need to be performed in the autoencoder 210). Further, such uncompressed information in the form of a spectrogram, for example, can be provided to the training logic 220, which can perform incremental training based on this incoming information.

[0037] Still referring to Figure 2 , the training logic 220 can also use the information in the dataset 230 to perform the initial training of the autoencoder 210, and this information can be various samples of the training dataset to enable pre-training to occur. It should be understood that although shown in this high-level configuration in the Figure 2 embodiment, many various alternatives are possible.

[0038] Now referring to Figure 3 , which shows a flowchart of a method according to an embodiment. More specifically, the method 300 is a method performed by an end node device for receiving incoming real-world information, minimally processing it, and sending it to a remote destination. Although for illustration, Figure 3 is in the context where the real-world information is audio information, it is understood that the embodiment is not limited to this aspect, and in other embodiments, the real-world information can have various different information types.

[0039] As shown, the method 300 starts by receiving an audio sample (block 310) in the end node device. For example, the end node device can include a microphone or other sensors to detect an audio input, such as the user's speech. At block 320, the audio sample can be digitized after certain analog front-end processing such as filtering, signal conditioning, etc. Next, at block 330, the digitized audio sample can be processed. In an embodiment, the sample can be processed into a spectrogram. Of course, other digital representations of the audio sample can also be generated in other embodiments, such as those that can be suitable for other types of input information. For example, in another case, a single fast Fourier transform (FFT) can be used for a time-stationary signal.

[0040] Still referring to Figure 3, control then proceeds to block 340 where the spectrogram can be compressed. In one embodiment, the spectrogram compression can be performed in a neural engine, which can be implemented in the hardware processor of the end node device. More particularly, in the embodiments herein, an autoencoder with a pre-trained model can perform the compression. Thereafter, at block 350, the compressed spectrogram can be sent to a cloud-based destination, e.g., a remote server. The remote server can further process the received audio sample, e.g., to determine various actions that the end node device is to take in response to the audio sample (e.g., a voice command). It should be understood that although shown in this high-level configuration in the Figure 3 embodiment, many variations and alternatives are possible.

[0041] Now referring to Figure 4 , a flowchart of a method according to another embodiment is shown. More specifically, Figure 4 method 400 is a method for identifying when it may be inappropriate to send compressed information from an end node device to a destination. Thus, method 400 can be performed within the end node device itself. In some cases, method 400 can be performed according to a periodic interval to periodically determine the quality of the compressed information sent from the end node device (where the periodic interval can be controlled to reduce power consumption). In other cases, method 400 can be performed in response to receiving an indication that incoming compressed information from a remote destination does not have an appropriate quality.

[0042] As shown, method 400 begins by decompressing the compressed spectrogram into a reconstructed spectrogram (block 410). In one embodiment, the autoencoder can perform this decompression, e.g., using a pre-trained decoder. It should be understood that for this operation to occur, the end node device includes a full autoencoder (both an encoder and a decoder). Next, at block 420, the end node device can generate a comparison result between its generated original spectrogram and the reconstructed spectrogram obtained by decompressing the compressed spectrogram. As an example, some type of difference calculation can be performed between the two spectrograms. Next, at diamond block 430, it is determined whether the comparison result exceeds a threshold. The threshold can take various forms, but can also be a threshold that measures the quality level of the compression performed within the end node device.

[0043] If the comparison result does not exceed the threshold, the current autoencoder model is appropriate, and thus control proceeds to block 440, where the end node device can continue to send the compressed spectrogram to the cloud-based destination. Otherwise, when it is determined that the comparison result exceeds the threshold, this indicates that the compression performed by the autoencoder did not provide a suitable result. Thus, at block 450, the full uncompressed spectrogram can be sent to the cloud-based destination. Additionally, a comparison indicator or other feedback information can also be sent to the cloud-based destination. Note that in response to receiving this notification, the cloud-based destination can perform an update to the autoencoder model, for example, by updating the training. At the end of such a training update, the updated model parameters can be sent to the end node device (and other end node devices) to effect an update to the autoencoder model. It should be understood that although shown in this high-level configuration in the Figure 4 embodiment, many variations and alternatives are possible.

[0044] Now referring to Figure 5 , which shows a flowchart of a method according to another embodiment. Specifically, method 500 is a method for performing the training of an autoencoder and a classifier as described. In one embodiment, method 500 can be performed by one or more cloud servers or other computing systems. Thus, method 500 can be performed by hardware circuitry, firmware, software, and / or a combination thereof.

[0045] Method 500 begins by generating an autoencoder and a classifier (block 510). Note that the generation of the autoencoder and the classifier can be according to a given model, where the autoencoder includes an encoder and a decoder. In different embodiments, the encoder and the decoder can be symmetric or asymmetric. Subsequently, the classifier can be generated as a machine learning classifier to classify the input into corresponding label categories. As an example, among many other classification problems, the input can be classified into a specific keyword, detection of a person within an image, pose detection.

[0046] Still referring to Figure 5 , next at block 520, the loss of the autoencoder and the loss of the classifier can be calculated independently. As an example, the autoencoder loss can be determined as one of the following: the sum of the absolute values of the differences between the input and the output (L1 norm); the sum of the squares of the differences between the input and the output (L2 norm); and the correlation between the input and the output (Pearson correlation). In one embodiment, the classifier loss can be determined as the binary cross-entropy between the predicted class and the true class. In some cases, a joint loss (e.g., L2 and cross-entropy) can be determined.

[0047] Next, at block 530, the autoencoder and the classifier can be jointly trained. This joint training can be performed to minimize a loss function. In one embodiment, the joint training can be based on a weighted sum of losses. Then, at block 540, the trained autoencoder and classifier can be stored in one or more non-transitory storage media, such as may be present in one or more cloud servers.

[0048] At this point, the model is properly trained and can be used to encode and decode information and then classify the resulting information. To enable cloud-based processing of sensed information from end-node devices, at block 550, the trained autoencoder can be sent to one or more such devices. In some cases, a full autoencoder including both an encoder and a decoder can be sent. In other cases, especially in the context of complexity-reduced IoT devices, only the encoder part of the autoencoder can be sent. At this time, the cloud server can start receiving compressed data from one or more devices, e.g., in the form of compressed spectrograms.

[0049] Note that over time, incremental training of one or more of the autoencoder and the classifier may be performed. As an example, such incremental training can be performed at a predetermined interval. Alternatively, when it is determined that lower-quality compressed data is being received from the terminal device, the classifier can trigger such incremental training. When this incremental training is triggered, control passes from diamond box 560 to block 570, where one or more uncompressed spectrograms can be received from one or more end-node devices in response to a request from the cloud server. These uncompressed spectrograms can be used to further train the autoencoder and / or the classifier. Control returns to block 520, where, as described above, the incremental training can be performed similar to the initial training. It should be understood that although shown in such a high-level configuration in the Figure 5 embodiment, many variations and alternatives are possible.

[0050] Now referring to Figure 6 , a block diagram of a system for performing combined training of an autoencoder and a classifier according to one embodiment is shown. As Figure 6 shown, the system 600 can be implemented within a computing device such as a server. The combined training can occur between an autoencoder formed by an encoder 610 and a decoder 630 and a classifier 640.

[0051] Figure 6The data flow through these components is also shown. More specifically, an input X, which can be an uncompressed spectrogram, is encoded within an encoder 610 to generate a code 620 corresponding to the compressed spectrogram. This compressed spectrogram is provided as an input to a decoder 630, which decompresses the code 620 to generate a reconstructed uncompressed spectrogram X. Subsequently, this value is provided as an input to a classifier 640, which generates a classification label y corresponding to a target word Y based on the original spectrogram X.

[0052] By performing training within the system 600, multiple loss functions can be minimized. In one embodiment, the loss function of the autoencoder can be one of the following: L autotcoder =|X - X|1 or |X - X|2 or ρ(X, X), where ρ is the Pearson correlation coefficient. Subsequently, the loss function of the classifier 640 can be: L classifier =binary_crossentropy(Y, y).

[0053] Note that these two resulting losses can be combined according to, for example, a weighted sum to perform combined training. For example, backpropagation can be performed for these multiple losses. In one embodiment, independent backpropagation can be performed for each loss. In another case, the combined loss can be backpropagated. The resulting trained parameters corresponding to the network structure and the weighting coefficients can be stored in a non-transitory storage medium for use both within a cloud-based destination and for providing at least the encoder of the autoencoder to one or more end-node devices.

[0054] Similarly as described above, multiple trainings can be performed, where each training is used to train the autoencoder (and the classifier) for different information types (such as speech, images, videos, etc.).

[0055] Embodiments can be implemented in many different types of end-node devices. Now referring to Figure 7 , a block diagram of a representative device 700 that can be incorporated as a node into a wireless network is shown. And as described herein, the device 700 can perform compression and transmit the compressed spectrogram to a remote server and subsequently receive commands from the remote server. In Figure 7 the embodiment shown, the device 700 can be a sensor, an actuator, a controller, or other device that can be used in various usage scenarios within a wireless control network, including sensing, metering, monitoring, embedded applications, communication applications, etc.

[0056] In the illustrated embodiment, device 700 includes a memory system 710 which, in one embodiment, may include non-volatile memory such as flash memory and volatile storage devices such as RAM. In one embodiment, the non-volatile memory may be implemented as a non-transitory storage medium capable of storing instructions and data. As described herein, such non-volatile memory may store code and data (e.g., trained parameters) for one or more autoencoders, and may also store code for performing methods including Figure 3 and 4 of the method.

[0057] Memory system 710 is coupled to digital core 720 via bus 750, which may include one or more cores and / or microcontrollers that act as the main processing unit of the device. As shown, digital core 720 includes neural network 725 which may perform compression / decompression of spectrograms as described herein. As further shown, digital core 720 may be coupled to clock generator 730 which may provide one or more phase-locked loops or other clock generation circuits to generate various clocks for use by the circuits of the device.

[0058] As further shown, device 700 also includes power circuit 740 which may include one or more voltage regulators. Depending on the particular implementation, additional circuitry may optionally be present to provide various functions and interaction with external devices. Such circuitry may include interface circuit 760 which may provide an interface to various off-chip devices, and sensor circuit 770 which may include various on-chip sensors, including digital and analog sensors, to sense desired signals such as voice input, image input, etc.

[0059] Additionally, as Figure 7 shown, transceiver circuit 780 may be provided to enable transmission and reception of wireless signals, for example, according to one or more local wireless communication schemes such as Zigbee, Bluetooth, Z-Wave, Thread, etc. It should be understood that while shown in this high-level view, many variations and alternatives are possible.

[0060] Now referring to Figure 8 , a block diagram of a representative wireless automation network according to one embodiment is shown. As Figure 8 shown, network 800 may be implemented as a mesh network in which various nodes 810 (810 0-n ) communicate with each other and may also convey messages from a particular source node to a given destination node. Such message-based communication may also be implemented between various nodes 810 and network controller 820.

[0061] It should be understood that while in Figure 8This is shown in a very high level, but network 800 can also take many forms other than a mesh network. For example, in other cases, a star network, a point-to-point network, or other network topologies may be possible. Additionally, although general nodes 810 are shown for ease of illustration, it should be understood that there can be many different types of nodes in a particular application. For example, in the context of a home automation system, nodes 810 can have many different types. As an example, nodes 810 can include sensors, lighting components, door locks or other actuators, HVAC components, security components, window covers and possibly entertainment components, or other interface devices to enable control of such components via network 800. Of course, in different embodiments, there can also be additional or different types of nodes.

[0062] In addition, different nodes 810 can communicate according to different wireless communication protocols. As an example, representative communication protocols can include Bluetooth, Zigbee, Z-Wave, and Thread, among other possible wireless communication protocols. In some cases, certain nodes may be able to communicate according to multiple communication protocols, while other nodes may only be able to communicate according to a given one of the protocols. Within network 800, certain nodes 810 can communicate with other nodes of the same communication protocol to provide direct message communication or to enable mesh-based communication with network controller 820 or other components. In other instances, for example, for some Bluetooth devices, communication can occur directly between a given node 810 and network controller 820.

[0063] Thus, in Figure 8 the embodiment, network controller 820 can be implemented as a multi-protocol general gateway. Network controller 820 can be the main controller within the network and can be configured to perform operations to establish the network and enable communication between different nodes, as well as update such establishment when devices enter and leave network 800.

[0064] In addition, network controller 820 can also be an interface for interacting with remote devices such as cloud-based devices. To this end, network controller 820 can also communicate with remote cloud server 840 via, for example, the Internet. Remote cloud server 840 can include a processor, a memory, and a non-transitory storage medium, which can be used to generate and pre-train an autoencoder and perform other operations described herein. As also shown, one or more user interfaces 850 that can be used to interact with network 800 can be located remotely and can communicate with network controller 820 via Internet 830. As an example, such a user interface 850 can be implemented within a mobile device (such as a smart phone, a tablet computer, etc.) of a user authorized to access network 800. For example, the user can be the homeowner of a household in which wireless network 800 is implemented as a home automation network. In other cases, the user can be an authorized employee, such as an IT individual, a maintenance individual, etc., who uses the remote user interface 850 to interact with network 800, for example, in the context of a building automation network. It should be understood that many other types of automation networks (such as industrial automation networks, smart city networks, agricultural crop / livestock monitoring networks, environmental monitoring networks, store shelf label networks, asset tracking networks, or health monitoring networks, etc.) can also utilize the embodiments described herein.

[0065] Although the invention has been described with reference to a limited number of embodiments, those skilled in the art should realize that many modifications and variations can be made therefrom. The appended claims are intended to cover all such modifications and variations that fall within the true spirit and scope of the invention.

Claims

1. A device, comprising: A sensor for sensing real - world information; A digitizer coupled to the sensor to digitize the real - world information into digitized information; A signal processor coupled to the digitizer to process the digitized information into a spectrogram; A neural engine coupled to the signal processor, the neural engine including an auto - encoder for compressing the spectrogram into a compressed spectrogram; And A radio circuit coupled to the neural engine to send the compressed spectrogram to a remote destination so that the remote destination can process the compressed spectrogram; Wherein, the decoder of the auto - encoder is used to decompress a first compressed spectrogram into a first reconstructed spectrogram, the first compressed spectrogram is compressed from a first spectrogram by the encoder of the auto - encoder, the device is used to compare the first spectrogram with the first reconstructed spectrogram, and at least partially based on the comparison, send the first spectrogram rather than the first compressed spectrogram to the remote destination.

2. The device according to claim 1, wherein The neural engine is used to store a model for the auto - encoder, the model including the structure of a neural network and a plurality of coefficients, wherein the model includes a pre - trained model generated using at least one correlation function according to the type of the real - world information.

3. The device according to claim 2, wherein The device is used to receive an updated model from the remote destination and update at least some of the coefficient weights of the model based on the updated model.

4. The device according to claim 1, wherein, The device is used to compare the first spectrogram with the first reconstructed spectrogram in response to a request from the remote destination.

5. The device according to claim 1, wherein The real - world information includes voice information, and the device includes a voice - controlled end - node device.

6. The device according to claim 5, wherein The voice - controlled end - node device is used to receive at least one command from the remote destination that is at least partially based on the compressed spectrogram.

7. The apparatus according to claim 1, wherein The real - world information includes image information, and the device is used to take an action at least partially based on the image information in response to a command from the remote destination and based on detecting a person in the image information at the remote destination.

8. A voice - controlled device, comprising: A microphone for receiving voice input; A digitizer coupled to the microphone to digitize the voice input into digitized information; A signal processor coupled to the digitizer to generate a spectrogram from the digitized information; A controller coupled to the signal processor, the controller including an encoder of an auto - encoder for compressing the spectrogram into a compressed spectrogram corresponding to the voice input; And A radio circuit coupled to the controller to send the compressed spectrogram to a remote server so that the remote server can process the compressed spectrogram, wherein in response to a command from the remote server, the controller is used to cause the radio circuit to send an uncompressed spectrogram corresponding to another voice input to the remote server.

9. The voice-controlled device according to claim 8, wherein, The controller further includes a decoder of the autoencoder for decompressing one or more compressed spectrograms compressed by the encoder of the autoencoder into one or more reconstructed spectrograms.

10. The voice-controlled device according to claim 8, wherein, The controller is configured to receive at least one second command from the remote server and implement an operation requested by the user in response to the at least one second command, wherein the voice input includes the operation requested by the user.

11. The voice-controlled device according to claim 8, wherein, It further includes a cache for storing the model of the encoder, the model including the structure of a neural network and a plurality of coefficient weights, wherein the model includes a pre-trained model generated using at least one correlation function of the spectrogram.

12. The voice-controlled device according to claim 11, wherein, It further includes a second cache for storing a second model of the encoder, the second model including the structure of a second neural network and a plurality of second coefficient weights, wherein the second model includes a pre-trained model generated using at least one correlation function of the image input.

13. The voice-controlled device according to claim 11, wherein, The encoder is asymmetric with the decoder of the autoencoder, and the decoder is present in the remote server to decompress the compressed spectrogram, wherein the decoder has a larger kernel and more filters or a different architecture than the encoder.

Citation Information

Patent Citations

  • Stereo audio signal encoder

    US20160027445A1