Method and computer-readable medium for automatic learning from sensors

The sensor data is processed through convolutional neural networks, and the co-occurrence loss function and shared feature space are used to solve the problem of insufficient training data in machine learning, realizing efficient and accurate learning of automated processing of sensor data.

CN113177572BActive Publication Date: 2025-08-15FUJIFILM BUSINESS INNOVATION CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011313639.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2020-11-20
Publication Date
2025-08-15
Estimated Expiration
2040-11-20

AI Technical Summary

Technical Problem

In the prior art, machine learning methods lack sufficient quantity and quality training data, especially the high cost and low efficiency of labeling visual and speech data, resulting in the inability to effectively utilize sensor data.

Method used

Convolutional neural network (CNN) is used to process sensor data of different modes, and automatically learn the association of visual and speech data through co-occurrence loss functions and shared feature spaces to generate shared feature spaces to provide classification or probability outputs.

Benefits of technology

It realizes automated processing of sensor data without manual tagging and storage, reducing dependence on human marking journalists, and improving the efficiency and accuracy of machine learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113177572B_ABST
    Figure CN113177572B_ABST
Patent Text Reader

Abstract

Method and computer-readable medium for automatically learning from sensors. A computer-implemented method includes: receiving a first input associated with a first modality and a second input associated with a second modality; processing the received first and second inputs using a convolutional neural network (CNN), wherein a first set of weights is used to process the first input and a second set of weights is used to process the second input; determining a loss for each of the first and second inputs based on a loss function that applies the first set of weights, the second set of weights, and the presence of co-occurrences; generating a shared feature space as an output of the CNN, wherein distances between cells associated with the first and second inputs in the shared feature space are determined based on the losses associated with each of the first and second inputs; and providing an output based on the shared feature space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of example implementations relate to methods, systems, and user experiences associated with learning directly from multiple inputs with sensor data of different modalities, and more particularly, to using audiovisual datasets in neural networks based on temporal co-occurrence of inputs. Background Art

[0002] Prior art machine learning methods for neural networks may lack sufficient quantity and quality of training data. This prior art problem may arise due to the shortage of human labelers, labeling costs, and labeling time constraints. In addition, prior art machine learning methods for neural networks rely primarily on text because available datasets utilize text labeling. In terms of providing speech and visual data for machine learning in neural networks, prior art methods still require text labels to learn speech and vision.

[0003] Shortages of training data can occur due to various factors, including but not limited to a shortage of human labelers, the cost of performing labeling, issues associated with verification and validation of labeling quality, and time constraints that result in delays in making labeled data available for machine learning in neural networks.

[0004] For example, existing machine learning methods may lack training data. On the other hand, existing sensors (e.g., Internet of Things (IoT) sensors) continuously transmit data. However, existing machine learning methods cannot use sensor data without recording and manual labeling. This existing recording and manual labeling activity can deviate from or undermine genuine human / machine learning heuristics.

[0005] For example, magnetic resonance imaging (MRI) and positron emission tomography (PET) of the prior art can be used for medical diagnosis. MRI scans can use magnetic fields and radio waves to form images of target tissues (e.g., organs and other structures inside the human body), and PET scans can use radioactive tracers to diagnose diseases by examining body functions at the cellular level. MRI scans are not as invasive as PET scans, and are cheaper, more convenient, and less harmful to people. However, PET scans provide better visualization of specific characteristics (e.g., metabolism, blood flow, oxygen use, etc.). Therefore, it may be necessary to perform a PET scan to obtain the necessary information for diagnosing a disease (e.g., dementia). Therefore, medical providers collect prior art PET / MRI image pairs. However, these prior art data pairs are not mapped in the same feature space.

[0006] Therefore, there exists an unmet need to overcome the problems associated with prior art techniques for obtaining training data and using sensor data in machine learning activities. Summary of the Invention

[0007] According to one aspect of an example implementation, a computer-implemented method is provided for: receiving a first input associated with a first modality and a second input associated with a second modality; processing the received first and second inputs in a convolutional neural network (CNN), wherein a first weight is assigned to the first input and a second weight is assigned to the second input; determining a loss for each of the first and second inputs based on a loss function applying the first weight, the second weight, and the presence of co-occurrence; generating a shared feature space as an output of the CNN, wherein a distance between cells associated with the first and second inputs in the shared feature space is determined based on the loss associated with each of the first and second inputs; and providing an output indicating a classification or a probability of a classification based on the shared feature space.

[0008] According to another aspect of an example implementation, a computer-implemented method is provided for: receiving historical pairing information associated with a pairing of a positron emission tomography (PET) image and a magnetic resonance imaging (MRI) image; providing the historical pairing information of the PET image and the MRI image to a neural network including a PET learning network and an MRI learning network to generate PET network output features and MRI network output features for a shared feature space and learn mapping weights of the PET network and the MRI network; providing an unseen MRI image; and generating an output providing a shared feature space of the PET image associated with the unseen MRI image based on the historical pairing information and a loss function.

[0009] Example implementations may also include a non-transitory computer-readable medium having storage and a processor capable of executing instructions for learning directly from multiple inputs having sensor data of different modalities. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] This patent or application file contains at least one drawing drawn in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0011] Figure 1 Various aspects of the architecture according to example implementations are shown.

[0012] Figure 2 Example results associated with implementation of the architecture according to example implementations are shown.

[0013] Figure 3 Mapping of MRI and PET images into the same feature space is shown according to an example implementation.

[0014] Figure 4Processing of PET / MRI images using a parallel architecture according to an example implementation is shown.

[0015] Figure 5 Example procedures are shown for some example implementations.

[0016] Figure 6 An example computing environment is shown with an example computer device suitable for use with some example implementations.

[0017] Figure 7 An example environment suitable for some example implementations is shown. DETAILED DESCRIPTION

[0018] The following detailed description provides further details of the drawings and example implementations of the present application. For clarity, the numbers and descriptions of redundant elements between the drawings are omitted. The terms used throughout the specification are provided as examples and are not intended to be limiting.

[0019] Various aspects of example implementations relate to systems and methods associated with automatic learning from sensors and neural network processing. Example implementations provide algorithms that can learn directly from sensors and that can reduce the reliance of machine learning on the availability of human labelers, reduce labeling costs, and / or facilitate machine learning for a variety of applications.

[0020] Therefore, in the example implementation, audio-visual datasets are used to test the learning model and training performance. The example implementation can be applied to sensor data (e.g., IoT sensor data) having one or more modalities (e.g., audio modality and visual modality). For example, but not limitation, medical image conversion from different sources and multiple patients can employ the example implementation based on medical imaging data.

[0021] IoT sensors have a large amount of data available continuously over time. The data is received in various modes, such as visual, audio, temperature, humidity, etc. The data received from IoT sensors is in different modes, and the data from different modes may not be comparable. For example, an image received from a camera may be sent as a photo image, while audio received from a microphone may be sent as a sound file. The formats of sensor data from different modes cannot be labeled in a standardized manner except using labor-intensive and time-consuming prior art methods that may require conversion to another format (e.g., text).

[0022] Example implementations include an architecture and one or more training methods for automating the learning and use of sensor data in neural networks within and across modalities. For example, and not by way of limitation, sensor input may involve audio-visual learning (e.g., learning of audio-visual relationships). Instead of requiring explicit labeling devices (e.g., keyboards or mice) to be associated with a user, sensors may be used to continuously provide audio-visual training data without requiring labeling devices to be attached to the user.

[0023] According to an example implementation, a time window approach is employed for training. For example, if data such as images and sounds appear sequentially within a close time interval, they are grouped together. On the other hand, if data appears far apart in time, it is not grouped together. Furthermore, by not requiring storage, conversion, and / or manual labeling, IoT sensors can more closely match real human behavior by allowing sensor input and its processing to be uninterrupted by storage, manual labeling, recording, or other processing activities.

[0024] Figure 1 An architecture for sensor-mediated autonomous learning according to an example implementation is shown. Architecture 100 includes multiple inputs 101, 103, 105 of a first modality (in this example implementation, visual). For example, visual input can be received from a sensor associated with image capture (e.g., a camera). Architecture 100 also includes inputs 107, 109 of a second modality (in this example implementation, audio). For example, audio input can be received from a sensor associated with audio capture (e.g., a microphone). Thus, inputs 101, 103, 105 can be image capture files and inputs 107, 109 can be audio files. As described herein, the inputs maintain their modality and are not converted to text or other modalities.

[0025] The neural networks receive inputs. For example, a first input 101 of the visual modality may be received by a corresponding first convolutional neural network (CNN) 111. Similarly, inputs 103 and 105 may be received by a second CNN 113 and a third CNN 115, respectively. Furthermore, a first input 107 of the audio modality may be received by a corresponding first CNN 117, and subsequent input 109 may be received by a second CNN 119.

[0026] The outputs of CNNs 111-119 are provided to fully connected layers 121-133, respectively. In addition to the fully connected layers or shared feature space, at least one additional CNN 135 and 137 are provided as anchors for each modality, respectively. Thus, outputs 139 and 141 are generated indicating the classification of the object in the input or the probability associated with the classification.

[0027] Weighting methods are also provided for different modalities. For example, 111, 113, and 115 share the same weight (as indicated by the dashed box 143), 117 and 119 share the same weight (as indicated by the dashed box 147), 123, 125, and 127 share the same weight (as indicated by the dashed box 145), and 131 and 133 share the same weight (as indicated by the dashed box 149).

[0028] exist Figure 1 The shared feature space is represented as line 157. Although the shared space is represented as a line at 157, one of ordinary skill in the art will understand that the representation is not limited to a linear representation and may alternatively be a higher dimensional space in the model according to example implementations.

[0029] As described below, methods according to example implementations determine co-occurrence and connection strengths between different inputs. Figure 1 As shown, inputs with stronger connections based on co-occurrence are marked with a "+", while those with weaker connections based on lack of co-occurrence are marked with a "-". The process of allowing strengthening and weakening is described in more detail below. In the shared representation space or fully connected layer, stronger connections can be marked by points showing connections at 155, while weaker connections can be marked by unconnected points such as 151 and 153.

[0030] Audio training can be achieved by converting the audio input (e.g., into a 2-D mel-spectrogram) so that images can be processed in a similar manner. Interference between image and audio channels can be avoided by using different weights in the audio encoder and image encoder. Similarly, different weights can be used at the audio decoder and image decoder.

[0031] like Figure 1 As shown, each network within a dashed box has common weights. For example, but not limitation, if three convolutional neural networks (CNNs) are within a dashed box due to their co-occurrence, then those three CNNs can be assigned common weights. If three fully connected networks are within a common dashed box, then those three fully connected networks can be assigned common weights. According to an example implementation, weights are not shared across dashed boxes.

[0032] As understood in the field of neuroscience, any two units or systems of units that are repeatedly active at the same time may become "associated", so that the activity in one unit or system promotes the activity in another unit or system. The example implementation herein uses a shared feature space to simulate the method so that the time span of the activity (e.g., a few seconds) can lead to promoting the activity in the network. More specifically, if two or more feature vectors become associated, the model of the example implementation will force those feature vectors to become closer in the feature space. For feature vectors that appear at different times, the model will force those feature vectors to move further away from each other in the feature space. Regardless of whether the data is from the same modality or from different modalities, this forcing can be performed.

[0033] Through this training, the model according to the example implementation can form an embedded feature space shared by the visual modality and the speech modality. This shared feature space learning process can correspond to the Siamese architecture feature space formation process. However, unlike the Siamese architecture feature space formation process, this example implementation can be used for cross-modality and time learning.

[0034] According to one example implementation of the architecture, the model can simulate a teaching process, such as that associated with teaching a child. For example, and not limitation, to teach a child to pronounce an object, the object and the corresponding pronunciation are provided to the child substantially simultaneously.

[0035] Thus, children are presented with successive “visual frames” of an object at different angles, scales, and lighting conditions; because the images appear essentially simultaneously, their features in this architecture are forced together to simulate the association process described above. This feature can be viewed as a self-supervised process or feature space formation, for example.

[0036] like Figure 1 As shown, two images are provided as input associated with product packaging, where the images are captured at different angles. Since these images are simultaneous relative to each other during the teaching process, their features are forced to a common place in the shared feature space indicated by point 155. After the shared feature space 157 is well formed, it can be an embedding space for both visual data and audio data.

[0037] For example, input associated with this example implementation can be received via an IoT sensor. In some example implementations, the IoT sensor can combine audio and visual modalities to provide training and machine learning while also avoiding the need for additional storage and communication costs because no data needs to be stored or manually labeled.

[0038] For example, and not limitation, the example implementations described above can be integrated into robotic systems to provide the robot with the ability to more closely match human behavior and provide better service. A robot acting as a caregiver for the elderly or injured may be able to use audio and visual IoT as inputs in the above methods to more quickly and accurately provide appropriate care and attention.

[0039] Furthermore, this example implementation may have other applications. For example, and not limitation, the example implementation may be integrated with electronic devices such as smartphones or home automation systems. Because this example implementation can be used on IoT sensors without the latency, storage, and manual labeling costs, these devices can be customized for specific users. This customization can be performed without providing data to third parties. For example, instead of storing and manually labeling data, the example implementation can perform learning and processing locally in a privacy-preserving manner.

[0040] In the example implementation described above, in addition to supervision based on the same modality, it is noted that the spoken name (e.g., Tylenol) may be pronounced when displayed. Therefore, the audio features in the representation space are also forced to the same position at 155. Figure 1 All sensor inputs associated with each other are shown with a "+" sign, for example 101, 103 and 109.

[0041] To simulate the long-term inhibition process that allows units to weaken and eventually eliminate port connections, the model can allocate memory that randomly samples past media data and forces features of past media data to trend away from features of the current media data. Those media inputs can be Figure 1 The images are marked with a "-" symbol, such as 105 and 107. Their representations are indicated at 151 and 153 in the shared representation space. For those media data as input, a contrastive loss function is used to simulate the neuron wiring process and the long-term inhibition process. Therefore, the loss function can be described as follows by formula (1):

[0042]

[0043] Note that L represents the loss function, W I and W A Represent the feature encoding weights of the image channel and audio channel respectively, is the input of the ith media in the sequence, Y i is the association index (0 indicates association, 1 indicates no association), m is the margin in the shared feature space, D i is the Euclidean distance between the i-th media representation and the anchor media representation.

[0044] In the training according to the example implementation, the image channel is used as the anchor channel. However, the example implementation is not limited thereto, and audio media segments can also be used as anchor points that change over time.

[0045] The distance between the i-th representation and the anchor representation can be described by formula (2) as follows:

[0046]

[0047] Please note, and are the shared spatial feature representation of the i-th input and the shared spatial feature representation of the anchor input, respectively. They are the high-dimensional outputs of the corresponding fully connected network.

[0048] According to example implementations, instead of using a ternary loss function or a ternary network loss function, a simple summation of contrastive losses is implemented as the loss function. This simple summation of contrastive losses allows the system to process data provided from a random number of inputs, rather than requiring the data to be formed into triplets before processing. As a result, example implementations can avoid this issue for online learning. Furthermore, because the contrastive loss pushes features of the same object and co-occurring sounds toward each other, adding a small perturbation to a representation can trigger many corresponding image or sound options related to that representation.

[0049] Compact representation spaces also enable systems to learn complex tasks. For example, and not by way of limitation, because contrastive and triplet loss labelers manually organize data into pairs or triplets, they require large amounts of data for training and can be slower than traditional classifiers. Requiring humans to prepare data before training can present challenges and drawbacks. On the other hand, according to example implementations, machines can continuously receive data from sensors and automatically learn from it. Thus, the obstacles associated with manually generated labels can be avoided.

[0050] According to this example implementation, two paths are provided to process "continuous" images associated with each other, and one path is provided to process voice associated with each other. Since the amount of time associated with the utterance of a word is greater, the longer utterance time may weaken the association effect associated with the audio channel compared to the image channel.

[0051] Furthermore, speech repetition may occur less frequently than image repetition (e.g., the same object viewed from different angles). More similar images can be processed in similar frames within the same amount of time. Since adding additional similar image channels can feed these images sequentially through the same image channel, adding additional similar image paths may not be necessary according to example implementations. However, example implementations are not limited thereto, and if the model is used to learn from other samples, the learning path arrangement may be adjusted accordingly.

[0052] In addition to sharing feature spaces across modalities, one or more media generators are provided to generate media output based on features in the shared feature space. According to this example method, an image or audio input can trigger an output that may have been generated with similar input in the past. For example, but not by way of limitation, an input image of an object with a less perturbed feature space can trigger an output image of the object with a different angle and lighting conditions.

[0053] Similarly, voice input can trigger various outputs of an object. The above method according to the example implementation can be similar to a human imagination model based on voice input. In addition, similar to the object naming process, image input can also trigger voice output associated with similar images in the past.

[0054] The example implementation includes fully connected layers between the autoencoder hidden layers and the shared feature space to account for the differences between audio and image latent spaces in prior art autoencoders. More specifically, the example implementation introduces three fully connected layers between each latent space and the shared space to achieve the shared space formation goal.

[0055] Furthermore, according to an example implementation, due to the significant feature space differences described above, the example implementation may require three fully connected layers to convert shared features into audio features or image features. Since fully connected layers may increase signal generation uncertainty, a shortcut is added between the latent spaces of the encoder and decoder to bypass the fully connected layers for training stability.

[0056] According to this example implementation, in a learning network, each artificial neuron can provide a weight vector for projecting its input data into its output space. An activation function can be used to select a portion of the projection or map the projection to a limited range. If each artificial neuron is characterized as a communication channel, then once its structure is fixed, its channel capacity is also fixed.

[0057] On the other hand, before the final network is tuned end-to-end, the example implementation can train each decoder layer as the inverse process of its corresponding encoder layer. During training, the example implementation can feed training data to both input and output in a manner similar to autoencoder training.

[0058] like Figure 2 Figure 2 shows a t-SNE visualization of the audio-visual embedding space trained using CIFAR-10 images and 10-category speech clips corresponding to CIFAR-10 labels. By training the system with paired audio-visual data without providing labels or the number of categories, the model automatically forms 10 clear clusters 201-210 in the embedding space 200.

[0059] As an example test of the above example implementation, a dataset was generated based on the Columbia Object Image Library (COIL-100) [9] (a database of color images of 100 objects). For the audio clips corresponding to these objects, 100 English names were created, such as "mushroom", "cetaphil", etc. The corresponding audio clips of different objects were generated using Watson text-to-speech by varying the speech model parameters (e.g., expressions, etc.). The new dataset has 72 images and 50 audio clips for each object. In this dataset, 24 images and 10 audio clips from each object category were randomly sampled and used as test data. The remaining images and audio clips were used as training data. The pairing of these images and audio clips was based on the 100 basic object states of the signal generating machine.

[0060] The example implementation described above was trained and tested using the above data. Using a standard binary classifier from the state of the art, the system can accurately identify whether an image pair is from the same category with 92.5%. If the images are fed into a Siamese network trained using the state of the art, the system can accurately identify whether the image pair is from the same category with 96.3%. When the example implementation described herein uses both the image and audio modalities, the binary classification accuracy on this dataset is 99.2%.

[0061] The example implementations described above can also be used to pair information from different modalities with a common feature space. For example, since there are many existing PET / MRI image pairs that have been collected (e.g., by medical providers), existing PET images and PET / MRI pairing relationships can be used to provide physicians with reference PET images based on MRI images. As a result, the need for future PET scans can be reduced, and unnecessary costs, procedures, radiation hazards, etc. can be avoided for patients. Furthermore, physicians can receive decision support or make decisions based on retrieval of similar PET cases based solely on MRI images.

[0062] More specifically, the PET / MRI pairing information can be used to map identical PET / MRI image pairs to identical locations in the learned feature space. Subsequent MRI mappings in the feature space can then be used to retrieve similar cases associated with closely related PET images that have similar features to the MRI image features.

[0063] Despite being generated using different technologies, having different risk factors, advantages, and disadvantages, MRI images and PET images can be used in combination to diagnose diseases (e.g., dementia). Thus, a combined PET / MRI machine provides paired images. However, PET scans generate significant radiation exposure, which can put patients at risk. Example implementations use paired PET / MRI images to reduce the need for PET scans.

[0064] Figure 3 A learning architecture 300 according to an example implementation is shown. For example, a slice of a PET image 301 and a corresponding MRI image 303 are passed to two networks (see 305 and 307) with different weights. At 309, the difference in output features is applied to guide the weight learning of one network. For example, but not by way of limitation, two independent networks can be used to learn mapping weights for MRI and PET images, respectively. This example method can weight the learned interference between MRI and PET images.

[0065] More specifically, according to an example implementation, formula (3) provides an example loss function for a neural network:

[0066]

[0067] Where W P and W M are the weights of the two mapping networks, X P and X M are the input images from the PET and MRI modalities, m is the margin setting, and D P-M is the absolute difference between the network outputs, Y indicates X P and X M Is it a paired image? If the input is a paired image, Y is 1, if the input is not a paired image, Y is 0.

[0068] Figure 4 An example implementation 400 using PET / MRI images is shown. For example, many PET / MRI images are grayscale images. Here, inputs 401 and 403 are fed into a neural network 405. In addition, many pre-trained CNN networks have RGB channels. Therefore, the example implementation packs consecutive PET / MRI slices in RGB channels and uses a parallel architecture to process consecutive images. In addition, using a parallel architecture such as described above and shown in FIG. Figure 1 A fully connected network is used to combine the outputs of all CNNs for the feature space.

[0069] Figure 4 The architecture of can be used to map PET or MRI images into feature space. Since this structure uses the same network multiple times, where the number of slices is greater than a specified number (e.g., 3), this example implementation is flexible for different system settings.

[0070] For example, if the system has limited memory and processing power (e.g., limited GPUs), the same network can be used multiple times on three slices. The final features can be combined and calculated at different times. On the other hand, if the system has many GPUs and a lot of memory, the computation can be parallelized by replicating the same network multiple times for different GPUs.

[0071] The network 405 may be implemented as described above. Figure 1 The architecture shown and described. For example, but not limitation, pairs of PET and MRI images are provided only for training, and during application, MRI images are generated to retrieve "similar" PET images of other patients to avoid the need to take more PET images. It should be noted that the mapping networks for PET and MRI images can be different networks.

[0072] As mentioned above Figure 2 As in the case of , the example implementations herein force the same feature space to be used for both PET and MRI images. Furthermore, both networks can be adapted to form the feature space to provide flexibility in optimizing the quality of the feature space. Thus, as also discussed above with respect to Figure 2 As illustrated, co-occurring features may have more flexibility to be closer together, and non-co-occurring features may have more freedom to be separated.

[0073] Figure 5 An example process 500 is shown according to an example implementation. As described herein, the example process 500 can be performed on one or more devices.

[0074] At 501, a neural network receives input. More specifically, the neural network receives a first type of input associated with a first modality and a second type of input associated with a second modality. For example, but not limitation, the first type of input may be an image received in association with a visual modality, such as an image received from a sensor having a camera. For another example, but not limitation, the second type of input may be a sound file or graphic received in association with an audio modality, such as output received from a sensor having a microphone.

[0075] Although a camera and a microphone are disclosed as sensor structures, this example implementation is not limited thereto and may be replaced by other sensors without departing from the scope of the present invention. Furthermore, since input is received from sensors such as IoT sensors, the input received by the neural network may be received continuously over time, including but not limited to real-time information.

[0076] At 502, a CNN layer of a neural network processes the received input. More specifically, one neuron of the CNN may process one input. The CNN may have one or more layers and may perform learning and / or training based on the input and historical information. For example, but not limitation, the CNN may include one or more convolutional layers, and optionally, pooling layers. Those skilled in the art will appreciate that hidden layers may be provided depending on the complexity of the task performed by the neural network. As described in more detail below, weights are assigned to the layers of the CNN. The CNN may receive one or more inputs and apply the functions described above (e.g., formulas (1) and (2)) to generate one or more feature maps provided for a shared feature space in a fully connected layer.

[0077] More specifically, at 505, weights are learned by the neurons of the CNN. For example, the weights are learned based on the modality of the input. The input of a first modality (e.g., vision) may be trained differently from the input of a second modality (e.g., audio). For example, but not limitation, this may be done in Figure 1 Shown as elements 143 and 147.

[0078] Furthermore, at 507, a determination is made as to whether co-occurrence exists. For example, but not limitation, as described above, when the timing of the sensor input associated with the word appearing on the packaging co-occurs with the word being pronounced via audio, this is determined to be a co-occurrence. For each input across each modality, a determination is made as to whether co-occurrence exists. The result of this determination, along with the encoding weights based on the modality, is applied to a loss function (e.g., (1) and (2)) to determine a loss associated with the input.

[0079] At 509, a shared feature space is generated as the output of the CNN layer. As described above, a loss function and weighting can be used to simulate the neuron wiring process and the long-term inhibition process, so that some units weaken their connections and have larger distances, while other units strengthen their connections and have shorter distances, or share a common position in the feature space.

[0080] For example, but not limitation, Figure 1 As shown, the outputs of the CNN layers are shown as reference numerals 121-133, and the feature space is represented as a plane at 157. In addition, for example, units with weaker connections due to a lack of co-occurrence are shown at 151 and 153, while units with stronger connections due to the presence of co-occurrence are shown at 155.

[0081] At 511, an output is provided. For example, but not limitation, the output may be a classification of the input or an indication of a probability class that provides the best-fit classification of the input. Those skilled in the art will appreciate that training may occur (e.g., via backpropagation).

[0082] Figure 6An example computing environment 600 is shown with an example computer device 605 suitable for use with some example implementations. The computing device 605 in the computing environment 600 may include one or more processing units, cores, or processors 610, memory 615 (e.g., RAM, ROM, etc.), internal storage 620 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interfaces 625, any of which may be coupled on a communication mechanism or bus 630 for communicating information or embedded in the computing device 605.

[0083] According to this example implementation, processing associated with neural activity may occur on processor 610, which is a central processing unit (CPU). Alternatively, other processors may be used without departing from the present invention. For example, but not by way of limitation, a graphics processing unit (GPU) and / or a neural processing unit (NPU) may be used in place of or in combination with the CPU to perform the processing of the above example implementation.

[0084] The computing device 605 may be communicatively coupled to an input / interface 635 and an output device / interface 640. Either or both of the input / interface 635 and the output device / interface 640 may be wired or wireless interfaces and may be detachable. The input / interface 635 may include any device, component, sensor, or interface (physical or virtual) that can be used to provide input (e.g., buttons, a touch screen interface, a keyboard, a pointing / cursor control, a microphone, a camera, Braille, a motion sensor, an optical reader, etc.).

[0085] Output devices / interfaces 640 may include displays, televisions, monitors, printers, speakers, Braille, etc. In some example implementations, input / interface 635 (e.g., a user interface) and output devices / interface 640 may be embedded in or physically coupled to computing device 605. In other example implementations, other computing devices may serve as or provide functionality for input / interface 635 and output devices / interface 640 of computing device 605.

[0086] Examples of computing devices 605 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices in vehicles and other machines, devices carried by people and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, server devices, other computers, kiosks, televisions embedded with and / or coupled to one or more processors, radios, etc.).

[0087] Computing device 605 can be communicatively coupled to external storage device 645 and network 650 (e.g., via I / O interface 625) for communicating with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. Computing device 605 or any connected computing device can function as, provide services for, or be referred to as a server, client, thin server, general-purpose machine, special-purpose machine, or another label. For example, and not by way of limitation, network 650 can include a blockchain network and / or a cloud.

[0088] I / O interface 625 may include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11xs, universal system bus, WiMAX, modems, cellular network protocols, etc.) for communicating information to and / or from at least all connected components, devices, and networks in computing environment 600. Network 650 may be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, etc.).

[0089] The computing device 605 may use and / or communicate with computer-usable or computer-readable media (including transitory and non-transitory media). Transitory media include transmission media (e.g., metal cables, optical fibers), signals, carrier waves, etc. Non-transitory media include magnetic media (e.g., magnetic disks and tapes), optical media (e.g., CD ROMs, digital video disks, Blu-ray discs), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage devices), and other non-volatile storage devices or memories.

[0090] The computing device 605 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some example computing environments. The computer-executable instructions can be retrieved from a transient medium, as well as stored on and retrieved from a non-transitory medium. The executable instructions can originate from one or more of any programming, scripting, and machine languages (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0091] The processor 610 can execute under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, including a logic unit 655, an application programming interface (API) unit 660, an input unit 665, an output unit 670, a sensor input processing unit 675, a machine learning unit 680, an output determination unit 685, and an inter-unit communication mechanism 695 for different units to communicate with each other, with the OS, and with other applications (not shown).

[0092] For example, the sensor input processing unit 675, the machine learning unit 680, and the output determination unit 685 may implement one or more of the processes described above for the above structures. The design, function, configuration, or implementation of the described units and elements may vary and are not limited to the description provided.

[0093] In some example implementations, when information is received or instructions are executed via the API unit 660, they may be communicated to one or more other units (e.g., the logic unit 655, the input unit 665, the sensor input processing unit 675, the machine learning unit 680, and the output determination unit 685).

[0094] For example, as described above, the sensor input processing unit 675 can receive and process information from one or more sensors. The output of the sensor input processing unit 675 is provided to the machine learning unit 680, which, for example, can process the information based on the data described above and shown in FIG. Figure 1 In addition, the output determination unit 685 can provide an output signal based on the output of the sensor input processing unit 675 and the machine learning unit 680.

[0095] In some cases, in some of the example implementations described above, the logic unit 655 can be configured to control the flow of information between the units and direct the services provided by the API unit 660, the input unit 665, the sensor input processing unit 675, the machine learning unit 680, and the output determination unit 685. For example, the flow of one or more processes or implementations can be controlled by the logic unit 655 alone or in conjunction with the API unit 660.

[0096] Figure 7 An example environment suitable for some example implementations is shown. Environment 700 includes devices 705-745, and each device is communicatively connected to at least one other device via, for example, a network 760 (e.g., via a wired and / or wireless connection). Some devices may be communicatively connected to one or more storage devices 730 and 745.

[0097] Examples of one or more devices 705-745 may be Figure 6 7. Devices 705-745 may include, but are not limited to, a computer 705 (e.g., a laptop computing device) with a monitor and associated webcam as described above, a mobile device 710 (e.g., a smartphone or tablet), a television 715, a device associated with a vehicle 720, a server computer 725, computing devices 735-740, and storage devices 730 and 745.

[0098] In some implementations, devices 705-720 can be considered user devices associated with users who can remotely obtain sensed input used as input for the example implementations described above. In this example implementation, as described above, one or more of these user devices can be associated with one or more sensors (e.g., a camera and / or a microphone) that can sense information required for this example implementation.

[0099] Thus, the present example implementations may have various benefits and advantages. For example, but not limitation, the example implementations involve learning directly from sensors, and thus may take advantage of the widespread presence of visual sensors such as cameras and audio sensors such as microphones.

[0100] Furthermore, example implementations do not require textual conversion from pre-existing sensor data to facilitate machine learning. Instead, example implementations learn sensor data relationships without the need for text. For example, example implementations do not require conversion of image grayscale values to audio or audio to text to facilitate other conversions from one modality or another to a common medium.

[0101] Instead of converting across modalities, this example implementation receives input from different modalities (e.g., images and audio) and processes the data across modalities with asymmetric dimensions and structures. Paired information between modalities can be generated (e.g., during image-audio training), similar to learning with the eyes and ears, and understanding corresponding image / audio pairs to correctly understand the connection. However, this example implementation does not require any manual human activity regarding this process.

[0102] To achieve the above, the example implementation can use a CNN-based autoencoder and a shared space across modalities. Thus, the example implementation can handle generation in both directions (e.g., audio to image and image to audio) with respect to modalities. In addition, the example implementation can generate audio spectrograms, which can result in a significantly smaller model size than using raw audio, for example.

[0103] By employing a neural network approach, the example implementation can learn to provide nonlinear interpolation in signal space. Compared to the prior art lookup table approach, the example implementation employs a neural network to provide a compact form for generating signals and provides substantial efficiency with respect to memory space allocation. The example implementation is consistent with neuroscience principles such as "wire together, fire together" and the long-term inhibition process described above. Furthermore, rather than using a classroom model that pairs one example and one modality with multiple examples in another modality, paired data can be fed into the example implementation architecture in a random manner, such as one-to-one pairings from a random order, in a manner similar to the human learning process, but performed in an automated manner.

[0104] While the example implementations described above are provided with respect to imaging technology for medical diagnosis, the example implementations are not limited thereto, and those skilled in the art will appreciate that other approaches may be employed. For example, and not by way of limitation, the example implementations may be employed in the architecture of systems for supporting people with disabilities, for autonomous robot training, for machine learning algorithms and systems requiring large amounts of low-cost training data, and for machine learning systems that are not constrained by the availability of manual text labelers.

[0105] Additionally, example implementations may involve language-independent devices that can be trained to display objects that others can physically hear to deaf people. Because this example implementation does not employ text itself, the training system can be language-independent. Furthermore, because the devices associated with the architecture are communicatively linked or connected to a network, people living in the same area and speaking the same language can train the system together.

[0106] Furthermore, while this example implementation involves vision and audio, other modalities may be added or replaced without departing from the present invention. For example, and not by way of limitation, a machine may include temperature or touch, and the inclusion of a new modality will not affect previously learned modalities. Instead, the new modality will learn on its own and gradually build more connections with the previous modalities.

[0107] Although some example implementations have been shown and described, these example implementations are provided in order to convey the subject matter described herein to those familiar with the art. It should be understood that the subject matter described herein can be implemented in various forms and is not limited to the example implementations described. The subject matter described herein can be practiced without those specifically defined or described matters or with other or different elements or matters that are not described. Those familiar with the art will understand that these exemplary implementations can be changed without departing from the subject matter defined in the appended claims and their equivalents as described herein.

[0108] The various aspects of certain non-limiting embodiments of the present disclosure address the features discussed above and / or other features not described above. However, the various aspects of the non-limiting embodiments need not address the features described above, and the various aspects of the non-limiting embodiments of the present disclosure may not address the features described above.

Claims

1. A computer-implemented method comprising the following steps: receiving a first input associated with a first modality and a second input associated with a second modality; Processing the received first input and second input in a convolutional neural network (CNN), wherein a first set of weights is assigned to the first input and a second set of weights is assigned to the second input; determining a loss for each of the first input and the second input based on a loss function applying the first set of weights, the second set of weights, and the presence of co-occurrences; generating a shared feature space as an output of the CNN, wherein distances between cells associated with the first input and the second input in the shared feature space are determined based on the loss associated with each of the first input and the second input; and providing an output indicating a classification or a probability of a classification based on the shared feature space, The first input and the second input are respectively an image capture file received from a camera associated with the first input and an audio file received from a microphone associated with the second input.

2. The computer-implemented method of claim 1 , wherein: A first anchor channel is associated with the first modality, and a second anchor channel is associated with the second modality.

3. The computer-implemented method of claim 1 , wherein: The first modality includes a visual mode and the second modality includes an audio mode.

4. The computer-implemented method of claim 1 , wherein: The co-occurrence is associated with the first input and the second input in sequence within a common time window.

5. The computer-implemented method of claim 1 , wherein: Text markup is not performed, the first input is not converted to the second modality, and the second input is not converted to the first modality.

6. The computer-implemented method of claim 1 , wherein: The computer-implemented method is executed in a neural processing unit of a processor.

7. The computer-implemented method of claim 1 , wherein: The computer-implemented method is executed in a processor of a mobile communication device, a home management device, and / or a robotic device.

8. A non-transitory computer-readable medium having a storage device storing instructions, the instructions being executable by a processor, the instructions comprising: receiving a first input associated with a first modality and a second input associated with a second modality; Processing the received first input and second input in a convolutional neural network (CNN), wherein a first set of weights is assigned to the first input and a second set of weights is assigned to the second input; determining a loss for each of the first input and the second input based on a loss function applying the first set of weights, the second set of weights, and the presence of co-occurrences; generating a shared feature space as an output of the CNN, wherein distances between cells associated with the first input and the second input in the shared feature space are determined based on the loss associated with each of the first input and the second input; and Based on the shared feature space, providing an output indicating whether there is co-occurrence, The first input and the second input are respectively an image capture file received from a camera associated with the first input and an audio file received from a microphone associated with the second input.

9. The non-transitory computer-readable medium of claim 8, wherein: The first modality includes a visual mode and the second modality includes an audio mode.

10. The non-transitory computer-readable medium of claim 8, wherein: The co-occurrence is associated with the first input and the second input in sequence within a common time window.

11. The non-transitory computer-readable medium of claim 8, wherein: Text markup is not performed, the first input is not converted to the second modality, and the second input is not converted to the first modality.

12. The non-transitory computer-readable medium of claim 8, wherein: The instructions are executed in a processor of a mobile communication device, a home management device, and / or a robotic device.

Citation Information

Patent Citations

  • Cross-modal retrieval method based on sketch retrieval three-dimensional model

    CN110188228A

  • Multi-task multi-modal machine learning model

    CN110574049A