Purified contrast learning for lightweight neural network training
By introducing loss functions of one-way attraction and pure negative repulsion in the contrast learning framework, clean samples and enhanced samples are generated, and the problems of low efficiency and reduced accuracy in lightweight neural network training are solved, and more efficient and accurate model training is achieved.
Patent Information
- Application Number
- CN202380073707.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-25
- Filing Date
- 2023-08-28
- Publication Date
- 2025-05-30
AI Technical Summary
The existing contrast learning methods are inefficient when training lightweight neural networks and rely on complex pre-tasks, resulting in reduced accuracy of the trained model.
A purified contrast learning framework is proposed to optimize the model to align clean and enhance the embeddings of clean and enhanced samples in a single embedding space by generating clean and enhanced samples and based on loss functions of one-way attraction and pure negative repulsion.
This method reduces the complexity of the contrast learning framework, makes it suitable for training of lightweight models, and improves the accuracy of lightweight models, even under limited labeled data conditions.
Smart Images

Figure CN120077383A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Patent Application No. 18 / 456,112, filed on August 25, 2023, entitled "PURIFIED CONTRASTIVE LEARNING FOR LIGHTWEIGHT NEURAL NETWORK TRAINING", which claims the benefit of U.S. Provisional Patent Application No. 63 / 419,272, filed on October 25, 2022, entitled "PURIFIED CONTRASTIVE LEARNING FOR LIGHTWEIGHT NEURAL NETWORK TRAINING", the disclosures of which are hereby incorporated by reference in their entireties. Technical Field
[0003] Aspects of the present disclosure generally relate to training artificial neural networks via purified contrastive learning. Background Art
[0004] An artificial neural network may include interconnected groups of artificial neurons (e.g., neuron models). An artificial neural network may be a computing device or represented as a method to be executed by a computing device. Some artificial neural networks can be trained in a supervised manner based on labeled data, allowing for the development of specialized models that excel in their designated tasks. However, in practice, it is impractical to label every possible element in the world. Additionally, certain tasks (such as training a speech recognition system in an ancient dialect) face the problem of scarce labeled data. Therefore, the reliance on supervised learning may hinder the development of more intelligent all - around models that can perform multiple tasks and / or acquire new skills. Thus, some artificial neural networks are trained on unlabeled data in a self - supervised manner.
[0005] Contrastive learning is an example of a framework for self - supervised learning used in various tasks. The goal of contrastive learning is to train an artificial neural network to learn a representation of the data without relying on explicit labels. The representation can be learned by contrasting positive and negative pairs of examples. During training, the artificial neural network learns to map similar augmented samples closer together in the feature space while separating dissimilar samples farther apart. This process encourages the artificial neural network to capture meaningful and discriminative representations of the data. Summary of the Invention
[0006] In some aspects of the present disclosure, a method includes: generating a clean sample and an augmented sample for each input in a set of inputs. The method further includes: associating the clean sample with the augmented sample for each input in the set of inputs to form a positive pair. The method further includes: associating the clean sample with another clean sample associated with another input among the multiple inputs for each input in the set of inputs to form a negative pair. The method further includes: learning one or more representations of the set of inputs based on the positive pairs and negative pairs for each input in the set of inputs.
[0007] Some aspects of the present disclosure relate to an apparatus that includes components for generating a clean sample and an augmented sample for each input in a set of inputs. The apparatus further includes components for associating the clean sample with the augmented sample for each input in the set of inputs to form a positive pair. The apparatus further includes components for associating the clean sample with another clean sample associated with another input among the multiple inputs for each input in the set of inputs to form a negative pair. The apparatus further includes components for learning one or more representations of the set of inputs based on the positive pairs and negative pairs for each input in the set of inputs.
[0008] In some aspects of the present disclosure, a non-transitory computer-readable medium having program code recorded thereon is disclosed. The program code is executed by a processor and includes program code for generating a clean sample and an augmented sample for each input in a set of inputs. The program code further includes program code for associating the clean sample with the augmented sample for each input in the set of inputs to form a positive pair. The program code further includes program code for associating the clean sample with another clean sample associated with another input among the multiple inputs for each input in the set of inputs to form a negative pair. The program code further includes program code for learning one or more representations of the set of inputs based on the positive pairs and negative pairs for each input in the set of inputs.
[0009] Some aspects of the present disclosure relate to an apparatus that has one or more processors and one or more memories coupled to the one or more processors and storing instructions that are operable, when executed by the one or more processors, to cause the apparatus to generate a clean sample and an augmented sample for each input in a set of inputs. Execution of the instructions further causes the apparatus to associate the clean sample with the augmented sample for each input in the set of inputs to form a positive pair. Execution of the instructions further causes the apparatus to associate the clean sample with another clean sample associated with another input among the multiple inputs for each input in the set of inputs to form a negative pair. Execution of the instructions further causes the apparatus to learn one or more representations of the set of inputs based on the positive pairs and negative pairs for each input in the set of inputs.
[0010] Overall, aspects include, for example, methods, apparatuses, systems, computer program products, non-transitory computer-readable media, user equipment, base stations, wireless communication devices, and processing systems substantially as described with reference to the figures and as illustrated in the figures and the specification.
[0011] The features and technical advantages of examples in accordance with the present disclosure have been outlined rather broadly above so that the detailed description below may be better understood. Additional features and advantages will be described. The disclosed concepts and specific examples may be readily used as a basis for modifying or designing other structures for achieving the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the disclosed concepts, both as to their organization and method of operation, as well as the associated advantages, will be better understood by considering the following description in conjunction with the accompanying figures. Each of the figures provided is for the purpose of illustration and description and not as a definition of the limits of the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The features, nature, and advantages of the present disclosure will become more apparent upon consideration of the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numerals refer to like elements throughout.
[0013] Figure 1 Illustrative embodiments of neural networks using a system-on-chip (SOC), including a general-purpose processor, in accordance with certain aspects of the present disclosure.
[0014] Figure 2A 、 Figure 2B and Figure 2C are diagrams illustrating neural networks in accordance with aspects of the present disclosure.
[0015] Figure 3 is a block diagram illustrating an example of a conventional contrastive learning framework.
[0016] Figure 4 is a block diagram illustrating an example of a contrastive learning framework in accordance with certain aspects of the present disclosure.
[0017] Figure 5 is a flowchart illustrating an example of a process for contrastive learning in accordance with aspects of the present disclosure. DETAILED DESCRIPTION
[0018] The following detailed description, taken in conjunction with the accompanying drawings, is intended as a description of various configurations and is not intended to represent the only configurations in which the described concepts may be practiced. To provide a thorough understanding of the various concepts, the detailed description includes specific details. It will be apparent, however, to one of ordinary skill in the art that the concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0019] Based on the teachings, those skilled in the art should recognize that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether that aspect is implemented independently of any other aspect of the present disclosure or in combination with any other aspect. For example, a device may be implemented or a method may be practiced using any number of the aspects set forth. In addition, the scope of the present disclosure is intended to cover such devices or methods practiced using other structures, functionalities, or a combination of structures and functionalities that complement or are different from the various aspects of the present disclosure set forth. It should be understood that any aspect of the present disclosure disclosed may be embodied by one or more elements of the claims.
[0020] The term "exemplary" is used to mean "serving as an example, instance, or illustration". Any aspect described as "exemplary" need not be construed as superior or better than other aspects.
[0021] Although specific aspects have been described, numerous variations and permutations of these aspects fall within the scope of the present disclosure. Although some benefits and advantages of the preferred aspects have been mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses, or purposes. Instead, the aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the accompanying drawings and the following description of the preferred aspects. The detailed description and the drawings are merely illustrative of the present disclosure and not limiting, and the scope of the present disclosure is defined by the appended claims and their equivalents.
[0022] As discussed, some artificial neural networks can be trained in a supervised manner based on labeled data, thereby allowing the development of specialized models that excel in their designated tasks. Nevertheless, the reliance on supervised learning may impede the development of more intelligent all-round models that can perform multiple tasks and / or acquire new skills. Therefore, if artificial neural networks can be trained in an unsupervised manner with respect to unlabeled data, the ability of the artificial neural networks to perform tasks can be improved.
[0023] Self-supervised learning is a training method that allows artificial neural networks to learn from unlabeled data by deriving supervision signals from the data itself. Contrastive learning is an example of a framework for self-supervised learning used in various tasks. The goal of contrastive learning is to train an artificial neural network to learn a representation of the data without relying on explicit labels. During training, the artificial neural network learns to map positive pairs (e.g., matching pairs) together while pushing negative pairs (e.g., non-matching pairs) apart, resulting in a discriminative embedding space. For example, in image recognition, contrastive learning can be used to train an image recognition model to identify the original image in an augmented version by maximizing the similarity between positive pairs (the original image and its augmentation) while minimizing the similarity to negative pairs (images from different classes or unrelated images). This process allows the image recognition model to learn meaningful representations without relying on explicit labels. Some contrastive learning methods have successfully achieved performance levels similar to supervised training of large models. Nevertheless, for lightweight models, contrastive learning methods may not achieve the same or similar performance as supervised training.
[0024] Aspects of the present disclosure relate to a purified contrastive learning framework that can be used to train lightweight models. In some examples, positive and negative pairs can be redefined to remove unnecessary information while retaining information that is discriminative from the unlabeled data.
[0025] Figure 1 An example embodiment of a system-on-chip (SOC) 100 is illustrated, which may include a central processing unit (CPU) 102 or a multi-core CPU configured to train an artificial neural network via purified contrastive learning. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., a neural network with weights), latencies, frequency slot information, and task information may be stored in a storage block associated with a neural processing unit (NPU) 108, a storage block associated with the CPU 102, a storage block associated with a graphics processing unit (GPU) 104, a storage block associated with a digital signal processor (DSP) 106, storage block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from storage block 118.
[0026] The SOC 100 may also include additional processing blocks customized for specific functions, such as the GPU 104, the DSP 106, the connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and the multimedia processor 112 that can detect and recognize poses, for example. In a specific implementation, the NPU 108 is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120, which may include a global positioning system.
[0027] The SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the general-purpose processor 102 may include code for performing the following operations: generating clean samples and enhanced samples for each of a plurality of inputs; associating the clean samples with the enhanced samples for each of the plurality of inputs to form positive pairs; associating the clean samples with another clean sample of another one of the plurality of inputs for each of the plurality of inputs to form negative pairs; and learning one or more representations of the plurality of inputs based on the positive and negative pairs for each of the plurality of inputs.
[0028] Object recognition is an example of a task performed by an artificial neural network. In some examples, the object recognition task learns to represent the input at successively higher levels of abstraction in each layer, thereby constructing a useful feature representation of the input data.
[0029] Deep learning architectures can learn hierarchies of features. For example, if presented with visual data, the first layer may learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer may learn to recognize spectral power in specific frequencies. The second layer takes the output of the first layer as input and may learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers may learn to represent complex shapes in visual data or words in auditory data. Even higher layers may learn to recognize common visual objects or spoken phrases.
[0030] Deep learning architectures can perform particularly well when applied to problems with a natural hierarchy. For example, the classification of motorized vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher layers to recognize cars, trucks, and airplanes.
[0031] Neural networks can be designed to have various connection patterns. In a feedforward network, information passes from lower layers to higher layers, where each neuron in a given layer communicates with neurons in the higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one block of input data presented sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the identification of high-level concepts can assist in discerning specific low-level features of the input.
[0032] The connections between the layers of a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, the neurons in the first layer can communicate their outputs to each neuron in the second layer, such that each neuron in the second layer will receive inputs from every neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, the neurons in the first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layer of the locally connected neural network 204 can be configured such that each neuron in the layer will have the same or a similar connection pattern, but the connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connection pattern can result in spatially distinct receptive fields in higher layers, because the higher layer neurons in a given region can receive inputs that are tuned through training to the characteristics of a restricted portion of the total input to the network.
[0033] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 can be configured such that the connection strengths associated with the inputs to each neuron in the second layer are shared (e.g., 208). Convolutional neural networks can be well-suited for problems where the spatial location of the input is meaningful.
[0034] A Deep Belief Network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract hierarchical representations of a training data set. A DBN can be obtained by stacking layers of Restricted Boltzmann Machines (RBMs). An RBM is a type of artificial neural network that can learn a probability distribution from a set of inputs. Since an RBM can learn a probability distribution without information about the class to which each input should be classified, an RBM is typically used for unsupervised learning. Using a hybrid paradigm of supervised and unsupervised, the bottom RBM of a DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the inputs from the previous layer and the target classes) and can be used as a classifier.
[0035] A Deep Convolutional Network (DCN) is a network of convolutional networks configured with additional pooling and normalization layers. A DCN has achieved state-of-the-art performance on many tasks. A DCN can be trained using supervised learning, where both input targets and output targets are known for many paradigms and are used to modify the weights of the network by using gradient descent methods.
[0036] A DCN can be a feedforward network. Additionally, as described above, the connections from the neurons in the first layer of a DCN to a set of neurons in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. For example, the computational burden of a DCN may be much smaller than that of a similarly sized neural network that includes recurrent or feedback connections.
[0037] As discussed, some artificial neural networks can be trained in a supervised manner based on labeled data, allowing for the development of specialized models that excel in their designated tasks. Nevertheless, the reliance on supervised learning may impede the development of more intelligent all-purpose models that can perform multiple tasks and / or acquire new skills. Therefore, if an artificial neural network can be trained in a self-supervised manner on unlabeled data, the ability of the artificial neural network to perform tasks can be improved.
[0038] Self-supervised learning is a training method that allows an artificial neural network to learn from unlabeled data by deriving a supervision signal from the data itself. Contrastive learning is an example of a framework for self-supervised learning used in various tasks. Contrastive learning is a technique in machine learning (such as artificial neural networks) that aims to learn meaningful representations by contrasting positive and negative pairs of examples. The purpose of contrastive learning is to bring similar examples closer together in the learned representation space while pushing dissimilar examples apart, resulting in a discriminative embedding space.
[0039] In conventional contrastive learning, a corresponding set of augmented samples is generated from each input sample in a group of input samples. These augmented samples share the same identity as the original input sample but are modified in some way. Positive pairs are formed by matching augmented samples from the same input sample, while negative pairs are composed of augmented samples from different input samples.
[0040] In most cases, a contrastive loss function quantifies the similarity between example pairs in the representation space such that the similarity between positive pairs is maximized or increased, while the similarity between negative pairs is minimized or decreased. During training, the model can be optimized to learn representations that effectively distinguish between positive and negative pairs. By doing so, the model learns to capture meaningful and informative features useful for downstream tasks such as (but not limited to) classification, retrieval, or clustering.
[0041] Contrastive learning has received significant attention and success, especially in the field of computer vision, where contrastive learning has been successfully applied to image recognition, object detection, image retrieval tasks, etc. By leveraging the structure and relationships within the data, contrastive learning enables the model to capture relevant patterns and representations without explicit labels, making contrastive learning a valuable tool for self-supervised learning and unsupervised representation learning.
[0042] For example, in image recognition, contrastive learning can be used to train an image recognition model to identify the original image in the augmented version by maximizing the similarity between positive pairs (the original image and its augmentation) while minimizing the similarity to negative pairs (images from different classes or irrelevant images). This process allows the image recognition model to learn meaningful representations without relying on explicit labels. Some contrastive learning methods have successfully achieved performance levels similar to those of supervised training for large models. Nevertheless, for lightweight models, contrastive learning methods do not achieve the same or similar performance as supervised training.
[0043] For example, self-supervised learning has demonstrated significant progress in the audio and speech domains. Some speech-related tasks often implement lightweight models that can operate on low-power devices such as edge devices. For example, applications such as keyword spotting for voice assistants can use lightweight models with an always-on behavior. These keyword spotting applications for voice assistants can operate on low-power edge devices.
[0044] An edge device is an example of a computing device located close to the data generation source. Edge devices can perform processing tasks and make real-time decisions at or near the location where data is being generated, rather than relying on sending data to a centralized server or cloud for processing. Edge devices are typically small, lightweight, and energy-efficient devices that can be deployed in various environments, including (but not limited to) Internet of Things (IoT) devices, smartphones, smart sensors, wearable devices, routers, embedded systems, hub systems, etc. In some cases, the processing resources and / or memory resources of edge devices may be less than those of non-edge devices.
[0045] Edge devices can be useful in scenarios where network bandwidth is limited or there are concerns about data privacy, network latency, or intermittent connectivity. Examples of edge computing applications include smart homes, industrial automation, autonomous vehicles, remote monitoring, healthcare monitoring devices, and smart cities. Edge devices can utilize machine learning models to achieve local decision-making and intelligence at the edge. This allows for real-time processing, predictive analytics, anomaly detection, and localized intelligence without relying heavily on cloud-based services.
[0046] When used to train lightweight models (e.g., small models), conventional contrastive learning frameworks are generally inefficient. These conventional contrastive learning frameworks rely heavily on complex pretext tasks to train lightweight models. This training results in lightweight models with reduced accuracy.
[0047] Figure 3 is a block diagram illustrating an example of a conventional contrastive learning framework 300. As Figure 3 shown in the example of, in the conventional contrastive learning framework 300, an augmentation module 302 can generate a pair of augmented samples from each input x 1 and x 2 . This augmentation can be an example of a pretext task. The augmentation module 302 can be a random augmentation module that randomly transforms a given input to generate two related views of the same input. Specifically, in Figure 3 the example of, for the pretext task, N examples in a mini-batch are transformed via the augmentation module 302 to form multiple pairs of augmented samples, resulting in 2N data points Each pair of augmented examples can be considered a positive pair. Augmentation can include, for example, cropping, resizing, distorting, and / or applying random Gaussian blur to each input x 1 and x 2 . For N examples in a mini-batch, i ∈ I = {1...2N} represents the subscripts of different pairs of augmented samples. The variable j represents the subscript of another augmented sample from the same source x i .
[0048] The sample output of the augmentation module 302 can be encoded via an encoder 304 associated with an encoding function f(·) that extracts a representation vector from the augmented data example. Each representation vector can be processed by a projection head 306 associated with a projection function g(·) that maps the corresponding representation to the space where the contrastive loss is applied. That is, the projection head 306 can transform the encoded data into a different feature space or dimension, typically with a reduced dimension, while preserving the relevant information. As Figure 3 shown in the example of 1 and x 2 a corresponding pair of augmented samples can be generated and Each augmented sample in a pair of augmented samples is a positive match for the corresponding augmented sample in that pair of augmented samples (as shown by the solid arrow). For example, the first augmented sample is a positive match for the second augmented sample within the pair of augmented samples associated with the first input x 1 In addition, each augmented sample within the pair of augmented samples is a negative match for another augmented sample within another pair of augmented samples (as shown by the dashed arrow). For example, the first augmented sample is a negative match for both the third augmented sample associated with the second input x 2 and the fourth augmented sample and vice versa.
[0049] The loss function for the conventional contrastive learning framework 300 is as follows:
[0050]
[0051] In Equation 1, denotes the indicator function, and is the latent representation of the augmented sample extracted from the encoder 304 (f(·)) and the projection head 306 (g(·)). The similarity function (sim(u,v)sim(u,v) = u T v / ∥u∥∥v∥) represents the l 2 normalized dot product between u and v (e.g., cosine similarity). The subscript i denotes the anchor, and the subscript j denotes the positive associated with the anchor. The other 2(N - 1) subscripts denote the negative samples with respect to the anchor and the positive. For example, the first augmented sample can be the anchor, and the second augmented sample can be the positive. The third augmented sample and the fourth augmented sample can be the negative samples with respect to the anchor and the positive. AsFigure 3 As shown in the example of, all positive and negative pairs are augmented samples. In Equation 1, τ represents the temperature parameter.
[0052] In the Figure 3 The pretext tasks used in the conventional contrastive learning framework 300 described increase the complexity of the conventional contrastive learning framework 300. Therefore, when training a lightweight model, the described pretext tasks (e.g., generating two augmented samples for each input) are not suitable. Aspects of the present disclosure relate to a contrastive learning framework suitable for lightweight models.
[0053] Figure 4 is a block diagram illustrating an example of a contrastive learning framework 400 according to aspects of the present disclosure. In Figure 4 the example of, the contrastive learning framework 400 incorporates unidirectional attraction (UDA) to align a single embedding space by increasing the similarity between the augmented samples and the clean samples associated with the input. The embedding space refers to a mathematical representation or feature space in which data points or entities are transformed into a lower-dimensional vector representation. In the context of machine learning, the embedding space can capture and represent the latent or underlying characteristics of the data.
[0054] As Figure 4 shown in the example of, for each input x 1 and x 2 , the contrastive learning framework 400 generates a clean sample z i and an augmented sample The clean sample z i refers to a sample generated by encoding the input via an encoder 304 associated with an encoding function f(·). The representation vector generated by the encoder 304 can be processed by a projection head 306 associated with a projection function g(·). As Figure 4 shown in, for the first input x 1 , the contrastive learning framework 400 can generate a first clean sample z 1 . For the second input x 2 , the contrastive learning framework 400 can generate a second clean sample z 2 .
[0055] The augmented sample refers to a sample generated by augmenting the input via an augmentation module 302. The output of the augmentation module 302 is processed by the encoder 304 and the projection head 306. As Figure 4 shown in, for the first input x 1 , the contrastive learning framework 400 can generate a first augmented sample For the second input x 2 , the contrastive learning framework 400 can generate a second augmented sample
[0056] In some examples, the embedding associated with the clean sample z i can be aligned or substantially aligned with the embedding associated with the augmented sample . The clean sample is regarded as the ground truth. In such examples, the similarity represented between the embedding of the clean sample and the augmented sample can be increased or maximized to align the embeddings. The complexity of the contrastive learning framework 400 can be reduced by considering unidirectional attraction, such that the contrastive learning framework 400 can be used to train lightweight models.
[0057] As Figure 3 shown, the conventional contrastive learning framework 300 uses augmented samples to generate one or more negative pairs. The residual between the augmented sample and the clean sample z i is Therefore, the similarity between these two pairs is defined as follows S neg = sim(z i - ∈ i , z k - ∈ k ).
[0058] In some examples, the discrimination between instances can be performed by reducing the similarity between these two pairs. The discriminative information depends on both the latent embedding z and the residual ∈. If the positive pair increases the robustness against data augmentation, then this residual ∈ should not be used as discriminative information. Therefore, in the example of Figure 4 , the contrastive learning framework 400 uses pure negative rejection that generates negative pairs using only clean samples (shown as the dashed line in the example of Figure 4 ). Additionally, the augmented sample is a positive match with the corresponding clean sample z i (shown as the solid line).
[0059] In the example of Figure 4 , the loss for the contrastive learning framework 400 is based on unidirectional attraction and pure negative rejection. The loss can be determined as follows:
[0060]
[0061] The objective of this loss function is to align the embeddings of the clean sample z i and the augmented sample . In this equation, represents the similarity between the stop-gradient operation applied to the clean sample z i and the augmented sample . SG(·) represents the stop-gradient operation for the anchor of the positive pair. The stop-gradient operation prevents the gradient from propagating through the anchor of the positive pair. The loss function in Equation 2 consists of two components. The numerator term Denote the clean sample as z i The stop-gradient operation with the augmented sample and the similarity, which is exponentiated with the temperature parameter τ and normalized, between the clean sample z i and the augmented sample The numerator term promotes the alignment of the embeddings of the clean sample z
[0062] The denominator term sums the similarities, which are exponentiated with the temperature parameter τ and normalized, between the augmented sample z i and all negative samples z in a mini-batch. The denominator term represents the dissimilarity between the augmented sample and the negative samples. As discussed, the negative sample z k is a clean sample associated with another input. k is a clean sample associated with another input.
[0063] By minimizing the loss function in Equation 2, the contrast learning framework 400 described in Figure 4 maximizes the similarity between the stop-gradient operation and the augmented sample while minimizing the similarity between the augmented sample and the negative samples. This promotes the alignment of the embeddings and helps to learn meaningful representations.
[0064] Refer to Figure 4 The contrast learning framework 400 described in
[0065] Figure 5 is a flowchart illustrating an example of a process 500 for contrast learning (CL) according to aspects of the present disclosure. As Figure 5 shown, the process 500 begins at block 502 by generating, for each input in a set of inputs, a clean sample and an augmented sample. At block 504, the process 500 associates the clean sample with the augmented sample for each input in the set of inputs to form positive pairs. At block 506, the process 500 associates the clean sample with another clean sample associated with another input in the plurality of inputs for each input in the set of inputs to form negative pairs. At block 508, the process 500 learns one or more representations of the set of inputs based on the positive and negative pairs for each input in the set of inputs.
[0066] Specific implementation examples are described in the following numbered clauses:
[0067] Clause 1. A processor-implemented method, comprising: generating a clean sample and an augmented sample for each of a plurality of inputs; associating the clean sample with the augmented sample for each of the plurality of inputs to form a positive pair; associating the clean sample with another clean sample of another one of the plurality of inputs for each of the plurality of inputs to form a negative pair; and learning one or more representations of the plurality of inputs based on the positive pairs and the negative pairs for each of the plurality of inputs.
[0068] Clause 2. The processor-implemented method according to any one of Clause 1, wherein: learning the one or more representations includes minimizing a loss for each of the plurality of inputs; the clean sample is the ground truth; and the stop gradient is a function of the embedding of the clean sample.
[0069] Clause 3. The processor-implemented method according to any one of Clauses 1 to 2, further comprising learning the one or more representations in a self-supervised manner via contrastive learning.
[0070] Clause 4. The processor-implemented method according to any one of Clauses 1 to 3, wherein each input is an audio input.
[0071] Clause 5. The processor-implemented method according to Clauses 1 to 4, further comprising receiving each input at a contrastive learning model.
[0072] Clause 6. The processor-implemented method according to Clause 5, wherein the contrastive learning model includes an augmentation module, an encoder, and a projection head.
[0073] Clause 7. The processor-implemented method according to Clause 6, wherein the augmented sample is generated via the augmentation module by augmenting the clean sample with noise.
[0074] The various operations of the above method can be performed by any suitable component capable of performing the corresponding functions. These components can include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Generally, in the case where an operation is illustrated in the drawings, these operations can have corresponding paired components plus functional components with similar numbers.
[0075] As used, the term "determine" encompasses a wide variety of actions. For example, "determine" can include calculating, computing, processing, deriving, researching, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, etc. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Furthermore, "determine" can include parsing, selecting, choosing, establishing, etc.
[0076] As used, the phrase "at least one of" in reference to a list of items refers to any combination of those items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c.
[0077] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed with a general - purpose processor, a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the described functions. A general - purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0078] The steps or algorithms of the methods described in connection with the present disclosure may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software modules may reside in any form of storage medium known in the art. Some examples of storage media that may be used include random access memory (RAM), read - only memory (ROM), flash memory, erasable programmable read - only memory (EPROM), electrically erasable programmable read - only memory (EEPROM), registers, hard disk, removable disk, CD - ROM, and the like. The software modules may include a single instruction, or many instructions, and may be distributed over several different code segments, among different programs, and across multiple storage media. The storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor.
[0079] The disclosed methods include one or more steps or acts for implementing the described methods. The steps and / or acts of the methods may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or acts is specified, the order and / or use of specific steps and / or acts may be modified without departing from the scope of the claims.
[0080] The described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnected buses and bridges. The bus may link together various circuits, including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect a network adapter, etc. to the processing system via the bus. The network adapter may be used to implement signal processing functions. For some aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits, such as a timing source, peripheral devices, voltage regulators, power management circuits, etc., which are well known in the art and will not be described further.
[0081] The processor may be responsible for managing the bus and general processing, including executing software stored on the machine-readable medium. The processor may be implemented using one or more general-purpose processors and / or dedicated processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuits that can execute software. Software should be broadly interpreted to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. By way of example, the machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard disk drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging materials.
[0082] In a hardware implementation, the machine-readable medium may be part of a processing system separate from the processor. However, as will be readily understood by those skilled in the art, the machine-readable medium or any part thereof may be external to the processing system. By way of example, the machine-readable medium may include a transmission line, a carrier modulated by data, and / or a computer product separate from the device, all of which may be accessed by the processor via the bus interface. Alternatively or in addition, the machine-readable medium or any part thereof may be integrated into the processor, such as in the case of having a cache and / or a general register file. Although the various components discussed may be described as having specific locations, such as local components, they may also be configured in various ways, such as some components being configured as part of a distributed computing system.
[0083] The processing system can be configured as a general-purpose processing system having one or more microprocessors providing processor functionality and an external memory providing at least a portion of the machine-readable medium, all of these components being linked together via an external bus architecture with other support circuitry. Alternatively, the processing system can include one or more neuromorphic processors for implementing the described neuron models and nervous system models. As yet another alternative, the processing system can be implemented with an application-specific integrated circuit (ASIC) having a processor, a bus interface, a user interface, support circuitry, and at least a portion of the machine-readable medium integrated on a single chip, or with one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functionality of the processing system depending on the particular application and the overall design constraints imposed on the system as a whole.
[0084] The machine-readable medium can include a plurality of software modules. These software modules include instructions that, when executed by the processor, cause the processing system to perform various functions. The software modules can include a sending module and a receiving module. Each software module can reside in a single storage device or be distributed across multiple storage devices. By way of example, when a triggering event occurs, the software module can be loaded from a hard disk drive into RAM. During the execution of the software module, the processor can load some of the instructions into a cache to improve access speed. One or more cache lines can then be loaded into the general register file for execution by the processor. When the functionality of a software module is referred to hereinafter, it will be understood that such functionality is implemented by the processor when executing instructions from that software module. Additionally, it should be understood that aspects of the present disclosure result in improvements to the functionality of a processor, computer, machine, or other system implementing such aspects.
[0085] If implemented in software, each function can be stored on or transmitted via a computer-readable medium as one or more instructions or code. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium accessible by a computer. By way of example and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and that is accessible by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and optical disc, where disks typically reproduce data magnetically, while discs reproduce data optically with a laser. Thus, in some aspects, computer-readable media can include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media can include transitory computer-readable media (e.g., signals). The above combinations should also be included within the scope of computer-readable media.
[0086] Accordingly, some aspects can include a computer program product for performing the operations of rendering. For example, such a computer program product can include a computer-readable medium having instructions stored (and / or encoded) thereon that are executable by one or more processors to perform the described operations. For some aspects, the computer program product can include packaging material.
[0087] Moreover, it should be understood that modules and / or other suitable components for performing the described methods and techniques can be downloaded and / or otherwise obtained by a user terminal and / or a base station, where applicable. For example, such devices can be coupled to a server to facilitate the transfer of components for performing the described methods. Alternatively, the described various methods can be provided via a storage component (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or floppy disk), such that once the storage component is coupled to or provided to the device, the user terminal and / or the base station can obtain the various methods. Additionally, any other suitable technology can be utilized that is adapted to provide the described methods and techniques to the device.
[0088] It should be understood that the claims are not limited to the exact configurations and components illustrated above. Various modifications, variations, and alterations can be made to the arrangements, operations, and details of the methods and apparatuses described above without departing from the scope of the claims.
Claims
1. A processor-implemented method, comprising: generating a clean sample and an augmented sample for each of a plurality of inputs; associating the clean sample with the augmented sample for each of the plurality of inputs to form a positive pair; associating the clean sample with another clean sample associated with another one of the plurality of inputs for each of the plurality of inputs to form a negative pair; and learning one or more representations of the plurality of inputs based on the positive pairs and the negative pairs for each of the plurality of inputs.
2. The processor-implemented method according to claim 1, wherein: learning the one or more representations includes minimizing a loss for each of the plurality of inputs; the clean sample is a ground truth; and the stop gradient is a function of the embedding of the clean sample.
3. The processor-implemented method according to claim 1, further comprising learning the one or more representations in a self-supervised manner via contrastive learning.
4. The processor-implemented method according to claim 1, wherein each of the plurality of inputs is an audio input.
5. The processor-implemented method according to claim 1, further comprising receiving each input at a contrastive learning model.
6. The processor-implemented method according to claim 5, wherein the contrastive learning model includes an augmentation module, an encoder, and a projection head.
7. The processor-implemented method according to claim 6, wherein the augmented sample is generated via the augmentation module by augmenting the clean sample with noise.
8. An apparatus, comprising: one or more processors; and one or more memories coupled to the one or more processors and storing instructions that, when executed by the one or more processors, are operative to cause the apparatus to: generate a clean sample and an augmented sample for each of a plurality of inputs; associate the clean sample with the augmented sample for each of the plurality of inputs to form a positive pair; associate the clean sample with another clean sample associated with another one of the plurality of inputs for each of the plurality of inputs to form a negative pair; and learn one or more representations of the plurality of inputs based on the positive pairs and the negative pairs for each of the plurality of inputs.
9. The apparatus according to claim 8, wherein: execution of the instructions further causes the apparatus to minimize a loss for each of the plurality of inputs based on learning the one or more representations; the clean sample is a ground truth; and the stop gradient is a function of the embedding of the clean sample.
10. The apparatus according to claim 8, wherein execution of the instructions further causes the apparatus to learn the one or more representations in a self-supervised manner via contrastive learning.
11. The apparatus according to claim 8, wherein each of the plurality of inputs is an audio input.
12. The apparatus according to claim 8, wherein execution of the instructions further causes the apparatus to receive each input at a contrastive learning model.
13. The apparatus according to claim 12, wherein the contrastive learning model includes an augmentation module, an encoder, and a projection head.
14. The apparatus according to claim 13, wherein the augmented samples are generated by the augmentation module based on augmenting the clean samples with noise.
15. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by one or more processors and comprising: program code for generating, for each of a plurality of inputs, a clean sample and an augmented sample; program code for associating, for each of the plurality of inputs, the clean sample with the augmented sample to form a positive pair; program code for associating, for each of the plurality of inputs, the clean sample with another clean sample associated with another one of the plurality of inputs to form a negative pair; and program code for learning, based on the positive pairs and the negative pairs for each of the plurality of inputs, one or more representations of the plurality of inputs.
16. The non-transitory computer-readable medium according to claim 15, wherein: the program code for learning the one or more representations includes program code for minimizing a loss for each of the plurality of inputs; the clean sample is the ground truth; and stop gradient is a function of the embedding of the clean sample.
17. The non-transitory computer-readable medium according to claim 15, wherein the program code further includes program code for learning the one or more representations in a self-supervised manner via contrastive learning.
18. The non-transitory computer-readable medium according to claim 15, wherein each of the plurality of inputs is an audio input.
19. The non-transitory computer-readable medium according to claim 15, wherein the program code further includes program code for receiving each input at a contrastive learning model.
20. The non-transitory computer-readable medium according to claim 19, wherein the contrastive learning model includes an augmentation module, an encoder, and a projection head.
21. The non-transitory computer-readable medium according to claim 20, wherein the augmented samples are generated by the augmentation module based on augmenting the clean samples with noise.
22. An apparatus, comprising: means for generating, for each of a plurality of inputs, a clean sample and an augmented sample; means for associating, for each of the plurality of inputs, the clean sample with the augmented sample to form a positive pair; means for associating, for each of the plurality of inputs, the clean sample with another clean sample associated with another one of the plurality of inputs to form a negative pair; and means for learning, based on the positive pairs and the negative pairs for each of the plurality of inputs, one or more representations of the plurality of inputs.
23. The apparatus according to claim 22, wherein: the means for learning the one or more representations includes means for minimizing a loss for each of the plurality of inputs; the clean sample is the ground truth; and The stop gradient is a function of the embedding of the clean sample.
24. The apparatus according to claim 22, wherein the execution of the instructions further causes the apparatus to learn the one or more representations in a self-supervised manner via contrastive learning.
25. The apparatus according to claim 22, wherein each of the plurality of inputs is an audio input.
26. The apparatus according to claim 22, further comprising: means for receiving each input at the contrastive learning model.
27. The apparatus according to claim 26, wherein the contrastive learning model comprises an augmentation module, an encoder, and a projection head.
28. The apparatus according to claim 27, wherein the augmented sample is generated via the augmentation module by augmenting the clean sample with noise.
Citation Information
Cited By
Prison break prompt generation method based on large model, electronic equipment, storage medium and computer program product
CN121744338A
Big model-based jailbreak prompt generation method, electronic device, storage medium and computer program product
CN121744338B