Unsupervised data clustering and object recognition with attention-based congruency mechanism

The attention-based congruency mechanism for unsupervised data clustering addresses inefficiencies in traditional attention mechanisms by reducing computational complexity and memory usage, and enhancing adaptability in dynamic environments, suitable for computer vision and NLP tasks.

WO2026035896A1PCT designated stage Publication Date: 2026-02-12STOWERS INST FOR MEDICAL RES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/040989
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-08
Filing Date
2025-08-06
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Traditional attention mechanisms in learning systems require multiple copies of data for embeddings, leading to high computational complexity and memory usage, and rely on backpropagation and supervised learning with labeled data, limiting scalability and applicability in resource-constrained environments.

Method used

An attention-based congruency mechanism for unsupervised data clustering that creates a clustering distribution by moving data points closer together based on similarity, without requiring backpropagation or labeled data, using a continuous learning mechanism to adapt to new data and update anchor points.

Benefits of technology

Reduces computational complexity and memory usage while maintaining high accuracy, enabling efficient data processing and adaptability in dynamic environments, suitable for computer vision and NLP tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025040989_12022026_PF_FP_ABST
    Figure US2025040989_12022026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates to improved computerized systems and methods for performing unsupervised data clustering using an attention-based congruency mechanism. A data clustering model can be configured to create a clustering distribution in multi-dimensional embedding space using embeddings extracted from visual or text inputs. The model generates an anchor point list from a subset of embeddings and maps remaining embeddings to data points around these anchors. The attention-based congruency mechanism moves data points closer to their nearest anchor points, tightening the clusters around the anchor points. Continuous learning is applied to update anchor points as new data is processed. The enhanced clustering distributions can be used to aid attention mechanisms used in performing various machine learning tasks, including computer vision or natural language processing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No.: 1065334.000206-PCT UNSUPERVISED DATA CLUSTERING AND OBJECT RECOGNITION WITH ATTENTION-BASED CONGRUENCY MECHANISM CROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims benefit of U.S. Provisional Patent Application Serial No.63 / 680,922, filed on August 8, 2024, the entire content of which is hereby incorporated by reference. TECHNICAL FIELD

[0002] This disclosure is related to improved systems, methods, and techniques for unsupervised data clustering and object recognition using an attention- based congruency mechanism. In certain embodiments, the systems, methods, and techniques described herein can be applied to improve attention mechanisms that are applied in connection with performing computer vision tasks and / or natural language processing (NLP) tasks executed by various types of learning systems. BACKGROUND

[0003] Learning systems, such as language model systems or computer vision systems, can be configured to perform various tasks. In some examples, computer vision systems can be configured to perform detection tasks, classification tasks, and segmentation tasks. These computer vision functions can be applied in many different contexts, such as facial recognition, medical image analysis, smart surveillance, and / or image analysis tasks. Language model systems, such as transformer models or other large language models (LLMs), can be configured to perform various types of NLP tasks, languagemodeling, question-answering, natural language generation, text interpretation or understanding, speech recognition, machine translation, optical character recognition, handwriting recognition, grammar induction, information retrieval, etc.

[0004] These learning systems typically rely on attention mechanisms to selectively focus on relevant parts of input data when performing tasks. Attention mechanisms allow the system to dynamically assign importance weights to different elements of the input, enabling it to prioritize and process the most relevant information for a given task. In computer vision applications, attention mechanisms may help the system focus on specific regions or objects within an image. For natural language processing tasks, attention can help the system identify key words or phrases in a sentence. By incorporating attention, learning systems can more effectively capture long-range dependencies and context information, leading to improved performance on various machine learning tasks such as image classification, object detection, machine translation, and text summarization.

[0005] Traditional attention mechanisms have several notable disadvantages. These mechanisms typically require multiple copies of the same data to form multiple embeddings, which can significantly increase computational complexity and memory usage. This redundancy in data representation may lead to inefficiencies in both training and inference processes, particularly when dealing with large-scale datasets or resource-constrained environments

[0006] Furthermore, traditional attention mechanisms often rely heavily on backpropagation for training, which can be computationally expensive and time-consuming, especially for deep neural networks. This dependence on backpropagation may limit the scalability of these models and make them less suitable for real-time or online learning scenarios.

[0007] Additionally, many conventional attention-based approaches require supervised learning with labeled data, which can be a significant bottleneck in practical applications. Obtaining high-quality labeled datasets is often costly, time-consuming, and may not be feasible in certain domains where data annotation is challenging or impractical. This reliance on labeled data may restrict the applicability of these models in scenarios where unlabeled data is abundant but labeled data is scarce. BRIEF DESCRIPTION OF DRAWINGS

[0008] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office, upon request and payment of the necessary fee.

[0009] To facilitate further description of the embodiments, the following drawings are provided, in which like references are intended to refer to like or corresponding parts, and in which:

[0010] FIG.1A is a diagram of an exemplary system for learning system(s) in accordance with certain embodiments.

[0011] FIG. 1B is a diagram demonstrating exemplary features of learning system(s) in accordance with certain embodiments.

[0012] FIG.2 is a diagram illustrating a process flow in accordance with certain embodiments.

[0013] FIG. 3A is a flowchart illustrating an exemplary method in accordance with certain embodiments.

[0014] FIG. 3B is a flowchart illustrating another exemplary method in accordance with certain embodiments.

[0015] FIG. 3C is a flowchart illustrating another exemplary method in accordance with certain embodiments.

[0016] FIG.4 illustrates graphs that demonstrate an effect of an attention-based congruency function on a clustering distribution in in accordance with certain embodiments.

[0017] FIG.5A is a diagram illustrating a typical attention model that does not apply attention-based congruency techniques.

[0018] FIG.5B is a diagram illustrating a typical attention model that does not apply attention-based congruency techniques.

[0019] FIG. 5C is a diagram illustrating how attention-based congruency techniques can be applied to increase nearness of a data point to a nearest anchor point in accordance with certain embodiments.

[0020] FIG. 5D is a diagram illustrating exemplary implementation details of attention-based congruency techniques.

[0021] FIG. 5E is a diagram illustrating details for implementing an unsupervised attention layer according to certain embodiments.

[0022] FIG. 5F is a flow diagram illustrating the attention-based congruency mechanism together with the anchor point selection process.

[0023] FIG.5G includes graphical depictions illustrating how data points can be selected for inclusion in an anchor point list using norm-based and / or similarity- based selection techniques in accordance with certain embodiments.

[0024] FIG.5H is a diagram illustrating how multiple congruence enhancement layers can be applied by the data clustering model according to certain embodiments. DEFINITIONS

[0025] The terms “first,” “second,” “third,” “fourth,” and the like in the description and in the claims, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that the embodiments described herein are, for example, capable of operation in sequences other than those illustrated or otherwise described herein.

[0026] The terms “left,” “right,” “front,” “rear,” “back,” “top,” “bottom,” “over,” “under,” and the like in the description and in the claims, if any, are used for descriptive purposes and not necessarily for describing permanent relative positions. It is to be understood that the terms so used are interchangeableunder appropriate circumstances such that the embodiments of the apparatus, methods, and / or articles of manufacture described herein are, for example, capable of operation in other orientations than those illustrated or otherwise described herein.

[0027] As used herein, “approximately” can, in some embodiments, mean within plus or minus ten percent of the stated value. In other embodiments, “approximately” can mean within plus or minus five percent of the stated value. In further embodiments, “approximately” can mean within plus or minus three percent of the stated value. In yet other embodiments, “approximately” can mean within plus or minus one percent of the stated value.

[0028] Certain data or functions may be described as “real-time,” “near real- time,” or “substantially real-time” within this disclosure. Any of these terms can refer to data or functions that are processed with a humanly imperceptible delay or minimal humanly perceptible delay. Alternatively, these terms can refer to data or functions that are processed within a specific time interval (e.g., in the order of milliseconds, microseconds, or nanoseconds). DETAILED DESCRIPTION OF EXEMPLARY EMBODIMENTS

[0029] The present disclosure relates to systems, methods, apparatuses, computer program products, and techniques for performing data clustering using an attention-based congruency mechanism on embeddings mapped in a multi-dimensional space. The attention-based congruency mechanism described herein provides an unsupervised approach to implementingattention-like functionality for clustering and enhancing similarity between data points in a clustering distribution.

[0030] In certain embodiments, a data clustering model is configured to create a clustering distribution from feature embeddings extracted from a batch of inputs and the attention-based congruency mechanism can enhance groupings in the distribution by moving data points closer together in scenarios where the data points demonstrate a certain degree of clustering. In contrast to many traditional techniques, the techniques described herein do not require learning multiple forms of embeddings through backpropagation or similar mechanisms, nor do they require labeled datasets during learning. Rather, these techniques can be applied efficiently to generate a clustering distribution in manner that is computationally inexpensive while achieving a high degree of accuracy.

[0031] The enhanced clustering distribution generated by the data clustering model can be leveraged by computer vision systems, language model systems, and / or other AI or machine-learning systems to apply attention mechanisms in connection with performing various tasks. In computer vision applications, embeddings extracted from visual inputs (e.g., pixels, images, objects, scenes, etc.) can be processed by the data clustering model to create an initial clustering distribution, and the attention-based congruency mechanism can operate to reduce the distance among data points that demonstrate similarity to each other. The enhanced clustering distribution can serve as an attention mechanism that is leveraged by the computer vision system to enhance performance of one or more computer vision tasks, such as detection,recognition, segmentation, and / or generative tasks. Along similar lines, for language model systems, embeddings extracted from text inputs (e.g., words, sentences, paragraphs, documents, etc.) can be processed by the data clustering model to create an initial clustering distribution, and the attention- based congruency mechanism can operate to reduce the distance among data points that demonstrate similarity to each other. The enhanced clustering distribution can serve as an attention mechanism that is leveraged by the language model system to focus attention on key words or phrases in text inputs and / or to enhance performance of one or more NLP tasks, such as language modeling, question-answering, natural language generation, etc. The clustering distributions output by the data clustering model described herein also can be applied to enhance tasks performed by other types of AI or machine-learning systems.

[0032] In certain embodiments, the data clustering model can be configured to execute an initialization procedure that creates an anchor point list for a clustering distribution. The model selects a data point with the maximum norm from the first batch of embeddings as the initial anchor point. Additional anchor points can be identified by selecting data points that are orthogonal or near orthogonal to the initial anchor points in a multi-dimensional embedding space.

[0033] Upon initialization of the anchor point list, the data clustering model can be configured to map embeddings to data points in the multi-dimensional embedding space where each dimension corresponds to a feature included in the embeddings. In some embodiments, the multi-dimensional space can be ahigh-dimensional space that includes 2,048 dimensions, 5,000 dimensions, and / or other large number of dimensions. The attention-based congruency function can be applied to move data points closer to their nearest anchor point. This can be achieved by first identifying the nearest anchor point for each data point, then adjusting the position of the data point to increase its nearness or similarity to that anchor point. The process effectively tightens the clusters around the anchor points, enhancing the overall clustering distribution in a manner that can be leveraged for attention purposes.

[0034] The data clustering model also can include a continuous learning mechanism that allows it to dynamically adapt and improve its clustering performance as new data is received and processed. This mechanism enables the model to update its anchor points and refine the clustering distribution, without requiring retraining on the entire dataset. When an embedding of a new sample or a new batch of embeddings are processed, the continuous learning mechanism evaluates whether to add new anchor points or replace existing ones based on similarity and norm calculations. This approach allows the model to capture evolving patterns in the data and maintain an up-to-date representation of the clustering structure. The continuous learning mechanism enhances the model's ability to handle non-stationary or dynamic data distributions and adapt to changes in the input space over time, making it particularly suitable for applications where data characteristics may shift, or new categories may emerge.

[0035] The continuous learning mechanism enables the data clustering model to create new anchor points corresponding to new clusters as new sets of data inputs are encountered. Additionally, it facilitates updating of the anchor point list through a replacement mechanism. This replacement mechanism evaluates newly received data points and may replace existing anchor points if a new point has a higher norm and satisfies certain similarity criteria. For example, the replacement measure may compute a similarity metric (e.g., a cosine similarity, Pearson similarity, Euclidian similarity, Hamming distance, and / or other similarity metric) indicating a similarity between a new data point and an existing anchor point that is nearest to the new data point and, if the new data point is within a predetermined distance of the existing anchor point, the norm of the new data point may be compared with the norm of the existing anchor point to determine whether new data point should replace the existing anchor point. For each batch of inputs, the process continues until the model has identified a set of maximally distant points that effectively represent distinct clusters in the embedding space.

[0036] The data clustering techniques described herein provide several notable advantages over conventional approaches. In contrast to traditional techniques, the techniques described herein do not require backpropagation for training (or resorting to using reconstruction error or credit assignment), which can significantly reduce computational complexity and training time. Additionally, the system does not need multiple copies of the same data to form embeddings, leading to more efficient memory usage and streamlined data processing. Theunsupervised nature of the clustering mechanism allows it to discover patterns and structures in data without the need for labeled training sets, making it particularly valuable in scenarios where labeled data is scarce or expensive to obtain. Furthermore, the improved continuous learning mechanism enables the system to dynamically adapt and refine its clustering performance as new data is encountered, without requiring complete retraining. This continuous learning capability allows the model to evolve and improve over time, maintaining its relevance and accuracy in dynamic data environments.

[0037] The technologies disclosed herein can be used in a variety of different contexts and environments. One useful application of these technologies is to improve attention mechanisms employed in the context of computer vision, which can be applied across a wide variety of different computer vision applications and tasks. Another useful application of these technologies is to improve attention mechanisms employed in the context of language model systems, which can be applied across a wide variety of different NLP applications and tasks.

[0038] In some exemplary use cases, the technologies can be applied to enhance computer vision applications that perform facial recognition, surveillance, scene analysis applications (e.g., which may be used in automated, unmanned, and / or autonomous vehicles that rely on automated, unmanned, and / or autonomous systems to control the vehicles), intelligent or automated traffic control, satellite imaging, quality control, medical imaging analysis, agricultural analysis, and / or image editing applications. In otherexemplary use cases, the technologies can be applied to enhance language model systems that perform AI chat functions, speech recognition, question answering, language modeling, and / or other NLP functions.

[0039] The technologies disclosed herein can also be applied to many other contexts as well. For example, they can be used to process and / or analyze DNA and RNA sequences, auditory data, sensory data, or data collected from other sources. In these contexts, the data clustering model can identify, categorize, or extract other information from the inputted data related to patterns or other features of the data. The data clustering model can generally perform the same functions related to extracting representations and / or classifying portions of the inputted data as it can with visual images. The data to be analyzed and / or processed by the data clustering model can be pre- processed in some way, such as by converting it into pixel values or pseudo image values that can be input into the data clustering model for processing. Other preprocessing steps, such as scaling and / or applying a wavelet or Fourier transform, can be applied to inputs of all types.

[0040] The embodiments described in this disclosure can be combined in various ways. Any aspect or feature that is described for one embodiment can be incorporated to any other embodiment mentioned in this disclosure. Moreover, any of the embodiments described herein may be hardware-based, may be software-based, or, preferably, may comprise a mixture of both hardware and software elements. Thus, while the description herein may describe certain embodiments, features, or components as being implementedin software or hardware, it should be recognized that any embodiment, feature and / or component referenced in this disclosure can be implemented in hardware and / or software.

[0041] FIG. 1A is a diagram of an exemplary system 100 in accordance with certain embodiments. FIG. 1B is a diagram illustrating exemplary features and / or functions associated with a learning system 125.

[0042] The system 100 comprises one or more computing device(s) 110 and one or more server(s) 120 that are in communication over a network 105. A learning system 125 is stored on, and executed by, the one or more server(s) 120. The learning system 125 can correspond to, or include, one or more computer vision systems 150 and / or one or more language model systems 155. Additionally, the learning system 125 can include a data clustering model 140 that generates clustering distributions based on received inputs, which can be used to enhance performance of machine-learning functions, such as computer vision functions and / or NLP functions, that are executed by the learning system 125.

[0043] The network 105 may represent any type of communication network, e.g., such as one that comprises a local area network (e.g., a Wi-Fi network), a personal area network (e.g., a Bluetooth network), a wide area network, an intranet, the Internet, a cellular network, a television network, a satellite communication network, and / or other types of networks.

[0044] All the components illustrated in FIGS. 1A and 1B, including the computing device(s) 110, server(s) 120, learning system(s) 125, data clusteringmodel 140, computer vision system(s) 150, and language model system(s) 155 can be configured to communicate directly with each other and / or over the network 105 via wired or wireless communication links, or a combination of the two. Each of the computing device(s) 110, server(s) 120, learning system(s) 125, data clustering model 140, computer vision system(s) 150, and language model system(s) 155 can include one or more computer storage devices 201, one or more processing devices 202, and / or one or more communication devices 203.

[0045] The one or more computer storage devices 201 may include (i) non- volatile memory, such as, for example, read only memory (ROM) and / or (ii) volatile memory, such as, for example, random access memory (RAM). The non-volatile memory may be removable and / or non-removable non-volatile memory. RAM may include dynamic RAM (DRAM), static RAM (SRAM), etc. Further, ROM may include mask-programmed ROM, programmable ROM (PROM), one-time programmable ROM (OTP), erasable programmable read- only memory (EPROM), electrically erasable programmable ROM (EEPROM) (e.g., electrically alterable ROM (EAROM) and / or flash memory), etc. ROM also can include cloud storage and / or flash drive storage. In certain embodiments, the one or more computer storage devices 201 include physical, non-transitory mediums. The one or more computer storage devices 201 can store instructions for implementing any of the functionalities associated with the learning system(s) 125 and / or data clustering model 140 described herein.

[0046] The one or more processing devices 202 may include one or more central processing units (CPUs), one or more microprocessors, one or more microcontrollers, one or more controllers, one or more complex instruction set computing (CISC) microprocessors, one or more reduced instruction set computing (RISC) microprocessors, one or more very long instruction word (VLIW) microprocessors, one or more graphics processor units (GPU), one or more digital signal processors, one or more application specific integrated circuits (ASICs), one or more neuromorphic processors, and / or any other type of processor or processing circuit capable of performing desired functions. The one or more processing devices 202 can be configured to execute any computer program instructions that are stored or included on the one or more computer storage devices 201 including, but not limited to, instructions associated with executing the functionalities of the learning system(s) 125 and / or data clustering model 140 described herein.

[0047] Each of the one or more communication devices 203 can include wired and wireless communication devices and / or interfaces that enable communications using wired and / or wireless communication techniques. Wired and / or wireless communication can be implemented using any one or combination of wired and / or wireless communication network topologies (e.g., ring, line, tree, bus, mesh, star, daisy chain, hybrid, etc.) and / or protocols (e.g., personal area network (PAN) protocol(s), local area network (LAN) protocol(s), wide area network (WAN) protocol(s), cellular network protocol(s), powerline network protocol(s), etc.). Exemplary PAN protocol(s) can comprise Bluetooth,Zigbee, Wireless Universal Serial Bus (USB), Z-Wave, etc. Exemplary LAN and / or WAN protocol(s) can comprise Institute of Electrical and Electronic Engineers (IEEE) 802.3 (also known as Ethernet), IEEE 802.11 (also known as Wi-Fi), etc. Exemplary wireless cellular network protocol(s) can comprise Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Evolution-Data Optimized (EV-DO), Enhanced Data Rates for GSM Evolution (EDGE), Universal Mobile Telecommunications System (UMTS), Digital Enhanced Cordless Telecommunications (DECT), Digital AMPS (IS-136 / Time Division Multiple Access (TDMA)), Integrated Digital Enhanced Network (iDEN), Evolved High-Speed Packet Access (HSPA+), Long-Term Evolution (LTE), WiMAX, etc. The specific communication software and / or hardware can depend on the network topologies and / or protocols implemented. In certain embodiments, exemplary communication hardware can comprise wired communication hardware including, but not limited to, one or more data buses, one or more universal serial buses (USBs), one or more networking cables (e.g., one or more coaxial cables, optical fiber cables, twisted pair cables, and / or other cables). Further exemplary communication hardware can comprise wireless communication hardware including, for example, one or more radio transceivers, one or more infrared transceivers, etc. Additional exemplary communication hardware can comprise one or more networking components (e.g., modulator-demodulator components, gateway components, etc.). In certain embodiments, the one or more communication devices 203can include one or more transceiver devices, each of which includes a transmitter and a receiver for communicating wirelessly. The one or more communication devices 203 also can include one or more wired ports (e.g., Ethernet ports, USB ports, auxiliary ports, etc.) and related cables and wires (e.g., Ethernet cables, USB cables, auxiliary wires, etc.).

[0048] In certain embodiments, the one or more communication devices 203 additionally, or alternatively, can include one or more modem devices, one or more router devices, one or more access points, and / or one or more mobile hot spots. For example, modem devices may enable the computing device(s) 110, server(s) 120, learning system(s) 125, computer vision system(s) 150 and / or language model system(s) 155 to be connected to the Internet and / or other networks. The modem devices can permit bi-directional communication between the Internet (and / or other network) and the computing device(s) 110, server(s) 120, learning system(s) 125, computer vision system(s) 150, and / or language model system(s) 155. In certain embodiments, one or more router devices and / or access points may enable the computing device(s) 110, server(s) 120, learning system(s) 125, computer vision system(s) 150 and / or language model system(s) 155 to be connected to a LAN and / or other more other networks. In certain embodiments, the one or more mobile hot spots may be configured to establish a LAN (e.g., a Wi-Fi network) that is linked to another network (e.g., a cellular network). The mobile hot spot may enable the computing device(s) 110, server(s) 120, learning system(s) 125, computervision system(s) 150, and / or language model system(s) 155 to access the Internet and / or other networks.

[0049] In certain embodiments, the computing device(s) 110 may represent desktop computers, laptop computers, mobile devices (e.g., smart phones, personal digital assistants, tablet devices, vehicular computing device(s), wearable devices, or any other device that is mobile in nature), gaming consoles and / or other types of devices. The one or more server(s) 120 may generally represent any type of computing device, including any of the computing device(s) 110 mentioned above. The one or more server(s) 120 also can comprise one or more mainframe computing device(s), one or more virtual server(s), one or more application server(s), and / or one or more cloud-based server(s). In some embodiments, the one or more server(s) 120 can be configured to execute web server(s) and can communicate with the computing device(s) 110, external data sources, client systems, and / or other devices over the network 105 (e.g., over the Internet).

[0050] As mentioned above, some or all of the computing device(s) 110 may represent mobile electronic devices in certain embodiments. Generally speaking, the mobile electronic devices can include any type of electronic device that is portable and / or transportable in nature. In some cases, a mobile electronic device can refer to a portable electronic device (e.g., an electronic device easily conveyable by hand by a person of average size) with the capability to present audio and / or visual data (e.g., text, images, videos, music, etc.). For example, a mobile electronic device can comprise at least one of acellular telephone (e.g., a smartphone), a personal digital assistant, a handheld digital computer device (e.g., a tablet personal computer device), a digital media player, a wearable device, and / or another portable computer device with the capability to present audio and / or visual data (e.g., images, videos, music, etc.). Thus, in many examples, a mobile electronic device can comprise a volume and / or weight sufficiently small as to permit the mobile electronic device to be easily conveyable by hand. For examples, in some embodiments, a mobile electronic device can occupy a volume of less than or equal to approximately 1790 cubic centimeters, 2434 cubic centimeters, 2876 cubic centimeters, 4056 cubic centimeters, and / or 5752 cubic centimeters. Further, in these embodiments, a mobile electronic device can weigh less than or equal to 15.6 Newtons, 17.8 Newtons, 22.3 Newtons, 31.2 Newtons, and / or 44.5 Newtons.

[0051] Exemplary mobile electronic devices can comprise (i) an iPod®, iPhone®, iTouch®, iPad®, and / or similar products offered by Apple Inc. of Cupertino, California, United States of America; (ii) a Blackberry® or similar product by Research in Motion (RIM) of Waterloo, Ontario, Canada; (iii) a Lumia® or similar product by the Nokia Corporation of Keilaniemi, Espoo, Finland, and / or (iv) a Galaxy® or similar product by the Samsung Group of Samsung Town, Seoul, South Korea. Further, in the same or different embodiments, a mobile electronic device can comprise an electronic device configured to implement one or more of (i) the iOS® or iPhone® operating system by Apple Inc. of Cupertino, California, United States of America, (ii) theBlackberry® operating system by Research In Motion (RIM) of Waterloo, Ontario, Canada, (iii) the Palm® operating system by Palm, Inc. of Sunnyvale, California, United States, (iv) the Android® operating system developed by the Open Handset Alliance, (v) the Windows Mobile™ operating system by Microsoft Corp. of Redmond, Washington, United States of America, or (vi) the Symbian™ operating system by Nokia Corp. of Keilaniemi, Espoo, Finland.

[0052] The mobile electronic devices can additionally, or alternatively, include wearable devices (e.g., wearable user computer devices) as mentioned above. Generally speaking, wearable devices can generally include any type of electronic device that is capable of be mounted to, worn by, and / or fixed to an individual. For example, in some cases, the wearable devices sometimes can be worn under or over clothing, and / or integrated with the clothing and / or other accessories (e.g., hats, eyeglasses, wristbands, watches, shoes, gloves, etc.). In some cases, wearable devices can be directly mounted or attached to individuals (e.g., the individuals’ head, wrist, arms, legs, or neck regions). The wearable devices can comprise a head mountable wearable user computer device (e.g., one or more head mountable displays, one or more eyeglasses, one or more contact lenses, one or more retinal displays, etc.) and / or a limb mountable wearable user computer device (e.g., a smart watch). In some configurations, the wearable devices can be configured to present audio and / or visual data (e.g., text, images, videos, audio, music, etc.) and / or to receive inputs from individuals (e.g., via one or more input devices such astouchscreens, switches, buttons, etc.). The mobile electronic devices can include additional types of devices other than those explicitly mentioned herein.

[0053] In certain embodiments, the learning system(s) 125 can be stored on, and executed by, the one or more server(s) 120. Additionally, or alternatively, the learning system(s) 125 can be stored on, and executed by, the one or more computing device(s) 110. Thus, in some embodiments, the learning system(s) 125 can be stored as server application on one or more server(s) 120 and, in other embodiments, can be stored as a local application on a computing device 110, or integrated with a local application stored on a computing device 110.

[0054] Additionally, in some embodiments, the learning system(s) 125 can be implemented as a combination of a front-end application (e.g., which is stored on a computing device 110) and a back-end application (e.g., which is stored on one or more server(s) 120). All functionalities of the learning system(s) 125 described herein can be executed by the front-end application, the back-end application, or a combination of both.

[0055] In some embodiments, the learning system(s) 125 can be implemented locally on a computing device 110. The computing device 110 can be trained to implement the learning system(s) 125 without communicating with the network 105. In some embodiments, the learning system(s) 125 can be used for an injury in a patient where transmission of sensory feeling is lost or compromised. The learning system(s) 125 can be implemented locally within the body of a patient and gather inputs from other parts of the body of the patient. The inputs can be clustered based on a similarity criterion of typicalresponses. The clustered information can be transmitted to the central nervous system, bypassing the injured part.

[0056] The learning system(s) 125 can additionally, or alternatively, be stored on, and executed by, the computing device(s) 110 and / or other devices. For example, in certain embodiments, the learning system(s) 125 can be integrated directly onto a camera device to enable the camera device to analyze images using the techniques described herein.

[0057] In some embodiments, the learning system(s) 125 also can be stored as a local application on a computing device 110, or integrated with a local application stored on a computing device 110, to implement the techniques and functions described herein. In certain embodiments, the learning system(s) 125 can be integrated with (or can communicate with) various applications including, but not limited to, facial recognition applications, automated vehicle applications, intelligent traffic applications, surveillance applications, security applications, industrial quality control applications, medical applications, agricultural applications, veterinarian applications, image editing applications, AI chatbot applications, social media applications, and / or other applications that are stored on a computing device 110 and / or server 120.

[0058] In certain embodiments, the one or more computing device(s) 110 can enable individuals to access the learning system(s) 125 over the network 105 (e.g., over the Internet via a web browser application).

[0059] The learning system(s) 125 of FIGS.1A and 1B can receive or extract embeddings 135 from one or more inputs 130, such as feature embeddingsextracted from visual content and / or word embeddings extracted from text content. The data clustering model 140 of the learning system 125 can utilize the embeddings 135 to compute a clustering distribution 142, by mapping the embeddings 135 to data points (also referred to as “data point indexes” herein) in a multi-dimensional space. In doing so, the data clustering model 140 may apply an attention-based congruency function 145 on the data points included in the clustering distribution 142, which operates to enhance the clustering distribution 142 by minimizing distances between embeddings 135 (or corresponding data points) that show a certain degree of similarity. The results of the data clustering model 140 can then be provided to and / or utilized by a computer vision system 150, a language model system 155, or other learning model associated with the learning system(s) 125.

[0060] The embeddings 135 can correspond to vectors and / or representations extracted from inputs 130, such as visual inputs (e.g., pixels, objects, scenes, images, etc.) and / or text inputs (e.g., letters, words, paragraphs, documents, etc.) received by the learning system(s) 125. For example, an embedding 135 can include numerical values that represent salient characteristics or attributes of a corresponding visual or text input in a high-dimensional space. The embedding 135 can be extracted from an image or text input as part of a feature extraction mechanism process. The specific contents of the embedding can depend on the feature extraction mechanism applied and / or the type of input data. In some examples, an embedding 135 extracted from an image or visual content can include vector values identifying or corresponding local features,global features, and content features of the image or visual content. For example, the vector values associated with the embedding 135 may identify or correspond to: a) edge features (presence and orientation of edges within the image); b) color information (e.g., distribution of colors in different parts of the image); c) texture features (e.g., repetitive patterns or variations in pixel intensity); d) shape features (e.g., geometric shape and structure of objects in the image); and / or e) object features (e.g., local or global objects included in a scene). In other examples, an embedding 135 extracted from a word or text content can include vector values identifying or corresponding to: a) semantic or context information; b) syntactic information; c) context proximity information; d) synonymy / antonymy information; e) grammatical relation information; and / or f) relationship and analogy information.

[0061] The data clustering model 140 can be configured to map embeddings 135 to data points in a multi-dimensional space where each dimension corresponds to a feature included in the embeddings 135. In some examples, the multi-dimensional space (also referred to as an “embedding space”) can include any number of dimensions (e.g., 50+ 100+, 300+, 2000+ dimensions). In some embodiments, the multi-dimensional space can be a high-dimensional space that includes 2,048 dimensions, 5,000 dimensions, and / or other large number of dimensions. Each value of the embedding 135 can correspond to a dimension in the multi-dimensional space. For example, if there are 2,048 dimensions there can be 2,048 separate values in the embeddings 135. Dimensionality reduction techniques can be used to project these high-dimensional vectors into 2D or 3D space for visualization to facilitate understanding of the structure and relationships within the data.

[0062] The data clustering model 140 can be configured to create a clustering distribution 142 based on a batch of embeddings 135 obtained from, or received by, the learning system 125 by mapping the embeddings 135 (or vectors) to data points in the multi-dimensional space. As new batches of embeddings 135 are obtained from the learning system(s) 125, the data clustering model 140 can continuously update the clustering distribution 142 to include additional anchor points 141 and / or data points corresponding to the newly received embeddings.

[0063] To create or initialize a clustering distribution 142, the data clustering model 140 may execute a process that identifies an initial set of anchor points 141 to be included in the clustering distribution 142. In some examples, the anchor point initialization may be performed according to the following process: 1. Create an empty list of anchor points. 2. If the anchor list is empty, initialize it with the point having maximum norm in a batch of datapoints. i. To make such selection, calculate the L2norm of vectors specifying the coordinates of points in the high-dimensional space. ii. Select the point with the highest L2norm of coordinates. 3. After initialization, update the anchor list with the points from the batch which are near orthogonal to all the points in the anchor list. 4. If any point is very similar to one of the points in the anchor list, keep the point with maximum norm.

[0064] As indicated by the exemplary initialization process provided above, the initial point on the anchor list 141 can be selected by calculating the L2norm of vectors specifying the coordinates of points in the high-dimensional space and selecting the point with the highest L2norm of coordinates. Additional anchor points 141 can be selected or identified based on the initial anchor point 141. In some examples, additional anchor points 141 can be data points that are determined to be orthogonal or near orthogonal to the initial anchor point in the high-dimensional space.

[0065] The last step of exemplary initialization process can represent an anchor point replacement mechanism or function that operates to replace or swap an anchor point for a cluster with another new data point that is received and processed by the data clustering model 140. After anchor points are initially identified, the data clustering model 140 may identify or flag new data points that are determined to be very similar to one of the anchor points included in the anchor point list. If a newly received data point has satisfied a threshold similarity to an existing anchor point, the norm of the newly received data point can be compared to the norm of the existing data point and, if the norm of the newly received data point exceeds the norm of the existing anchor point, the newly received data point may replace the existing anchor in the anchor point list.

[0066] In certain embodiments, the anchor points 141 can have an orthogonality threshold to and a similarity threshold ts. For a first batch of data points B, the empty list of anchor points Q can be populated with a data point bisuch that‖^^^^‖ is the maximum, or furthest, data point in the batch of data points B. For a second data point bj, if the similarity of bj and qk is less than the orthogonality threshold to, the anchor point list Q can be populated with bj. If the similarity of bj and qk is greater than the threshold similarity ts and bj is further from the anchor point than qk, qk can be replaced with bj in the list of anchor points Q.

[0067] The anchor point updating mechanism can be carried out in a continuous learning fashion, continuously updating Q, as new batches of data points are received and added to the clustering distribution 142. The anchor point list can be configured so that it will no longer update Q when it has found all of the most distinct points (e.g., points with similarity less than to) or near orthogonal points. In certain embodiments, the similarity threshold for the cosine or Pearson similarity between all of the anchor points can be 0.8, and if any similarity is greater than 0.8 the data point can become a member of the cluster with the corresponding anchor point. If one of the points is very close or in a similar direction to the known anchor points and another point is farther from the origin, the farther point can be kept on the anchor point list.

[0068] An exemplary technique for selecting anchor points 141 is described in Pseudocode Example 1. Pseudocode Example 1: 1. Initialize an empty list of Q, ^^^^, and ^^^^2. If ^^ is first batch: Populate ^^ with ^^^^such that ‖^^^^‖ is maximum in ^^ For ^^^^ in ^^ where ^^ ≠ ^^:For ^^^^in ^^: If similarity(^^^^, ^^^^) < ^^^^: a. Populate ^^ with ^^^^If similarity(^^^^, ^^^^) > ^^^^and ‖^^^^‖ > ‖^^^^‖:where: Q denotes a list of anchor points including ^^^^; ^^^^denotes orthogonality threshold; ^^^^denotes similarity threshold; and Bdenotes a batch of data points including ^^^^ ^^^^^^ ^^^^.

[0069] The anchor point list can be initialized with a first anchor point 141 that is the point having a maximum norm in a batch of data points since there may not be an anchor point stored in the memory. The maximum norm can be the maximum L2norm. The L2norm of a real vector v = (v1, v2, v3) can be given by ‖^^‖ = √^^2 21 + ^^2 + ^^32.135 utilized to create the clustering distribution 142 can be extracted or derived from different types of inputs 130. For embodiments in which the data clustering techniques described herein are applied to a computer vision system 150, the inputs 130 can include visual inputs, such as one or more pixels, one or more images, one or more image regions, one or more objects, one or more scenes, etc. In certain embodiments, each pixel received as an input to the learning system(s) can be an independent axis or dimension and each image can be represented by a vector in the high-dimensional space. For embodiments in which the data clustering techniques are applied to a language model system(s) 155, the inputs 130 can comprise text inputs, such as words, sentences, sentence fragments, paragraphs, etc.

[0071] The data clustering model 140 can map the anchor points 141 included in the anchor point list to a multi-dimensional space, along with the remaining data points derived from the other embeddings 135 (which are not included in the anchor list). Each of the anchor points 141 may correspond to a centroid or reference point for a corresponding cluster included the clustering distribution 142, and the remaining data points may be mapped to the multi-dimensional space to form the clusters around the anchor points 141.

[0072] In computing the clustering distribution 142 for a batch of inputs 130 (or corresponding embeddings 135 derived therefrom), an objective function 143 can be applied to adjust the positions of the data points in the multi-dimensional embedding space. In particular, the objective function 143 can include an attention-based congruency function 145 that operates to move each of the data points in the clustering distribution 142 closer to a nearest anchor point and / or minimize the distance between each data point and the nearest anchor point.

[0073] In certain embodiments, the objective function 143 can be configured to identify a nearest anchor point in the multi-dimensional space for each of the data point index that is not included on the anchor point list. The nearest anchor point can be computed by identifying the index of the anchor point in the anchor point list which is nearest to a given data point. This calculation can be doneacross all data point indices to find the convex surrogate of the maximum by computing the maximum of the log-sum exponential of ^^(^^, ^^^^):^^ = max ^^(^^,^^ ) = log(∑ ^^^^(^^,^^^^)^^^^ ^^) Theusing the following objective function 143: m ^a^x (m ^a^ x^^(^^, ^^^^)where: ^^(^^, ^^^^) =1 ^^ ^^ ^^^^^^the loss can be computed by taking the derivative of the loss with respect to q: ^^^^= ^^^^^^^^^^^^^^^^^^^^^^^^ (^^ ) where:q denotes the query point which is not on the anchor list; i denotes an index of points; s denotes the similarity metric; vi denotes the value point across the index; C denotes a normalizing constant; and T denotes a transpose matrix operation. The normalizing constant C can denote a normalizing factor that normalizes s between 0 and 1.

[0074] The attention-based congruency function 145 can be used to increase the nearness of a data point (corresponding to embedding “q”) with the maximum data point index on the anchor point list. The attention-basedcongruency function 145 can be part of, or integrated into, the objective function 143. The nearness can be increased by moving the points representing embeddings 135 in the multi-dimensional space to a corresponding nearest point in the anchor list.

[0075] Below is an exemplary process that can be applied to map the remaining data points (not included in the anchor point list) to the multi-dimensional while applying the attention-based congruency function 145 to the distribution: 1. Denote points in the anchor list as {^^1, ^^2, ^^3 … ^^^^}2. Denote any point in the batch that is not in anchor list as ^^ 3. Denote the similarity (nearness) between q and any point in anchor list as ^^(^^, ^^^^) where ^^ ∈ [1,^^].4. To find the nearest point, identify the index of the point in K which is nearest to q. As finding the maximum over indices is not feasible analytically, this can be achieved this by finding the maximum of log- sum-exponential of ^^(^^, ^^ ) i ∑ ^^(^^,^^^^)^^ .e. m^a^x ^^^^^^( ^^ ^^ )5. Once the maxof q with the max index point. Thus, results in objective function m ^a^x (m ^a^ x^^(^^, ^^^^) )

[0076] As indicated above, the points on the anchor list 141 can be denoted as {^^1, ^^2, ^^3 … ^^^^} and any point in the batch of data points that is not on theanchor point list can be denoted as q. The similarity or nearness between an embedding 135 and an anchor point 141 can be denoted as ^^(^^, ^^^^) where ^^ ∈[1,^^]. The nearest point can be the index of the point in K which is nearest to q. The maximum index can be found by finding the maximum of the log-sum- exponential of ^^(^^, ^^^^). The log-sum-exponential of ^^(^^, ^^^^) can bem ^a^ x^^^^^^(∑^^ ^^^^(^^,^^^^) ). The nearness of q to the max index point can be increasedaccording to the objective function 143: m ^a^x(m ^a^ x^^(^^, ^^^^) ). The attention-based congruency function 145 can travel the gradient of the task to optimize the task. For the list of anchor points 141, the similarity does not need to be computed for each normalized vector. Any data point that is not in the set of anchor points 141 can be multiplied across all of the data points such that the dot product with all of the anchor points can be taken to indicate the similarity. The softmax function can provide the index of the most similar data point. The attention-based congruency function 145 can operate to achieve classification of the embeddings 135.

[0077] Similarity can be enhanced by moving a data point (q) in the clustering distribution 142 closer to the nearest anchor point 141 with the maximum similarity. The loss can be optimized by following the gradient descent. The data clustering techniques can be implemented with continuous learning, such that the model continuously updates ^^ , for as long as the network isencountering new data points (e.g., received in one or more subsequently received input batches). A skip connection is used to add q to the attention applied to q. The congruence function 145 can employ multiple skip connections. To maximize the congruence function 145, a fraction of the gradient can be added to the actual value: ^^ ≔ ^^ +^^^^ ^^where:q denotes a query point which is not on the anchor list; ^^^^ ^^^^ denotes the gradient of the loss and can further be denoted as:^^^^^^(^^) ^^;to be multiplied by the gradient of the loss; C denotes a normalization constant; V denotes a matrix of the anchor points, or value points, expressed as coefficients, or weights, in the high dimensional space; and T denotes a transpose matrix operation.

[0078] Pseudocode Example 2 describes an exemplary technique that may be utilized for implementing the attention-based congruency function 145 to enhance similarity between points in the clustering distribution 142. As mentioned above, similarity can be enhanced by moving the points in the clustering distribution 142 closer to the nearest anchor point 141 by gradient descent or ascent. The softmax function can be used to turn a vector of real values into a vector of real values that sum to one. The inputs to the softmax function can be positive, negative, zero. The softmax converts scores to a normalized probability distribution. The data clustering model can employ continuous learning, which provides for the continuous updating of the anchor point list Q (e.g., such as by adding / removing anchor points and / or replacing anchor points in the list), which in turn enhances adjustments to the data points as their distances are minimized with respect to the anchor points. Pseudocode Example 2:^^^^^^ ^^ ^^^^^^_ =√^^^^^^(^^^^) (^^^^^^_)^^^^^^ = (1 − ^^)^^^^ + ^^. ^^^^^^_where: bi denotes a data point in a batch of data points B; Q denotes a query point that is not on the anchor list; T denotes a transpose matrix operation; len denotes a function that returns the length of a vector based on the number of entries in the vector and can be used to determine the number of data points in the batch; and α denotes an initialized strength.

[0079] The data clustering model 140 can operate in a first-order pixel space or in multi-order pixel space. The attention-based congruency function 145 can group the inputs 130 using the data clustering model 140. The operation of the attention-based congruency function 145 of the data clustering model 140 can provide more accurate classification of inputs 130.

[0080] In certain examples, inputs 130 corresponding to visual content may appear in the first-order pixel space as similar. An example of similar inputs can be images containing animals with similar poses (e.g., cats and dogs sitting upright). However, despite their similarities, cats and dogs are two different species that may belong to different clusters. The data clustering model 140 can distinguish species (e.g., cats and dogs) by relying on identifying features and feature combinations that are independent of the feature embeddings 135 identified in the first-order pixel space (e.g., identifying poses). Smaller imageregions can be used to determine the independent features including lines, curves, angles, or other identifying features that can be represented by the embeddings 135. The data clustering model 140 can determine the features present in the training data. In certain embodiments, the features present in the training data can include ears, head shape, faces, tails, or other identifying features. The data clustering model 140 can identify feature combinations that commonly occur in certain objects, text, or images. The data clustering model 140 can use the correlation between a previously identified feature combination and a detected feature combination from an input 130 in the batch of inputs to classify the input 130 with the previously identified feature combination or anchor point 141.

[0081] The clustering model 140 can be performed on multiple levels. In certain embodiments, with visual inputs 130, small features can be clustered according to their shape (e.g., ear shape, eye shape, face shape, nose shape, mouth shape, arm shape, leg shape). Shapes of different features for different species can be different and can be clustered by similarity. The features clustered by similarity can be used for species identification. In certain embodiments, the shape of the head or body in a visual input 130 can comprise feature combinations. The feature combinations can be drawn from a class of lower- level features. The lower-level features can include ear shape, eye shape, etc. The feature combinations can be clustered at a higher level. The higher level can include identification of certain poses. For poses, the clustering can rely on a body position of the image. The species identification information for theimage, which is being analyzed at a higher level to identify pose, can be embedded in the earlier feature clusters.

[0082] In certain examples, the inputs 130 corresponding to text content may appear in the first-order pixel space as similar. An example of similar inputs can be text inputs with similar length, sentence structure, semantic or context information, syntactic information, context proximity information, synonymy / antonymy information, grammatical relation information, and / or relationship and analogy information. An example of similar inputs can be text inputs with synonymous meaning. However, despite their similarities, the text inputs can include different words that belong to different clusters. The clustering model 140 can rely on identifying features and feature combinations that are independent of the feature embeddings 135 identified in the first-order pixel space (e.g., identifying meaning). Smaller image regions can be used to determine the independent features including sentence structure, grammatical relation information, or other identifying features that can be represented by the embeddings 135. The data clustering model 140 can determine the features present in the training data. In certain embodiments, the features present in the training data can include similar length, sentence structure, semantic or context information, syntactic information, context proximity information, synonymy / antonymy information, grammatical relation information, and / or relationship and analogy information. The data clustering model 140 can use the correlation between a previously identified feature combination and a detected feature combination from an input 130 in the batch of inputs to classifythe input 130 with the previously identified feature combination or anchor point 141.

[0083] For text inputs 130, the data clustering model 140 can be performed to cluster words according to their spelling, or by their presence in certain context to extract similarity in meaning. Synonyms, for example, often do not have similar spelling. For text inputs 130, similarity used for clustering can derive from context. Context can include word combinations in sentences. Each word in the text input 130 can be clustered according to the data clustering model 140. Each sentence or paragraph in the text input 130 can have its own combination of words or features.

[0084] The clustering distribution 142 can represent a mapping of the embeddings 135 in a multi-dimensional space. The feature or word embeddings 135 that are mapped in a multi-dimensional space according to the clustering distribution 142 can include points from a batch of data points that are very similar to the points on the anchor list. Very similar can refer to the distance of the points in the high-dimensional space or the L2norm. The very similar points that are mapped according to the clustering distribution 142 can be removed from the anchor point list. The anchor point list can include maximally distant points in the high-dimensional space.

[0085] The system configurations described herein are provided as examples to demonstrate environments in which embodiments described herein can be deployed. Numerous modifications and variations to the disclosedembodiments are possible, and the techniques described herein can be implemented in many other contexts and environments.

[0086] FIG. 2 illustrates an exemplary process flow 200 according to certain embodiments.

[0087] At block 210, a batch of inputs 130 are initially received. In some examples, the inputs may include visual inputs 130A (e.g., images) and / or text inputs 130B (e.g., words, sentences, paragraphs etc.).

[0088] For embodiments that process visual inputs, the inputs 130 provided to, and analyzed by, the learning system(s) 125 can include any type of image, video, or visual content. The inputs 130 may be captured in any digital or analog format, and using any color space or color model. The inputs 130 can be portions excerpted from a video.

[0089] In certain embodiments, the inputs 130 can include one or more two- dimensional (2D) images. In certain embodiments, the inputs 130 may include one or more three-dimensional (3D) images. Exemplary image formats can include, but are not limited to, bitmap (BMP), JPEG (Joint Photographic Experts Group), TIFF (Tagged Image File Format), GIF (Graphics Interchange Format), PNG (Portable Network Graphics), STEP (Standard for the Exchange of Product Data), etc. Exemplary color spaces or models can include, but are not limited to, sRGB (standard Red-Green-Blue), Adobe RGB, gray-scale, etc. Further, in some embodiments, some or all of the inputs 130 can be preprocessed and / or transformed prior to being analyzed by the learning system(s) 125. For example, the inputs 130 can be split into different colorelements and / or processed via a transform, such as a Fourier or wavelet transform. Other preprocessing and transformation operations also can be applied.

[0090] Further, the inputs 130 can be created from non-visual data sources, such as DNA or RNA sequences, auditory data, sensory data, and other types of data. These non-visual data can be converted to a visual domain by pixelizing the data (e.g., converting the non-visual data into an “image” including one or more “pixel” values representing portions of the non-visual data).

[0091] The inputs 130 received by the learning system(s) 125 include inputs that can be captured by any type of camera device. The camera devices can include any devices that include an imaging sensor, camera, or optical device. For example, the camera devices may represent still image cameras, video cameras, and / or other devices that include image / video sensors. The camera device can capture and / or store both visible and invisible spectra including, but not limited to, ultraviolet (UV), infrared (IR), or positron emission tomography (PET), Magnetic resonance imaging (MRI), x-ray, ultrasound, other types of medical and nonmedical imaging. The camera devices also can include devices that comprise imaging sensors, cameras, or optical devices and which are capable of performing other functions unrelated to capturing images. For example, the camera devices can include mobile devices (e.g., smart phones or cell phones), tablet devices, computing device(s), desktop computers, etc. The camera devices can be equipped with analog-to-digital (A / D) converters and / or digital-to-analog (D / A) converters based on the configuration or designof the camera devices. In certain embodiments, the computing device(s) 110 shown in FIG.1A can include any of the aforementioned camera devices, and other types of camera devices.

[0092] As mentioned above, the inputs 130 also may comprise text inputs according to some embodiments. The text inputs may comprise words, sentences, sentence fragments, paragraphs, documents, etc.

[0093] At block 220, one or more feature extraction functions 160 are executed to extract or derive embeddings 135 from the inputs 130. Any appropriate feature extraction function can be utilized. In scenarios where the inputs 130 correspond to visual inputs, the feature extraction function(s) 160 may include one or more deep learning networks and / or one or more convolutional neural networks (CNNs) that are configured to extract feature embeddings or vectors from the visual inputs. In scenarios where the inputs 130 correspond to text inputs, the feature extraction function(s) 160 may include any appropriate text- based feature extraction mechanism (e.g., Word2Vec, Doc2Vec, FastText, USE (Universal Sentence Encoder), BERT (Bidirectional Encoder Representations from Transformers), etc.

[0094] At block 230, a data clustering model 140 of a learning system 125 creates a clustering distribution 142 using the embeddings 135. The clustering distribution 142 can be generated using any of the techniques described in this disclosure. As explained throughout this disclosure, the clustering distribution 142 generated by the data clustering model 140 can be enhanced by theattention-based congruency function 145, which operates to migrate data points in the distribution closer to their nearest anchor point.

[0095] At block 240, the clustering distribution 142 is utilized by a learning system 125 to perform one or more machine learning tasks (e.g., one or more computer vision tasks and / or one or more NLP tasks). The data clustering techniques described herein are applicable to various types of learning system(s) 125 including, but not limited to, computer vision system(s) 150 and language model system(s) 155.

[0096] Thus, in some certain embodiments, the learning system(s) 125 may comprise a computer vision system 150 that utilizes the clustering distribution 142 to enhance or optimize performance of one or more computer vision tasks 180. In some examples, the computer vision system(s) 150 may utilize the clustering distribution 142 to aid performance of a detection task 181 (e.g., which may include predicting or determining whether content of an image includes certain objects or features), a classification task 182 (e.g., which may include predicting or determining whether images, or objects in the images, belong to one or more target semantic classes and / or predicting or determining labels for the images or objects), and / or a segmentation task 183 (e.g., which may include predicting or identifying precise locations of objects or features within images). In some instances, a computer vision system 150 may utilize the clustering distribution 142 to more accurately classify one or more images (or regions within the visual content) into one or more semantic classes corresponding to the clusters included in the clustering distribution 142. Thecomputer vision system(s) 150 also can be trained to perform other types of computer vision tasks that are enhanced using the clustering distribution 142 generated by the data clustering model 140.

[0097] Additionally, or alternatively, the learning system(s) 125 may comprise a language model system 155 that utilizes the clustering distribution 142 to perform one or more NLP tasks 185. Exemplary NLP tasks 185 may include natural language generation tasks, question-answering tasks, understanding tasks, speech recognition tasks, machine translation tasks, etc. In some instances, a language model system 155 may utilize the clustering distribution 142 to more accurately classify text content (e.g., words, sentences, etc.) into one or more semantic classes corresponding to the clusters included in the clustering distribution 142. The language model system 155 also can be trained to perform other types of NLP tasks that are enhanced using the clustering distribution 142 generated by the data clustering model 140.

[0098] In some embodiments, the learning system(s) 125 can utilize the data clustering model 140 and the clustering distribution 142 generated by the data clustering model 140 to perform a cluster labeling task, which applies labels to the clusters in the high-dimensional space based on the similarity of the pixels, images, or objects depending upon the level of granularity of the inputs 130 and / or extracted embeddings 135. The labels applied during the cluster labeling task can identify the pixels, images, or objects or the type of images that the objects or pixels correspond to. Once the clusters are labeled with the pre- existing labels, additional labels can be created and applied to replace or addto the pre-existing labels. The cluster labeling task can be achieved using matching algorithms to match a cluster with a label to identify the clusters or the type of images that the clusters correspond to.

[0099] For computer vision system(s) 150, the cluster labeling task can be performed as part of the classification task(s) 182. The cluster labeling task can include using matching algorithms to match each cluster with a label. The matching algorithms can include matching n number of labels with a corresponding n number of clusters.

[0100] As mentioned above, a continuous learning mechanism 146 can update the cluster distributions 142 generated by the data clustering model 140 based on new inputs 130. The continuous learning mechanism 146 enables the data clustering model 140 to create new anchor points 141 for a given clustering distribution 142. The creation of new anchor points 141 can occur when the new inputs 130 (or their corresponding embeddings 135) are not similar enough to existing anchor points 141 (e.g., do not satisfy a similarity threshold), or when the new inputs 130 (or their corresponding embeddings 135) are orthogonal to existing anchor points 141. The data point indexes corresponding to these new, non-similar inputs can be appended to the anchor point list and mapped to the clustering distribution 142.

[0101] The new anchor points 141 added to the clustering distribution 142 by the continuous learning mechanism 146 can impact subsequent processing steps. As new inputs 130 are added to the clustering distribution 142, they can be mapped to the multi-dimensional space creating a cluster around the mostsimilar anchor point 141 (whether the anchor point 141 is a new or pre-existing). The attention-based congruency function 145 is applied, the data points corresponding to the new inputs will be migrated closer to anchor points included in the updated anchor point list, which includes the newly created anchor points that are added by the continuous learning mechanism 146.

[0102] Previous learning systems 125 are trained at capacity and do not allow for updates to the data clustering model 140, computer vision system 150, or the language model system 155 when new inputs 130 are encountered. Traditional computer vision systems 150 and language model systems 155 new inputs 130 are updated during the learning phase only. During inference and execution, the inputs 130 are kept constant such that the previous systems are not able to learn indefinitely. The new inputs 130 received by previous learning systems 125 can be mapped to anchor points 141 without considering whether to add or adjust an anchor point on the anchor point list so that an anchor point is similar to the new input 130. The number of anchor points 141 in previous learning systems 125 can exceed the dimensionality of the learning system 125. A disadvantage of previous learning systems includes inaccurate mapping of data points that do not have the desired degree of similarity to the anchor points that were selected based on the initial batch of inputs.

[0103] FIG. 3A is a flowchart illustrating an exemplary method 300A for a learning system in accordance with certain embodiments. In some embodiments, the procedures, the processes, steps, and / or the activities of method 300A can be performed in the order presented. In other embodiments,the procedures, the processes, steps, and / or the activities of method 300A can be performed in any suitable order. In still other embodiments, one or more of the procedures, the processes, steps, and / or the activities of method 300A can be combined or skipped.

[0104] Method 300A of FIG.3 can include an activity 310 of receiving a batch of embeddings extracted from one or more inputs. The one or more inputs 130 can be on the pixel level, the object level, or the image level. The feature or word embeddings 135 can be extracted from the one or more inputs 130.

[0105] Method 300A also can include an activity 320 of determining data points in a high-dimensional space corresponding to the embeddings 135. One or more feature or word extraction functions 160 can be used to extract one or more points from the one or more feature or word embeddings 135.

[0106] Method 300A further can include an activity 330 of creating an anchor point list for a clustering distribution that includes a first subset of the data points.

[0107] Method 300A further can include an activity 355 of computing the data clustering distribution of a second subset of data points that are not included on the anchor point list.

[0108] Method 300A further can include an activity 360 of using the data clustering distribution to perform one or more machine-learning tasks.

[0109] FIG.3B is a flowchart illustrating an exemplary method 300B for creating an anchor point list for a clustering distribution that includes a first subset of the data points in accordance with certain embodiments. In some embodiments,method 300B can be executed to implement the activity 330 in FIG. 3A for creating an anchor point list for a clustering distribution that includes a first subset of the data points.

[0110] Method 300B can include an activity 330A of selecting one of the data points having a maximum norm to be an initial anchor point.

[0111] Method 300B of FIG.3B further can include an activity 330B of selecting additional anchor points by identifying the data points that are near orthogonal to the initial anchor point in the high-dimensional space. As newly received data points are added to the clustering distribution, an anchor point replacement function can be configured to update the anchor points included in anchor point list as part of activity 330B. Activity 330B further can include, at least in part, an activity 330C of determining a similarity between one or more of the newly received data points and a nearest anchor point in the high-dimensional space. Activity 330B further can include, at least in part, an activity 330D of replacing the nearest anchor point on the anchor point list with a newly received data point having a higher norm than the nearest anchor point.

[0112] FIG. 3C is a flowchart illustrating an exemplary method 300C for computing the data clustering distribution of a second subset of data points that are not included on the anchor point list. In some embodiments, method 300C of FIG. 3C can be executed to implement the activity 355 of FIG. 3A for computing the data clustering distribution.

[0113] Method 300C can include an activity 355A of mapping the data points included in the second subset of data points to the high-dimensional space.

[0114] Method 300C further can include an activity 355B of executing an attention-based congruency function that identifies a nearest anchor point in the high-dimensional space for each of the data points included in the second subset of data points, and moves each of the data points included in the second subset of data points closer to its nearest anchor point in the high-dimensional space.

[0115] FIG. 4 is a diagram illustrating the effect of an attention-based congruency function 145 on a clustering distribution 142. A first clustering distribution 142A is computed without using an attention-based congruency function 145, while a second clustering distribution 142B is computed using an attention-based congruency function 145. Both clustering distributions 142A, 142B are computed on embeddings 135 extracted from the same batch of inputs.

[0116] The first clustering distribution 142A is computed by directly by mapping the embeddings 135 extracted from the inputs to data points in a multi- dimensional space. The second clustering distribution 142B is computed after the attention-based congruency function 145 to applied to increase the nearness of the data points to their nearest anchor point 141. In both clustering distributions, clusters are illustrated in different colors. As shown, the clusters included in the second clustering distribution 142B are more tightly compacted relative to the clusters included in the first clustering distribution 142A.

[0117] FIG.5A is a diagram illustrating a typical attention model that does not apply the attention-based congruency techniques described in this disclosure. The scaled dot-product attention can be computed using the following equation: ^^^^^^Softmax( )^^ √^^^^where: q denotes a vector of dimension dk containing the queries k denotes a vector of dimension dk containing the keys v denotes a vector of dimension dv containing the values Q, K, and V denote matrices packing together sets of queries, keys, and values, respectively QKTdenotes a matrix where if Q is the size of m × dk, and the matrix, K, is the size of n × dk, then the resulting matrix will be of the size m × n.

[0118] FIG.5B is a diagram 600 illustrating a typical attention model that does not apply the attention-based congruency techniques described in this disclosure. In traditional machine learning, attention computation includes three steps. First, at step 610, attention is computed by matrix multiplication of a query matrix “Q” and key matrix KTas shown in FIG.5B. Second, at step 615, the product resulting from the matrix multiplication is normalized with the dimensionality of keys “dk” as shown in FIG.5B. Third, at step 620, a softmax operation is conducted over the normalized product as shown in FIG.5B.

[0119] FIG.5C is a diagram 500 illustrating an application of the attention-based congruency techniques described in this disclosure. The diagram illustrateshow a data point 510 in high-dimensional space can be moved closer to its nearest anchor point (141a, 141b, 141c, or 141d) that is determined through operation of the equations illustrated on the right side of the figure. As explained above, the attention-based congruency function 145 can be incorporated in the objective function 143 of the data clustering model 140 to move the data point 510 closer to the nearest anchor point 141 in computing a clustering distribution 142.

[0120] FIG.5D is a diagram 700 illustrating implementation details of attention- based congruency techniques. The attention-based congruency function 145 can incorporate the addition of keys to the attention outputs. FIG.5D illustrates the attention-based congruency function 145 with keys 710 and queries 720 matrices. A simplified version is shown in the right portion of FIG.5D, illustrating the same mechanism where keys have been incorporated in the attention block of the attention-based congruency function 145.

[0121] FIG. 5E is a diagram 800 is a diagram illustrating an unsupervised attention layer in accordance with certain embodiments. The attention-based congruency function 145 can be used to move a query point “q” 720 closer to one of the anchor points 141. Anchor points 141 can be selected by determining the similarity of an anchor point 141 with other points in the batch. A point “q” can undergo the similarity determination, as indicated by the “select” block 810 in Fig.5E. IF the point “q” passes the criterion, it can be put as an anchor point 141, or “Key” 710. The points “q” that do not pass the criterion, can undergo theattention-based congruency function 145 with respect to the new anchor points 141.

[0122] FIG.5F is a flow diagram 900 illustrating the attention-based congruency mechanism together with the anchor point 141 selection process. Any query point “q” 720 can be normalized. The query point “q” can go through the selection process indicated by the “select” block 810 in flow diagram 900. If the query point 720 passes the selection criterion set forth in the “select” block 810, the query point or key 710 can be included as an anchor point 141. If the query point 720 does not pass the selection criterion set forth in the “select” block, the query point can go through the attention-based congruency function 145. The key matrix for the congruency mechanism can include the anchor points 141. The final output of the flow diagram 900 can be calculated according to the attention-based congruency function 145.

[0123] FIG. 5G includes graphical depiction 950 and 970 showing how data points can be selected for inclusion in an anchor point list using norm-based and / or similarity-based selection techniques. Data points 960 in FIG. 5G represent the norms of data points in a batch. The data point with the largest norm can be used to initialize the list of anchor points. The subsequent anchor points can be selected based on similarity. Graphical depiction 970 in FIG.5G illustrates the similarity matrix of data points. The dark portions 980 can indicate low similarity values. The bright points 990 can indicate high similarity values. Data points that have low similarity value with respect to most of the points 985can be selected as anchor points. Data points that have high similarity with respect to other points 995 can be discarded.

[0124] FIG. 5H is a diagram 1000 illustrating how multiple congruence enhancement layers can be applied by the data clustering model 140 according to certain embodiments. Each layer represents moving the data point one step closer to the anchor point. Data points can be moved closer to the anchor point through operation of the objective function 143 of the data clustering model 140. The number of steps “n” can continue until each data point starts converging to a single point in the cluster. The output of one congruence step can serve as an input query to a second congruence step. Multiple steps can be sequenced together, in series or in parallel, to form a multistep congruence process.

[0125] The configurations of the learning system(s) 125 described herein can be varied in different ways.

[0126] The structure or configuration of the learning system(s) 125 can vary. In certain embodiments, the extraction mechanisms can include one or more recurrent neural networks (RNNs) and / or one or more CNNs. Additionally, in some cases, the extraction mechanism can include a Hopfield network that has been modified and optimized to perform extraction tasks using an unsupervised approach. In certain embodiments, the modified Hopfield network is a shallow, bi-layer RNN that comprises a first layer of input nodes (or input neurons) and a second layer of representation nodes (or representation neurons). Each of the representation nodes can be connected to each of the input nodes in an all- to-all configuration, and feedforward weights between the input andrepresentation nodes can be chosen to minimize the chances that two representation nodes are active at the same time. Additionally, the representation nodes can be connected to each other using recurrent connections. In some embodiments, the biased connectivity among the nodes, coupled with a stochastic gradient descent (SGD) based learning mechanism, enable the extraction mechanism to sequentially identify multiple inputs without catastrophic forgetting. The biased connectivity and lateral inhibition in the data clustering model 140 enable the representation nodes to encode structures that uniquely identify individual objects. Other extraction mechanisms also may be utilized, including any of those mentioned in this disclosure.

[0127] In certain embodiments, the learning system(s) 125 and / or computer vision system 150 may additionally comprise a convolutional neural network (CNN), or a plurality of convolutional neural networks. Each CNN may represent an artificial neural network, and may be configured to analyze images and to execute deep learning functions and / or machine learning functions on the images. Each CNN may include a plurality of layers including, but not limited to, one or more input layers, one or more output layers, one or more convolutional layers (e.g., that include learnable filters), one or more ReLU (rectifier linear unit) layers, one or more pooling layers, one or more fully connected layers, one or more normalization layers, etc. The configuration of the CNNs and their corresponding layers can be configured to enable the CNNs to learn and execute various functions for analyzing, interpreting, and understanding the images, including any of the functions described in thisdisclosure (including, but not limited to the detection functions 181, classification functions 182, segmentation functions 183, cluster labeling functions, etc.).

[0128] Each of the inputs 130 can include one or more objects. Generally speaking, any type of object or scene corresponding to an image may be included the input, and the types of objects included in an image can vary greatly. The objects included in an input may correspond to various types of inanimate articles (e.g., vehicles, beds, desks, windows, tools, appliances, industrial equipment, curtains, sporting equipment, fixtures, etc.), living things (e.g., human beings, faces, animals, plants, etc.), structures (e.g., buildings, houses, etc.), and / or the like. Thus, in some use cases, the inputs received by the learning system(s) 125 can be provided to the data clustering model 140 for processing and / or analysis, and the clustering distribution generated by the data clustering model 140 can be utilized to detect and / or classify the objects included in the inputs 130.

[0129] Additionally, or alternatively, the learning system(s) 125 and / or language model system(s) 155 may one or more language models. In some embodiments, the one or more language models may include one or more transformer models, such as one or more BERT (Bidirectional Encoder Representations from Transformers) models, one or more GPT (Generative Pre-trained Transformer) models, one or more T5 (Text-to-Text Transfer Transformer) models, etc. Additionally, or alternatively, the one or more language models may include one or more n-gram models, one or more HMMs(Hidden Markov Models), one or more VAE (Variational Autoencoders) models, one or more GANs (Generative Adversarial Networks), and / or one or more reinforced learning models.

[0130] In some scenarios, the extraction functions 160 can extract feature or word embeddings 135 from inputs. The embeddings 135 may represent embeddings, encodings, vectors, features, and / or the like, and each feature or word embedding 135 may include encoded data that represents and / or identifies one or more objects included in an input 130. In some embodiments, the feature or word embeddings 135 can be fed to the data clustering model 140. The data clustering model can be trained to utilize the feature or word embeddings 135 to execute one or more computer vision task(s) 180 (e.g., object detection, object classification, and / or instance segmentation functions) or one or more language model task(s) 185.

[0131] The data clustering model of the learning system(s) 125 can be configured to generate and output analysis information based on an analysis of the inputs 130. The analysis information for an input 130 can generally include any information or data associated with analyzing, interpreting, understanding, and / or classifying the images or other objects included in the inputs 130. In certain embodiments, the analysis information can include information or data representing the feature or word embeddings 135 that are extracted from the inputs 130. Additionally, or alternatively, the analysis information can include information or data that indicates the results of the learning system(s) 125 performed by the data clustering model. For example, the analysis informationmay include the predictions and / or results associated with performing computer vision task(s) 180 including object detection, object classification, and / or other computer vision functions. The analysis information also can include the predictions and / or results associated with performing language model task(s) 185. Language model task(s) 185 can include sentence completion, content searching, content interpretation, context dependent tasks, and / or sentiment analysis.

[0132] In certain embodiments, one or more training procedures may be executed to train the learning system(s) 125 to perform the computer vision task(s) 180 and / or the NLP task(s) 185 described in this disclosure. The training procedures can enable the data clustering model 140 to learn to optimize the clustering of inputs 130 received by the learning system 125. The specific procedures that are utilized to train the data clustering model 140 can vary. In some cases, one or more supervised training procedures, one or more unsupervised training procedures, and / or one or more semi-supervised training procedures may be applied to train the learning system.

[0133] As evidenced by the disclosure herein, the inventive techniques set forth in this disclosure are rooted in computer technologies that overcome existing problems in known learning system(s) 125, including problems dealing with computational efficiency and accuracy of attention mechanism used by computer vision and / or language model systems. The techniques described in this disclosure provide a technical solution (e.g., one that utilizes continuous learning and unsupervised training approaches) for overcoming the limitationsassociated with known techniques. This technology-based solution marks an improvement over existing capabilities and functionalities related to learning systems.

[0134] In many embodiments, the techniques described herein can be used continuously at a scale that cannot be reasonably performed using manual techniques or the human mind.

[0135] In a number of embodiments, the techniques described herein can solve a technical problem that arises only within the realm of computer networks, as data clustering models and learning systems do not exist outside the realm of computer networks.

[0136] In some embodiments, a computerized method for applying an attention- based congruency mechanism on a data clustering distribution can be implemented via execution of computing instructions stored on one or more non-transitory storage devices by one or more processing devices. The method can include receiving embeddings extracted from one or more inputs. The method also can include executing a data clustering model that creates a clustering distribution of the embeddings in a multi-dimensional space. The executing the data clustering method can include creating a list of anchor points for the multi-dimensional space from a first subset of the embeddings. The executing the data clustering model further can include mapping a second subset of the embeddings to data point indexes in the multi-dimensional space to create a plurality of clusters. The executing the data clustering model also can include executing an attention-based congruency function on the data pointindexes included in the clustering distribution. The executing the attention- based congruency function on the data point indexes included in the clustering distribution can include, for each of the data point indexes, identifying a nearest anchor point in the multi-dimensional space. The executing the attention-based congruency function on the data point indexes included in the clustering distribution further can include, for each of the data point indexes, adjusting a nearness of the corresponding data point index to the nearest anchor point in the clustering distribution. The method further can include executing one or more machine learning tasks using the clustering distribution.

[0137] In some embodiments, a method for generating and creating a clustering distribution can be implemented via execution of computing instructions configured to run at one or more processing devices and stored on non- transitory computer-readable media. The method can include (a) receiving, at a computing device, a batch of embeddings extracted from one or more inputs. The method also can include (b) determining data points in a high-dimensional space corresponding to the embeddings. The method further can include (c) creating an anchor point list for a clustering distribution that includes a first subset of the data points. The creating the anchor point list can include selecting one of the data points having a maximum norm to be an initial anchor point. The creating the anchor point list also can include selecting additional anchor points by identifying the data points that are near orthogonal to the initial anchor point in the high-dimensional space. As newly received data points are added to the clustering distribution, an anchor point replacement function canbe configured to update the anchor points included in the anchor point list. The anchor point replacement function can include, at least in part, determining a similarity between one or more of the newly received data points and a nearest anchor point in the high-dimensional space. The anchor point replacement function also can include, at least in part, replacing the nearest anchor point on the anchor point list with a newly received data point having a higher norm than the nearest anchor point. The method further can include (d) computing the data clustering distribution of a second subset of data points that are not included on the anchor point list. The computing the data clustering distribution of a second subset of data points can include mapping the data points included in the second subset of data points to the high-dimensional space. The computing the data clustering distribution of a second subset of data points also can include executing an attention-based congruency function that identifies a nearest anchor point in the high-dimensional space for each of the data points included in the second subset of data points, and moves each of the data points included in the second subset of data points closer to its nearest anchor point in the high-dimensional space. The method further can include (e) using the data clustering distribution to perform one or more machine-learning tasks.

[0138] Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. A computer-usable or computer-readable medium may include any apparatus that stores, communicates, propagates, or transports the program for use by orin connection with the instruction execution system, apparatus, or device. The medium can be a magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium. The medium may include a computer-readable storage medium, such as a semiconductor or solid-state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disk and an optical disk, etc.

[0139] A data processing system suitable for storing and / or executing program code may include at least one processor coupled directly or indirectly to memory elements through a system bus. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memories that provide temporary storage of at least some program code to reduce the number of times code is retrieved from bulk storage during execution. Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) may be coupled to the system either directly or through intervening I / O controllers.

[0140] Network adapters may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the currently available types of network adapters.

[0141] It should be recognized that any features and / or functionalities described for an embodiment in this application can be incorporated into any otherembodiment mentioned in this disclosure. Moreover, the embodiments described in this disclosure can be combined in various ways. Additionally, while the description herein may describe certain embodiments, features, or components as being implemented in software or hardware, it should be recognized that any embodiment, feature, or component that is described in the present application may be implemented in hardware, software, or a combination of the two. * * * * *

[0142] While various novel features of the invention have been shown, described, and pointed out as applied to particular embodiments thereof, it should be understood that various omissions and substitutions, and changes in the form and details of the systems and methods described and illustrated, may be made by those skilled in the art without departing from the spirit of the invention. Amongst other things, the steps in the methods may be carried out in different orders in many cases where such may be appropriate. Those skilled in the art will recognize, based on the above disclosure and an understanding of the teachings of the invention, that the particular hardware and devices that are part of the system described herein, and the general functionality provided by and incorporated therein, may vary in different embodiments of the invention. Accordingly, the descriptions of system components are for illustrative purposes to facilitate a full and complete understanding and appreciation of the various aspects and functionality of particular embodiments of the invention asrealized in system and method embodiments thereof. Those skilled in the art will appreciate that the invention can be practiced in other than the described embodiments, which are presented for purposes of illustration and not limitation. Variations, modifications, and other implementations of what is described herein may occur to those of ordinary skill in the art without departing from the spirit and scope of the present invention and its claims.

Claims

WHAT IS CLAIMED IS:

1. A computerized method for applying an attention-based congruency mechanism on a data clustering distribution, the computerized method implemented via execution of computing instructions stored on one or more non-transitory storage devices by one or more processing devices, the method comprising: receiving embeddings extracted from one or more inputs; executing a data clustering model that creates a clustering distribution of the embeddings in a multi-dimensional space, wherein executing the data clustering model includes: creating a list of anchor points for the multi-dimensional space from a first subset of the embeddings; mapping a second subset of the embeddings to data point indexes in the multi-dimensional space to create a plurality of clusters; and executing an attention-based congruency function on the data point indexes included in the clustering distribution by: for each of the data point indexes, identifying a nearest anchor point in the multi-dimensional space; and for each of the data point indexes, adjusting a nearness of the corresponding data point index to the nearest anchor point in the clustering distribution; and executing one or more machine learning tasks using the clustering distribution.

2. The computerized method of claim 1, wherein creating a list of anchor points for the multi-dimensional space includes: selecting a first anchor point based on a data point index determined to have a maximum norm; and selecting additional anchor points by identifying embeddings having data point indexes that are orthogonal or near orthogonal to the first anchor point.

3. The computerized method of claim 2, wherein the data clustering model includes an anchor point replacement function that is configured to: in response to determining that a first norm computed for a newly received embedding is greater than a second norm of an anchor point included in the list of anchor points, replacing the anchor point with the newly received embedding.

4. The computerized method of claim 1, wherein adjusting the nearness of the corresponding data point index to the nearest anchor point in the clustering distribution includes moving the corresponding data point index closer to the nearest anchor point in the clustering distribution.

5. The computerized method of claim 4, wherein the corresponding data point index is moved closer to the nearest anchor point in the clustering distribution using an objective function that increase a similarity of the data point index to the nearest anchor point.

6. The computerized method of claim 1, wherein the method further includes executing a continuous learning mechanism that continuously updates the list of anchor points as new batches of embeddings are processed.

7. The computerized method of claim 1, wherein the one or more machine learning tasks include one or more computer vision tasks executed by a computer vision system, and the one or more computer vision tasks include one or more of: a detection task, a recognition task, or a segmentation task.

8. The computerized method of claim 1, wherein the one or more machine learning tasks include one or more natural language processing (NLP) tasks executed by a language model system, and the one or more NLP tasks include one or more of: a as language modeling task, question-answering task, or natural language generation task.

9. A method for generating and creating a clustering distribution, the method implemented via execution of computing instructions configured to run at one or more processing devices and stored on non-transitory computer-readable media, the method comprising: (a) receiving, at a computing device, a batch of embeddings extracted from one or more inputs; (b) determining data points in a high-dimensional space corresponding to the embeddings;(c) creating an anchor point list for a clustering distribution that includes a first subset of the data points by: selecting one of the data points having a maximum norm to be an initial anchor point; and selecting additional anchor points by identifying the data points that are near orthogonal to the initial anchor point in the high-dimensional space; wherein, as newly received data points are added to the clustering distribution, an anchor point replacement function is configured to update the anchor points included in anchor point list, at least in part, by: determining a similarity between one or more of the newly received data points and a nearest anchor point in the high- dimensional space; and replacing the nearest anchor point on the anchor point list with a newly received data point having a higher norm than the nearest anchor point; (d) computing the data clustering distribution of a second subset of data points that are not included on the anchor point list by: mapping the data points included in the second subset of data points to the high-dimensional space; and executing an attention-based congruency function that identifies a nearest anchor point in the high-dimensional space for each of the data points included in the second subset of data points, and moves each of the datapoints included in the second subset of data points closer to its nearest anchor point in the high-dimensional space; and (e) using the data clustering distribution to perform one or more machine- learning tasks.

10. The method of claim 9, further comprising a continuous learning mechanism that updates the anchor point list for the clustering distribution based on newly received data points, and adjusts the data points in the clustering distribution based on updates to the anchor point list.

11. The method of claim 9, wherein moving each of the data points included in the second subset of data points closer to its nearest anchor point in the high-dimensional space comprises using an objective function that increases the similarity between the data points included in the second subset of data points and the corresponding nearest anchor points, and adjusting locations of the data points included in the second subset of data points based on the similarity adjustment.

Citation Information

Patent Citations

  • System and Method for Creating Timbres

    US20180342258A1

  • Natural language generation using pinned text and multiple discriminators

    US20190236139A1

  • Searching multidimensional indexes using associated clustering and dimension reduction information

    US6134541A

  • Method and apparatus for similarity retrieval from iterative refinement

    US7272593B1