Electronic device and method for controlling electronic device

Through multimodal prompt tuning technology, visual and text encoder generation layer-specific learnable prompt marks are solved, and the CZSL model lacks the ability to recognize combinations in an open world environment is achieved efficient model training and optimization.

CN120345010APending Publication Date: 2025-07-18SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085571.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-23
Filing Date
2023-12-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing Combined Zero Sample Learning (CZSL) models lack the ability to recognize unseen combinatorials in an open-world environment, and the fine-tuning of large pre-trained visual language models is expensive.

Method used

Using multimodal prompt tuning technology, a pre-trained visual language model is used to combine vision, first text and second text encoder to generate layer-specific learnable prompt marks for selecting attribute tags and object tags in the image to reduce the training needs for the entire model.

Benefits of technology

While maintaining resource requirements low, the model's recognition performance for unseen combinations is significantly improved, training costs are reduced, and the recognition ability of unseen combinations is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345010A_ABST
    Figure CN120345010A_ABST
Patent Text Reader

Abstract

A method includes obtaining an image, a set of attribute tags, and a set of object tags; and performing hint tuning of the pre-trained visual language model having first and second text encoders and a visual encoder. The model is trained during hint tuning to select an attribute tag and an object tag that match content contained in the image. Performing hint tuning includes, for each attribute tag-object tag pair: generating, using a first text encoder, an object text feature associated with an object tag; generating an attribute textual feature associated with the attribute tag using a second text encoder; and generating an image feature associated with the image using the visual encoder. Intermediate outputs from initial layers of the text encoders and the visual encoders are integrated to generate layer-specific learnable cue indicia that are appended to inputs of specified layers of the first and second text encoders and the visual encoders.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to machine learning systems and processes. More specifically, the present disclosure relates to prompt tuning for zero-shot compositional learning in machine learning systems. Background Art

[0002] Humans can generally understand new concepts by composing or combining parts of other previously learned concepts, which is often referred to as "compositional generalization ability". Compositional generalization ability is a fundamental feature that allows humans to understand new concepts by combining learned knowledge. Compositional zero-shot learning (CZSL) aims to simulate this human intelligence in a machine learning environment by enabling a machine learning model to learn a relatively small number of known compositions and generalize its recognition ability to unseen compositions. Summary of the Invention

[0003] Solution to the Problem

[0004] The present disclosure relates to prompt tuning for zero-shot compositional learning in machine learning systems.

[0005] In a first embodiment, a method includes: obtaining an image, a set of attribute labels, and a set of object labels; and performing prompt tuning of a pre-trained vision-language model having a first text encoder, a second text encoder, and a vision encoder. The pre-trained vision-language model is trained during the prompt tuning to select one of the attribute labels and one of the object labels that match the content included in the image. Performing the prompt tuning includes, for each of a plurality of attribute label-object label pairs: using the first text encoder to generate an object text feature associated with the object label of the attribute label-object label pair; using the second text encoder to generate an attribute text feature associated with the attribute label of the attribute label-object label pair; and using the vision encoder to generate an image feature associated with the image. Intermediate outputs from initial layers of the first text encoder, the second text encoder, and the vision encoder are integrated to generate layer-specific learnable prompt tokens that are appended to the input of a specified layer in the first text encoder, the second text encoder, and the vision encoder during the prompt tuning. In another embodiment, a non-transitory machine-readable medium includes instructions that, when executed, cause at least one processor to perform the method of the first embodiment.

[0006] In a second embodiment, an apparatus includes: at least one processing device configured to: obtain an image, a set of attribute tags, and a set of object tags; and perform prompt tuning of a pre-trained vision-language model having a first text encoder, a second text encoder, and a vision encoder. The pre-trained vision-language model is trained during the prompt tuning to select one of the attribute tags and one of the object tags that match the content included in the image. To perform the prompt tuning, the at least one processing device is configured to, for each of a plurality of attribute tag-object tag pairs: generate, using the first text encoder, an object text feature associated with the object tag of the attribute tag-object tag pair; generate, using the second text encoder, an attribute text feature associated with the attribute tag of the attribute tag-object tag pair; and generate, using the vision encoder, an image feature associated with the image. The at least one processing device is configured to integrate intermediate outputs from initial layers of the first text encoder, the second text encoder, and the vision encoder to generate layer-specific learnable prompt tokens that are appended to the input of a specified layer in the first text encoder, the second text encoder, and the vision encoder during the prompt tuning.

[0007] In a third embodiment, a method includes: obtaining an image of an object and obtaining a set of attribute tokens and a set of object tags. The method further includes: using a vision-language model having a first text encoder, a second text encoder, and a vision encoder to select one of the attribute tags and one of the object tags associated with the object. Selecting one of the attribute tags and one of the object tags associated with the object includes: generating, using the first text encoder, an object text feature associated with one of the object tags; generating, using the second text encoder, an attribute text feature associated with one of the attribute tags; and generating, using the vision encoder, an image feature associated with the image. Each of one or more layers in the first text encoder, one or more layers in the second text encoder, and one or more layers in the vision encoder is associated with a layer-specific multimodal shared prompt that is concatenated to the input of the layer. In another embodiment, an apparatus includes at least one processing device configured to perform the method of the third embodiment. In yet another embodiment, a non-transitory machine-readable medium includes instructions that, when executed, cause at least one processor to perform the method of the third embodiment.

[0008] Based on the following figures, description, and claims, other technical features may be apparent to those skilled in the art.

[0009] Before proceeding with the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "send," "receive," and "communicate" and their derivatives cover both direct and indirect communication. The terms "include" and "comprise" and their derivatives mean including but not limited to. The term "or" is inclusive and means and / or. The phrase "associated with" and its derivatives mean including, being included within, interconnecting with, containing, being contained within, connected to or coupled with, capable of communicating with, cooperating with, interlacing, juxtaposing, adjacent to, bound to or coupled with, having, having the property of, having a relationship to or with, etc.

[0010] In addition, the various functions described below can be implemented or supported by one or more computer programs, each formed from computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof, suitable for implementation in appropriate computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory. A "non-transitory" computer-readable medium excludes wired, wireless, optical, or other communication links that transmit transitory electrical or other signals. Non-transitory computer-readable media include media in which data can be stored permanently and media in which data can be stored and later rewritten, such as rewritable optical discs or erasable memory devices.

[0011] As used herein, terms and phrases such as "having", "may have", "including", or "may include" a feature (such as a number, a function, an operation, or a component such as a part) indicate the presence of the feature, without excluding the presence of other features. In addition, as used herein, the phrase "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may indicate (1) including at least one A, (2) including at least one B, or (3) including all of at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate user devices different from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, the first component may be represented as the second component, and vice versa.

[0012] It should be understood that when an element (such as a first element) is referred to as being "coupled" / "coupled to" or "connected" / "connected to" another element (such as a second element) (operatively or communicatively), it may be directly or via a third element connected or connected to the other element. In contrast, it will be understood that when an element (such as a first element) is referred to as being "directly coupled" / "directly coupled to" or "directly connected" / "directly connected to" another element (such as a second element), no other element (such as a third element) intervenes between the element and the other element.

[0013] As used herein, the phrase "configured (or set) to" may be used interchangeably with the phrases "suitable for", "capable of", "designed to", "adapted to", "manufactured to", or "able to", as the case may be. The phrase "configured (or set) to" does not essentially mean "specially designed in hardware". Instead, the phrase "configured to" may indicate that a device may perform an operation together with another device or part. For example, the phrase "a processor configured (or set) to perform A, B, and C" may mean a general-purpose processor (such as a CPU or an application processor) that can perform operations by executing one or more software programs stored in a memory device, or a dedicated processor (such as an embedded processor) for performing the operations.

[0014] The terms and phrases used herein are for describing only some embodiments of the present disclosure and are not intended to limit the scope of other embodiments of the present disclosure. It should be understood that, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural referents. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as is commonly understood by one of ordinary skill in the art to which the embodiments of the present disclosure pertain. It will be further understood that terms and phrases (such as those defined in common dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In some instances, the terms and phrases defined herein may be interpreted to exclude embodiments of the present disclosure.

[0015] Examples of an "electronic device" according to embodiments of the present disclosure may include at least one of a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, electronic accessories, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the electronic device include smart home appliances. Examples of the smart home appliances may include at least one of a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker or a speaker having an integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (such as XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camera, or an electronic photo frame. Other examples of the electronic device include at least one of the following: various medical devices (such as various portable medical measurement devices (e.g., a blood glucose measurement device, a heartbeat measurement device, or a body temperature measurement device), a magnetic resource angiography (MRA) device, a magnetic resource imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an in-vehicle infotainment device, marine electronic devices (such as a marine navigation device or a gyrocompass), avionics, a security device, an in-vehicle head unit, an industrial or domestic robot, an automated teller machine (ATM), a point of sale (POS) device, or an Internet of Things (IoT) device (such as a light bulb, various sensors, a water meter, an electricity meter, or a gas meter, a sprinkler, a fire alarm, a thermostat, a street lamp, an oven, a fitness device, a hot water tank, a heater, or a boiler). Other examples of the electronic device include at least a part of a piece of furniture or a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measurement devices (such as a device for measuring water, electricity, gas, or electromagnetic waves). Note that, according to various embodiments of the present disclosure, the electronic device may be one or a combination of the devices listed above. According to some embodiments of the present disclosure, the electronic device may be a flexible electronic device. The electronic device disclosed herein is not limited to the devices listed above and may include new electronic devices depending on technological development.

[0016] In the following description, various embodiments according to the present disclosure are described with reference to the accompanying drawings. As used herein, the term "user" may refer to a person using an electronic device or another device (such as an artificial intelligence electronic device).

[0017] Throughout this patent document, definitions may be provided for certain other words and phrases. Those of ordinary skill in the art should understand that, in many if not most cases, such definitions apply to both the prior and future use of the words and phrases so defined.

[0018] None of the descriptions in this application should be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the patent subject matter is defined only by the claims. Additionally, no claim is intended to invoke 35 U.S.C. § 112(f) unless the exact phrase "means for" is followed by a participle. The use of any other term within the claims, including but not limited to "mechanism", "module", "device", "unit", "component", "element", "member", "apparatus", "machine", "system", "processor", or "controller", is understood by the applicant to refer to a structure known to those skilled in the relevant art and is not intended to invoke 35 U.S.C. § 112(f). BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals represent like parts:

[0020] Figure 1 An example network configuration including an electronic device according to the present disclosure is shown;

[0021] Figure 2 An example architecture for prompt tuning to support zero-shot compositional learning in a machine learning system according to the present disclosure is shown;

[0022] Figures 3 to 6 An example use case of zero-shot compositional learning in a machine learning system according to the present disclosure is shown;

[0023] Figure 7 An example method for training a machine learning model to support zero-shot compositional learning according to the present disclosure is shown; and

[0024] Figure 8 An example method for using a machine learning model trained to perform zero-shot compositional learning according to the present disclosure is shown. DETAILED DESCRIPTION

[0025] The following discussion is described with reference to the accompanying drawings Figures 1 to 8and various embodiments of the present disclosure. However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalents or substitutions thereof also fall within the scope of the present disclosure.

[0026] As described above, humans can generally understand new concepts by combining or integrating parts of other previously learned concepts, which is commonly referred to as the "combinatorial generalization ability". The combinatorial generalization ability is a fundamental feature that allows humans to understand new concepts by combining learned knowledge. Combinatorial zero-shot learning (CZSL) aims to simulate this human intelligence in a machine learning environment by enabling a machine learning model to learn a relatively small number of known combinations and generalize its recognition ability to unseen combinations.

[0027] In the CZSL setting, each known combination can include a pair of labels that identify an object and an attribute of the object. For example, in an image of an old elephant, "old" is the attribute label and "elephant" is the object label. A CZSL model can attempt to identify unseen combinations by decomposing the known combinations contained in the training data, such as learning the new concept of "old truck" after learning the concepts of "old elephant" and "new truck" from the training data. However, this is a very challenging machine learning task. Among other reasons, different integrations of the same attribute label and object label can be associated with images having different shapes, colors, and textures. As a specific example, two instances of the same type of object can be fundamentally different in terms of their visual features (such as "raw chicken" and "sliced chicken"). Similarly, the same attribute may look very different in two different objects (such as "old truck" and "old elephant").

[0028] Existing CZSL models are typically limited to a closed output space where prior knowledge of unseen combinations is assumed during testing. As a result, when applying these CZSL models to more realistic settings, a significant performance degradation can be observed without the limitation of the output search space. Large pre-trained machine learning models have shown potential in providing "common sense" knowledge to numerous downstream vision and language tasks. Large pre-trained vision-language models (such as the Contrastive Language-Image Pretraining (CLIP) neural network) are good candidates for handling CZSL tasks because CZSL tasks are generally multimodal and involve text- and image-based inputs. However, although they are effective, fine-tuning these large pre-trained vision-language models is costly or prohibitively expensive due to their large scale.

[0029] The present disclosure provides various techniques for prompt tuning of zero-shot compositional learning in machine learning systems. As described in more detail below, an image, a set of attribute labels, and a set of object labels can be obtained, and prompt tuning of a pre-trained vision-language model can be performed. The image can include at least one object, the set of attribute labels can identify possible attributes of the object, and the set of object labels can identify possible types of the object. The pre-trained vision-language model can include a first text encoder, a second text encoder, and a vision encoder. The pre-trained vision-language model can be trained during prompt tuning to select one of the attribute labels and one of the object labels that match the content included in the image. Performing prompt tuning can include, for each of a plurality of attribute label-object label pairs, using the first text encoder to generate object text features associated with the object label of the attribute label-object label pair, using the second text encoder to generate attribute text features associated with the attribute label of the attribute label-object label pair, and using the vision encoder to generate image features associated with the image. Intermediate outputs from initial layers of the first text encoder, the second text encoder, and the vision encoder can be integrated to generate layer-specific learnable prompt tokens that are appended to the input of a specific layer in the first text encoder, the second text encoder, and the vision encoder during prompt tuning.

[0030] The vision-language model trained as a result of prompt tuning can be used in any suitable manner. For example, an image of an object can be obtained, and a set of attribute labels and a set of object labels can be obtained. A vision-language model including a first text encoder, a second text encoder, and a vision encoder can be used to select one of the attribute labels and one of the object labels associated with the object. Selecting one of the attribute labels and one of the object labels associated with the object can include using the first text encoder to generate object text features associated with each of the object labels, using the second text encoder to generate attribute text features associated with each of the attribute labels, and using the vision encoder to generate image features associated with the image. One or more layers in the first text encoder, one or more layers in the second text encoder, and one or more layers in the vision encoder can each be associated with a layer-specific multimodal shared prompt that is concatenated to the input of that layer.

[0031] In this way, multimodal "prompt tuning" can be used to adapt a pre-trained vision-language model for one or more downstream tasks, such as one or more CZSL tasks. CZSL tasks typically involve using rich knowledge to identify unseen combinations, and tuning a large-scale pre-trained vision-language model can help achieve high identification performance and reduce or minimize training costs. Additionally, the number of parameters in some large pre-trained vision-language models can be huge, and performing full model fine-tuning can be costly or prohibitively expensive. The techniques described provide an effective mechanism for fine-tuning large pre-trained vision-language models using multimodal prompts, which can achieve good performance while maintaining low resource requirements for training.

[0032] Note that while some of the embodiments discussed below are described in the context of use in consumer electronic devices such as smart phones, ovens, microwave ovens, refrigerators, washing machines, dryers, and steam cabinets, these are merely examples. It should be understood that the principles of the present disclosure can be implemented in any number of other suitable contexts and can use any suitable one or more devices. Also note that while some of the embodiments discussed below are described assuming the training of a machine learning model on one device, such as a server, for deployment to one or more other devices, such as one or more consumer electronic devices, this is also merely an example. It should be understood that the principles of the present disclosure can be implemented using any number of devices, including a single device that both trains and uses the machine learning model. Generally, the present disclosure is not limited to use with any particular type of device.

[0033] Figure 1 An example network configuration 100 including an electronic device in accordance with the present disclosure is shown. Figure 1 The illustrated embodiment of the network configuration 100 is for illustrative purposes only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.

[0034] In accordance with an embodiment of the present disclosure, an electronic device 101 is included in the network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or a sensor 180. In some embodiments, the electronic device 101 may exclude at least one of these components or may add at least one other component. The bus 110 includes circuitry for connecting the components 120 - 180 to each other and for transmitting communications, such as control messages and / or data, between the components.

[0035] The processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In some embodiments, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), or a graphics processing unit (GPU). The processor 120 is capable of performing control and / or executing operations or data processing related to communication or other functions on at least one of the other components of the electronic device 101. As described in more detail below, the processor 120 may perform various operations related to prompt tuning using a machine learning model for zero-shot composition learning and / or performing zero-shot composition learning using a trained machine learning model.

[0036] The memory 130 may include volatile and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one other component of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or a program 140. The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or “app”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).

[0037] The kernel 141 may control or manage system resources (such as the bus 110, the processor 120, or the memory 130) for executing operations or functions implemented in other programs, such as middleware 143, API 145, or app 147. The kernel 141 provides an interface that allows middleware 143, API 145, or app 147 to access various components of the electronic device 101 to control or manage system resources. The app 147 may support various functions related to the training and / or use of a machine learning model. These functions may be performed by a single app or multiple apps, with each app performing one or more of these functions. For example, the middleware 143 may act as a repeater to allow the API 145 or app 147 to communicate data with the kernel 141. Multiple apps 147 may be provided. The middleware 143 is capable of controlling work requests received from the app 147, such as by assigning priorities for using system resources of the electronic device 101, such as the bus 110, the processor 120, or the memory 130, to at least one of the multiple apps 147. The API 145 is an interface that allows the app 147 to control functions provided from the kernel 141 or middleware 143. For example, the API 145 includes at least one interface or function (such as a command) for archive control, window control, image processing, or text control.

[0038] The I / O interface 150 serves as an interface that can transfer, for example, commands or data input from a user or other external devices to other components of the electronic device 101. The I / O interface 150 can also output commands or data received from other components of the electronic device 101 to the user or other external devices.

[0039] The display 160 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a microelectromechanical systems (MEMS) display, or an electronic paper display. The display 160 can also be a depth perception display, such as a multi-focus display. The display 160 is capable of displaying various contents (such as text, images, videos, icons, or symbols) to the user. The display 160 can include a touch screen and can receive, for example, touch, gesture, proximity, or hover inputs using an electronic pen or a body part of the user.

[0040] For example, the communication interface 170 is capable of establishing communication between the electronic device 101 and external electronic devices (such as the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can be connected to the network 162 or 164 through wireless or wired communication to communicate with external electronic devices. The communication interface 170 can be a wired or wireless transceiver or any other component for transmitting and receiving signals.

[0041] Wireless communication can use, for example, at least one of WiFi, Long Term Evolution (LTE), Advanced Long Term Evolution (LTE-A), Fifth Generation Wireless System (5G), millimeter wave or 60 GHz wireless communication, Wireless USB, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM) as a communication protocol. Wired connections can include, for example, at least one of Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Recommended Standard 232 (RS-232), or Plain Old Telephone Service (POTS). The network 162 or 164 includes at least one communication network, such as a computer network (such as a Local Area Network (LAN) or a Wide Area Network (WAN)), the Internet, or a telephone network.

[0042] The electronic device 101 also includes one or more sensors 180, which can measure physical quantities or detect the activation state of the electronic device 101 and convert the measured or detected information into an electrical signal. For example, the one or more sensors 180 may include one or more cameras or other imaging sensors for capturing images of a scene. The sensors 180 may also include one or more buttons for touch input, a gesture sensor, a gyroscope or gyroscopic sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (such as a red, green, blue (RGB) sensor), a biophysical sensor, a temperature sensor, a humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an infrared (IR) sensor, an ultrasonic sensor, an iris sensor or a fingerprint sensor. The sensors 180 may also include an inertial measurement unit, which may include one or more accelerometers, gyroscopes and other components. Additionally, the sensors 180 may include control circuitry for controlling at least one of the sensors included herein. Any one of these sensors 180 may be located within the electronic device 101.

[0043] In some embodiments, the first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device installable on an electronic device (such as an HMD). When the electronic device 101 is installed in the electronic device 102 (such as an HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network. The electronic device 101 may also be an augmented reality wearable device including one or more imaging sensors, such as glasses.

[0044] The first external electronic device 102, the second external electronic device 104, and the server 106 can each be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. Additionally, according to certain embodiments of the present disclosure, all or some of the operations performed on the electronic device 101 can be performed on another or more other electronic devices (such as the electronic devices 102 and 104 or the server 106). Further, according to certain embodiments of the present disclosure, when the electronic device 101 is supposed to automatically or upon request perform some function or service, the electronic device 101 can request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some functions associated therewith, rather than performing the function or service itself, or additionally performing the function or service. The other electronic devices (such as the electronic devices 102 and 104 or the server 106) are capable of performing the requested function or additional functions and transmitting the result of the execution to the electronic device 101. The electronic device 101 can provide the requested function or service by processing the received result as it is or additionally. For this purpose, for example, cloud computing, distributed computing, or client-server computing technologies can be used. Although Figure 1 FIG. shows that the electronic device 101 includes a communication interface 170 that communicates with the external electronic device 104 or the server 106 via the network 162 or 164, but according to some embodiments of the present disclosure, the electronic device 101 can operate independently without a separate communication function.

[0045] The server 106 can include the same or similar components 110 - 180 (or a suitable subset thereof) as the electronic device 101. The server 106 can support driving the electronic device 101 by performing at least one of the operations (or functions) implemented on the electronic device 101. For example, the server 106 can include a processing module or a processor that can support the processor 120 implemented in the electronic device 101. As described in more detail below, the server 106 can perform various operations related to prompt tuning using a machine learning model for zero-shot combinatorial learning and / or performing zero-shot combinatorial learning using a trained machine learning model.

[0046] Note that the electronic device 101 here can take various forms according to the implementation. For example, in certain cases discussed in this patent document, it can be assumed that the electronic device 101 represents a portable consumer device, such as a smart phone, a tablet computer, or a virtual reality (VR) / augmented reality (AR) / extended reality (XR) headset. However, the electronic device 101 can take any other suitable form, such as an oven, a microwave oven, a refrigerator, a washing machine, a dryer, or a steam cabinet. Generally, the present disclosure is not limited to use with any specific type of device.

[0047] Although Figure 1 an example of a network configuration 100 including an electronic device 101 is shown, various changes can be made to Figure 1 it. For example, the network configuration 100 can include any number of each component in any suitable arrangement. Generally, computing and communication systems have a wide variety of configurations, and Figure 1 the scope of the present disclosure is not limited to any specific configuration. Additionally, although Figure 1 an operating environment in which various features disclosed in this patent document can be used is shown, these features can be used in any other suitable system.

[0048] Figure 2 An example architecture 200 for prompting tuning that supports zero-shot compositional learning in a machine learning system according to the present disclosure is shown. For ease of explanation, Figure 2 the illustrated architecture 200 is described as being used by the server 106 in the Figure 1 illustrated network configuration 100 to participate in prompting tuning of a large pre-trained vision-language model. However, the architecture 200 can be used by any other suitable device (such as the electronic device 101) and in any other suitable system, and the architecture 200 can be used to train any other suitable machine learning model.

[0049] As Figure 2 shown, the architecture 200 is used to support multimodal prompting tuning applied to a pre-trained vision-language model, which allows the pre-trained vision-language model to be effectively applied to one or more CZSL tasks. The pre-trained vision-language model here can represent any suitable vision-language model, such as a CLIP model, a Context Optimization (CoOp) model, or a Conditional Context Optimization (CoCoOp) model. As Figure 2 shown, the architecture 200 uses a machine learning model architecture based on a three-branch transformer to implement the pre-trained vision-language model, which can undergo multimodal prompting tuning to align the pre-trained vision-language model with at least one specific CZSL task. In this example, branch 202 of the architecture 200 implements a vision encoder, branch 204 of the architecture 200 implements a first text encoder, and branch 206 of the architecture 200 implements a second text encoder.

[0050] As shown here, the branch 202 of the architecture 200 receives the image 208 as an input. The image 208 can optionally be divided into multiple patches 210a, which can be processed in parallel in some cases. As described below, the visual patch prompt 210b can also optionally be used to improve the generalization ability of multi-modal prompt tuning. The visual encoder implemented by the branch 202 generally operates to extract visual features 212 from the image 208. For example, the visual encoder implemented by the branch 202 can be trained to learn which visual features 212 of the image 208 are relevant to a given task, and the visual encoder can extract those visual features 212 from each image 208 received by the visual encoder. The resulting visual features 212 are processed using a plurality of transformer encoder layers 214a - 214n, and the transformer encoder layers 214a - 214n generate embeddings of the visual features 212 within the associated feature space. Each transformer encoder layer 214a - 214n is associated with a set of learnable parameters 216, such as weights applied to the vectors processed by the transformer encoder layer 214a - 214n. As described below, one or more initial transformer encoder layers 214a - 214k in the visual encoder also receive one or more layer-specific shared prompts 218.

[0051] The branch 204 of the architecture 200 receives an input 220 that includes or identifies a set of object labels 222. The object labels 222 can represent a set of known object types, which can be generated or otherwise obtained in any suitable manner. For example, the set of object labels 222 can be retrieved or derived from a set of text categories in a knowledge base, which may have been previously learned and updated during an optimization process. The first text encoder implemented by the branch 204 generally operates to extract object text features 224 from the set of object labels 222. For example, the first text encoder implemented by the branch 204 can be trained to learn which object text features 224 are relevant to a given task, and the first text encoder can extract those object text features 224 from the set of object labels 222 received by the first text encoder. The resulting object text features 224 are processed using a plurality of transformer encoder layers 226a - 226n, and the transformer encoder layers 226a - 226n generate embeddings of the object text features 224 within the associated feature space. Each transformer encoder layer 226a - 226n is associated with a set of learnable parameters 228, such as weights applied to the vectors processed by the transformer encoder layer 226a - 226n. As described below, one or more initial transformer encoder layers 226a - 226k in the first text encoder also receive one or more layer-specific shared prompts 230.

[0052] The branch 206 of the architecture 200 receives an input 232 that includes or identifies a set of attribute labels 234. The attribute labels 234 can represent a set of known attribute types, which can be generated or otherwise obtained in any suitable manner. For example, the set of attribute labels 234 can be retrieved or derived from a set of text categories in a knowledge base, which may have been previously learned and updated during an optimization process. The second text encoder implemented by the branch 206 generally operates to extract attribute text features 236 from the set of attribute labels 234. For example, the second text encoder implemented by the branch 206 can be trained to learn which attribute text features 236 are relevant to a given task, and the second text encoder can extract those attribute text features 236 from the set of attribute labels 234 received by the second text encoder. The resulting attribute text features 236 are processed using a plurality of transformer encoder layers 238a - 238n, which generate embeddings of the attribute text features 236 within an associated feature space. Each transformer encoder layer 238a - 238n is associated with a set of learnable parameters 240, such as weights applied to vectors processed by the transformer encoder layer 238a - 238n. As described below, one or more initial transformer encoder layers 238a - 238k in the second text encoder also receive one or more layer - specific shared cues 242.

[0053] In some embodiments, an object score 244 can be generated for each object label 222, such as by identifying the similarity between the embedding of the visual features 212 generated by the final transformer encoder layer 214n of the visual encoder (branch 202) and the embedding of the object text features 224 associated with the object label 222 generated by the final transformer encoder layer 226n of the first text encoder (branch 204). The object label 222 associated with the highest object score can be selected as the object label 222 representing the type of the object included in the image 208. Similarly, an attribute score 246 can be generated for each attribute label 234, such as by identifying the similarity between the embedding of the visual features 212 generated by the final transformer encoder layer 214n of the visual encoder (branch 202) and the embedding of the attribute text features 236 associated with the attribute label 234 generated by the final transformer encoder layer 238n of the second text encoder (branch 206). The attribute label 234 associated with the highest attribute score can be selected as the attribute label 234 representing the attribute of the object included in the image 208. The selected object label 222 and the selected attribute label 234 can be integrated or otherwise used to generate the output 248 of the architecture 200.

[0054] In architecture 200, various branches 202, 204, 206 support the use of prompt tuning, which typically involves appending additional data to the input provided to one or more initial transformer encoder layers 214a - 214k, 226a - 226k, 238a - 238k in each of the branches 202, 204, 206. The additional data can be defined based on the downstream task to be performed, such that a large pre-trained vision-language model can better understand the downstream task. Examples of downstream tasks can include image-text retrieval, visual question answering, or visual grounding. The use of prompt tuning allows the pre-trained vision-language model to achieve substantial performance improvements without requiring additional training of the entire pre-trained vision-language model. As can be seen in Figure 2 Instead of using single-modal prompts for each of the branches 202, 204, and 206, shared prompts 218, 230, and 242 can be used across the visual and text encoders, which support multi-modal prompt tuning that can bridge the gap between different visual and language modalities. Again, visual chunk prompts 210b can be added to further improve the generalization ability of multi-modal prompt tuning.

[0055] The following now provides an explanation of how to use architecture 200 to support multi-modal prompt tuning in some embodiments. Note that this explanation pertains to a specific implementation of architecture 200, and other implementations of architecture 200 can be used. One goal of combined zero-shot learning can be to identify combinations of multiple concepts (such as objects and attributes) contained in a given image 208. For example, each sample in the training dataset may contain an image sample and a composition , where is an element in the image space , and is a composite label that includes an attribute label and an object label . Given a set 222 of object labels and a set 234 of attribute labels , the complete combination set can be constructed as . All combinations that appear in the training dataset form the set of "seen" combinations, which is a subset of the complete combination set . Combinations that do not appear in the training dataset form the set of "unseen" combinations, which is a subset of the complete combination set and may be larger than the set A much larger subset. Depending on the output space of the predictions to be made using the machine learning model, the difficulty of various CZSL tasks can vary widely. For example, given a new image , the prediction for the image x generated by the machine learning model can be expressed as , and the various learning scenarios can be defined as follows:

[0056] Supervised learning: ;

[0057] Zero-shot learning: ;

[0058] Generalized zero-shot learning: ; and

[0059] Open-world zero-shot learning: .

[0060] In the standard zero-shot learning setting, only unseen combinations are predicted during testing. However, recent CZSL work has considered generalization scenarios where test samples can come from the set of seen combinations or the set of unseen combinations . This is more challenging compared to standard zero-shot learning because the machine learning model will naturally be biased towards the seen combinations. The most challenging case is open-world combination zero-shot learning (OW-CZSL), where test samples can be drawn from the entire set of combinations . In this case, the output space is so large that it is almost impossible to generalize from a small number of seen combinations (meaning: ). The most challenging OW-CZSL scenario where there are no prior assumptions about the set of unseen combinations during testing is discussed here.

[0061] Multimodal Prompt Tuning (MMPT) performed using architecture 200 involves the use of both visual and text prompts. More specifically, MMPT involves using three branches 202, 204, and 206, one for vision and two for object and attribute text. In some embodiments, a Vision Transformer (ViT) can be used as the backbone in branch 202, and both branches 204 and 206 can use a language transformer as the encoder. To enable multimodal prompt tuning, shared prompts 218, 230, and 242 are introduced into the respective input spaces at one or more initial layers of each of branches 202, 204, and 206. Here, for the i-th transformer encoder layer 214a - 214k, 226a - 226k, 238a - 238k of each of branches 202, 204, 206, the input learnable prompt is represented as , which is of dimension The vectors in the visual encoder. The image embeddings have dimensions represented as (which can have a value of 768 in some cases), and the word embeddings in each object and attribute text encoder can have dimensions represented as (which can have a value of 512 in some cases). The mutual cues can be shared across multiple modalities and can have a length of 6 in some cases. To share the mutual cues across multiple modalities, the projection function is used to map the shared cues to the learnable parameters in the visual encoder (branch 202). Similarly, the projection functions and are used to map the shared prompts to the learnable parameters in the object and attribute text encoders (branches 204 and 206) respectively.

[0062] In the input layer of the visual encoder, the j-th patch 210a of the image can be represented as and is embedded into a -dimensional vector. Additionally, a visual patch cue 210b can be added, which can represent a learnable vector parameterized by . In some cases, the visual patch cue 210b can be initialized using random sizes and / or random positions within the image . As a specific example, the image can have a resolution of 224 pixels by 224 pixels, and the visual patch cue 210b can have a resolution of 16 pixels by 16 pixels (although these values are only examples). One or more initial embeddings generated by the visual encoder can be represented as follows.

[0063] (1)

[0064] Here, represents the number of patches 210a in the image 208. The learnable shared cue can be injected into the input space of the i-th Transformer encoder layer of branch 202 together with the learned embeddings from the previous layer and the hidden features encoding the learnable attribute tokens and object tokens. This process can be represented as follows.

[0065] (2)

[0066] Here, represents the concatenation operation on multiple vectors, Represents the i-th transformer encoder layer in branch 202. Note that only the first k transformer encoder layers 214a - 214k in branch 202 have layer-specific shared cues 218. Each subsequent transformer encoder layer in branch 202 (starting from transformer encoder layer k + 1) can process one or more embeddings and one or more cues from the previous layer, which can be expressed as follows.

[0067] (3)

[0068] Here, represents the total number of transformer encoder layers 214a - 214n in branch 202, and is equal to k. Since starts at , the initial value of is equal to

[0069] As described above, the shared cue can project the text encoders (branches 204 and 206) using the projection functions and . The initial transformer encoder layer 238a of branch 206 generates word embeddings from the text. For each subsequent transformer encoder layer in branch 206, the input includes the layer-specific shared cue , the fixed embedding and the attribute tokens from the previous layer. This can be expressed as follows.

[0070] (4)

[0071] Similar to the operation of branch 202 shown in equations (2) and (3) above, after the k-th transformer encoder layer 238k, branch 206 can reuse one or more cues learned from the previous layer as its input cues, as well as one or more embeddings from the previous layer. This can be expressed as follows.

[0072] (5)

[0073] Here, represents the total number of transformer encoder layers 238a - 238n in branch 206, and A hint indicating direct forwarding. In a particular embodiment, branch 206 may include a total of 12 transformer encoder layers 238a - 238n, and 9 of the transformer encoder layers 238a - 238k may receive layer - specific shared hints 242 (although these values are only examples).

[0074] Branch 204 may operate in a similar manner. For example, the initial transformer encoder layer 226a of branch 204 generates word embeddings from the text . For each subsequent transformer encoder layer in branch 204 , the input includes layer - specific hints , fixed embeddings and object tokens from the previous layer This can be represented as follows.

[0075] (6)

[0076] After the k - th transformer encoder layer 226k in branch 204, branch 204 may reuse one or more hints learned from the previous layer as its input hints, as well as one or more embeddings from the previous layer. This can be represented as follows.

[0077] (7)

[0078] Here, represents the total number of transformer encoder layers 226a - 226n in branch 204, and represents the hint for direct forwarding. In a particular embodiment, branch 204 may include a total of 12 transformer encoder layers 226a - 226n, and 9 of the transformer encoder layers 226a - 226k may receive layer - specific shared hints 230 (although these values are only examples).

[0079] The outputs of the final transformer encoder layers 214n and 226n in branches 202 and 204 can be integrated as described above to produce the object scores 244 of the object labels 222. Additionally, the outputs of the final transformer encoder layers 214n and 238n in branches 202 and 206 can be integrated as described above to produce the attribute scores 246 of the attribute labels 234. In some cases, the object scores 244 and the attribute scores 246 represent the probabilities that a single object label 222 and an attribute label 234 respectively exist within the image 208. As a specific example, take the prediction of which attribute label 234 should be selected as existing in image x. For a particular attribute , its probability on the image can be determined as follows.

[0080] (8)

[0081] Here, represents the attribute score 246 of the specific attribute label 234, and represents the cosine similarity. In addition, and represent the projection function, and represents the output vector from the final transformer encoder layer 214n in the visual encoder. In addition, represents the attribute token from the final transformer encoder layer 238n of the second text encoder, and represents the fixed temperature parameter. Similar calculations can be used to determine the object score 244 for each object label 222, such as in the following manner.

[0082] (9)

[0083] Here, represents the object score 244 of the specific object label 222, and represents the object token from the final transformer encoder layer 226n of the first text encoder.

[0084] In this way, during multi-modal prompt tuning, various images 208 containing objects associated with a subset of all possible attribute label-object label pairs can be used to train a model with a pre-trained vision-language backbone. Here, multi-modal prompt tuning is performed by (among other things) obtaining intermediate outputs from one or more initial transformer encoder layers 214a - 214k of the visual encoder, one or more initial transformer encoder layers 226a - 226k of the first text encoder, and one or more initial transformer encoder layers 238a - 238k of the second text encoder, and integrating these intermediate outputs to create layer-specific learnable prompt tokens (representing or based on the shared prompts 218, 230, and 242). These layer-specific learnable prompt tokens can be appended to the inputs of subsequent layers in each of the visual encoder, the first text encoder, and the second text encoder during prompt tuning. The visual encoder here can also use the learnable visual chunking prompt 210b to process each image 208, which can be added to the image 208 and moved within the image 208 during prompt tuning. After multi-modal prompt tuning is completed, the vision-language model can be configured to evaluate additional images 208 in a zero-shot manner to identify objects within the additional images 208, including those associated with attribute label-object label pairs not seen during multi-modal prompt tuning.

[0085] Thus, the architecture 200 can use shared multi-modal cues, which can help facilitate better model guidance. For example, the architecture 200 supports the use of cues from both visual and text modalities. For the visual modality, visual cues can be adopted in the form of visual chunk cues 210b in the image 208, which can help provide guidance for specific local regions within each image 208 that contain the most distinguishable visual features for object attribute understanding. Thus, for example, the architecture 200 can perform cueing in both pixel space and the fused embedding space to guide the pre-trained vision-language model and better understand where the focus should be placed within the image 208. For the text modality, object and attribute labels (which can represent object and attribute category names) can be concatenated at the initial stage in the branches 202, 204, 206 and can generally be updated to pseudo-soft cues during training to provide semantic meanings aligned with the perceived visual features. The cues from the visual and text modalities can be fused and re-injected into the corresponding feature representation learning networks (branches 202, 204, and 206) for improved cross-modal alignment. Through this multi-modal shared cueing process, better guidance can be provided to the pre-trained vision-language model to find optimized object attribute labels.

[0086] In addition, the architecture 200 supports progressive depth cueing, which enables efficient and parameter-efficient model tuning. That is, the described technique can be used to inject additional learnable tokens into each of the various transformer encoder layers 214a - 214k, 226a - 226k, 238a - 238k in the branches 202, 204, 206. The remaining parameters in the pre-trained vision-language model can remain unchanged and can be not updated during cue tuning. In the three branches 202, 204, 206, the intermediate outputs from the earlier layers in the branches 202, 204, 206 can be fused and used as inputs to the subsequent layers. The fusion of visual and language features can be done in this progressive manner by learning from the training data, such that the subsequent layers provide a more refined description of the cross-modal feature representation, thereby enhancing the model performance when identifying unseen combinations. On the other hand, although cueing can be performed in each transformer encoder layer 214a - 214n, 226a - 226n, 238a - 238n, only one token can be tuned at a time during this process, thus allowing a low learning cost and resulting in a novel parameter-efficient model tuning process that maintains high performance.

[0087] In addition, the architecture 200 supports the use of a three-branch model and the use of multi-task learning strategies, which implement open-world exploration for improved or optimal object property identification. For example, in each of the three branches 202, 204, 206 of the three-branch framework, clues from other branches can be used to assist in feature understanding in that branch. In addition, the final object label prediction and the final attribute label prediction can be independent of each other, and each possible object label-attribute label combination can be considered based on the given input image 208. This gives the architecture 200 the ability to discover rarely seen object-attribute combinations, thus enhancing its performance for difficult situations.

[0088] Note that in the architecture 200, branches 204 and 206 can be regarded as two separate machine learning classifiers, where one performs classification to select one of the object labels 222, and the other performs classification to select one of the attribute labels 234 of the image 208. For example, branch 204 can be used to select the object label 222 with the highest object score 244, and branch 206 can be used to select the attribute label 234 with the highest attribute score 246. In some embodiments, a loss function that integrates the losses associated with the machine learning classifiers can be used to jointly train the two machine learning classifiers. As a specific example, the following loss function can be used during prompt tuning.

[0089] (10)

[0090] Here, a refers to the total loss, represents the loss associated with object label classification, and represents the loss associated with attribute label classification. As shown in equation (10), during prompt tuning, the cross-entropy losses of the attribute and object branches 204 and 206 can be minimized.

[0091] After prompt tuning is completed and the machine learning model is put into use for inference, additional images 208 can be received and processed to identify the object label-attribute label pairs of the objects contained in the additional images 208. During inference, all possible combinations (meaning all possible object label-attribute label pairs) can be considered, and the object label-attribute label pair with the highest score can be used as the prediction for any given additional image 208. In some embodiments, the operations during inference can be expressed as follows.

[0092] (11)

[0093] Here, represents using the visual encoder from the given image The visual features 212 extracted therefrom. For OW-CZSL, the output search space is the full combination set , where most of the combinations are not seen in the training data used to train the machine learning model.

[0094] Note that it is often assumed above that when processing the image 208 and generating predictions of the image content, the complete set of object labels 222 and the complete set of attribute tags 234 are considered. However, this may not necessarily be the case. For example, a user of the electronic device 101 can set user preferences that indicate that some (but not all) of the object labels 222 and / or some (but not all) of the attribute labels 234 should be used when identifying the content in the image 208. As a specific example, the user can be allowed to customize his or her preferences regarding the main attribute categories that the user wants the machine learning model to identify, such as when the user prefers the model to identify the material, state, and texture attributes of the object while ignoring the color and pose attributes of the object. This can help save time when generating classification results and make the entire system clearer and more concise. This is an example of how the proposed machine learning model can be extended to identify attributes at a customized granularity level to improve the user experience.

[0095] It should be noted that Figure 2 the functions shown or described above can be implemented in any suitable manner in the electronic devices 101, 102, 104, the server 106, or other devices. For example, in some embodiments, one or more software applications or other software instructions executed by the processor 120 of the electronic devices 101, 102, 104, the server 106, or other devices can be used to implement or support Figure 2 at least some of the functions shown or described above. In other embodiments, dedicated hardware components can be used to implement or support Figure 2 at least some of the functions shown or described above. Generally, any suitable hardware or any suitable combination of hardware and software / firmware instructions can be used to perform Figure 2 the functions shown or described above. Moreover, Figure 2 the functions shown or described above can be performed by a single device or by multiple devices.

[0096] Although Figure 2 an example of the architecture 200 for prompt tuning that supports zero-shot combination learning in a machine learning system is shown, various changes can be made to Figure 2 it. For example, the various components or functions shown in Figure 2 can be combined, further subdivided, rearranged, copied, or omitted, and additional components or functions can be added according to specific needs. Moreover, although in Figure 2It is assumed in [description] that each of the branches 202, 204, 206 includes the same number (n) of transformer encoder layers, but this need not be the case, and each of the branches 202, 204, 206 may include any suitable number of transformer encoder layers. Additionally, while in [description] Figure 2 it is assumed that each of the branches 202, 204, 206 includes the same number (k) of initial transformer encoder layers that receive a shared prompt, this need not be the case, and each of the branches 202, 204, 206 may include any suitable number of transformer encoder layers that receive a shared prompt.

[0097] Additionally, while shown as having three branches 202, 204, 206, the architecture 200 may include more than three branches, such as by including one or more additional branches that consider one or more additional modalities. As a specific example, another branch for processing audio signals may be included in the architecture 200, where the architecture 200 may be used to fine-tune a pre-trained vision-language model to integrate information from the visual, text, and audio branches in order to understand more fine-grained details of objects and attributes. For example, by deploying the pre-trained vision-language model in a smart microwave oven, the pre-trained vision-language model may be able to identify fully cooked popcorn in the microwave oven based on inputs that include an image and a snippet of an audio recording (such as the audio of a "popping" sound) (and potentially automatically alert the user of the food status).

[0098] Figures 3 to 6 An example use case of zero-shot compositional learning in a machine learning system according to the present disclosure is shown. For ease of explanation, Figures 3 to 6 the use case shown in [description] is described as an example of how the architecture 200 may be used after prompt tuning has been completed. However, the architecture 200 may be used in any other suitable manner.

[0099] As Figure 3As shown, a food cooking device such as a microwave oven 302 or a conventional oven 304 can use the architecture 200 (or be used in combination with the architecture 200) to estimate the food cooking state and perform one or more actions based on the food cooking state. For example, one or more cameras or other imaging sensors 180 can be used inside the microwave oven 302 or the conventional oven 304 or otherwise associated with the microwave oven 302 or the conventional oven 304 to capture images of the food being cooked or to be cooked. The architecture 200 can be used to select object tags 222 and attribute tags 234 associated with the current state of a particular type of food. For example, the selected object tag 222 can identify the particular type of food being cooked, and the attribute tag 234 can identify the current cooking state of the food. The architecture 200 can be trained using various basic combinations such as "raw potato", "thawed pizza", and "baked chicken", and the architecture 200 can automatically understand novel concepts about the fine-grained state of food such as "baked bread", "thawed chicken", or "baked pizza". This enables the architecture 200 to perform dynamic identification 306 of food categories and states. Thereby, the processor 120 can execute or initiate the execution of one or more actions, such as dynamically controlling the cooking of the food (e.g., changing the remaining or total cooking time, adjusting the temperature, or ending the cooking). In this example, the processor 120 can automatically adjust the power setting 308 of the microwave oven 302 or the conventional oven 304. The processor 120 can also or alternatively control at least one notification of the cooking state, such as by updating a cooking progress bar 310 to indicate how well the food is cooked at the current time. Note that while the microwave oven 302 and the conventional oven 304 are shown here, any other suitable cooking device can be used, such as an air fryer, a bread maker, an oven, or other devices. This type of method can significantly enhance the user experience of the cooking device.

[0100] As Figure 4As shown, a food storage device such as refrigerator 402 can use architecture 200 (or be used in conjunction with architecture 200) to identify stored food and perform one or more actions based on the identified food. For example, one or more cameras or other imaging sensors 180 can be used within refrigerator 402 or otherwise associated with refrigerator 402 to capture images of the food within refrigerator 402. Architecture 200 can be used to select object tags 222 and attribute tags 234 associated with each identified food item. For example, the object tag - attribute tag integration selected for each food item can identify the type of food and its current state. This enables architecture 200 to perform dynamic identification 404 of the available food items. Thereby, processor 120 can generate one or more recipes 406 that can be presented to the user, such as when presented on the display of refrigerator 402 or when sent to a mobile device used by the user. Recipes 406 can include various ingredients, at least some of which are available within refrigerator 402. Note that using only object tags here may be insufficient because certain types of food items may be present in refrigerator 402 while having attributes that do not allow the food item to be used in certain recipes. This type of approach can improve cooking efficiency and reduce food waste.

[0101] As Figure 5 As shown, a clothing cleaning device such as washer / dryer 502 (individually or in combination) or steam cabinet 504 can use architecture 200 (or be used in conjunction with architecture 200) to identify a specific clothing item to be cleaned and perform one or more actions based on the identified clothing item. For example, one or more cameras or other imaging sensors 180 can be used within washer / dryer 502 or steam cabinet 504 or otherwise associated with washer / dryer 502 or steam cabinet 504 to capture images of the clothing item to be cleaned. Architecture 200 can be used to select object tags 222 and attribute tags 234 associated with each identified clothing item. For example, the object tag - attribute tag integration selected for each clothing item can identify the type of the clothing item and the fabric or material used in the clothing item, such as a cotton shirt, a wool coat, a leather jacket, or a silk dress. This enables architecture 200 to perform dynamic identification 506 of the clothing items to be cleaned. Thereby, processor 120 can generate one or more recommendations 508 that identify one or more settings of washer / dryer 502 or steam cabinet 504. The one or more recommendations 508 can be presented to the user (such as when presented on the display of washer / dryer 502 or steam cabinet 504), or used to automatically set or change one or more settings of washer / dryer 502 or steam cabinet 504. This can help promote the use of appropriate settings for different materials of different clothing.

[0102] As Figure 6As shown, a user's mobile device (such as a smart phone 602 or a tablet computer) can use architecture 200 (or in combination with architecture 200) to identify a specific object in a specific state and perform one or more actions based on the identified object and state. For example, one or more cameras or other imaging sensors 180 can be used to capture an image of the object, and architecture 200 can be used to select object labels 222 and property labels 234 associated with each object. For example, the object label-property label integration selected for each object can identify the type of the object and the state of the object. This enables architecture 200 to perform dynamic identification 604 of the object and its state. Thereby, the processor 120 can generate one or more recommendations 606 based on the object and its state. As a specific example, architecture 200 can be used to identify "fresh salmon", "damp clothing", or "deflated tire", and these visual understanding results can be used to make intelligent recommendations 606, such as "find the nearest auto repair shop" or connect to a smart home control system to ask the user if they want to "preheat" the oven to a specified temperature or perform other actions. One or more recommendations 606 can be presented to the user (such as when presented on the display of smart phone 602) or used to automatically execute or initiate one or more actions.

[0103] Although Figures 3 to 6 illustrates an example use case of zero-shot compositional learning in a machine learning system, various changes can be made to Figures 3 to 6 it. For example, zero-shot compositional learning can be used in any other suitable way, and Figures 3 to 6 the use of zero-shot compositional learning is not limited to the specific example use cases shown and described above.

[0104] Figure 7 illustrates an example method 700 for training a machine learning model to support zero-shot compositional learning according to the present disclosure. For ease of explanation, method 700 is described as being performed by Figure 1 the server 106 in the network configuration 100 shown in Figure 2 using architecture 200 shown. However, method 700 can be performed using any other suitable device (such as the electronic device 101) and in any other suitable system, and method 700 can be used with any other suitable machine learning architecture.

[0105] As Figure 7 shown, at step 702, an image, a set of property labels, and a set of object labels are obtained. This can include, for example, the processor 120 of the server 106 obtaining the image 208 from a repository of training data or other suitable source. This can also include the processor 120 of the server 106 obtaining one or more sets of text categories from a knowledge base, where the text categories include or are used to derive the set of object labels 222 and the set of property labels 234.

[0106] Perform prompt tuning of a pre-trained vision-language model, where the pre-trained vision-language model includes a first text encoder, a second text encoder, and a vision encoder. For example, at step 704, use the first text encoder to generate object text features associated with the object label of each of the plurality of attribute label-object label pairs, and at step 706, use the second text encoder to generate attribute text features associated with the attribute label of each of the attribute label-object label pairs. This can include, for example, the processor 120 of server 106 using the first text encoder implemented by branch 204 of architecture 200 to generate object text features 224 from object label set 222. This can also include the processor 120 of server 106 using the second text encoder implemented by branch 206 of architecture 200 to generate attribute text features 236 from attribute label set 234. At step 708, use the vision encoder to generate image features associated with the image. This can include, for example, the processor 120 of server 106 using the vision encoder implemented by branch 202 of architecture 200 to generate vision features 212 from image 208.

[0107] At step 710, use the transformer encoder layers of the text and vision encoders to process the generated features. This can include, for example, the processor 120 of server 106 using the transformer encoder layers 214a-214n of branch 202 to process vision features 212, using the transformer encoder layers 226a-226n of branch 204 to process object text features 224, and using the transformer encoder layers 238a-238n of branch 206 to process attribute text features 236. As part of the processing, at step 712, generate layer-specific learnable prompt tokens and append them to the input of the specified layers in the text and vision encoders. This can include, for example, the processor 120 of server 106 using the above techniques to generate shared prompts 218, 230, and 242, which can be provided to the transformer encoder layers 214a-214k, 226a-226k, 238a-238k in each of branches 202, 204, 206.

[0108] In step 714, an object score is generated based on the similarity between the object text features and the image features, and an attribute score is generated based on the similarity between the attribute text features and the image features. This can include, for example, the processor 120 of server 106 using the outputs of the final transformer encoder layers 214n and 226n to generate the object score 244, and using the outputs of the final transformer encoder layers 214n and 238n to generate the attribute score 246. In step 716, during prompt tuning, the pre-trained vision-language model is trained to select one of the attribute labels and one of the object labels that match the content included in the image. This can include, for example, the processor 120 of server 106 selecting the object label 222 associated with the highest object score 244 and selecting the attribute label 234 associated with the highest attribute score 246.

[0109] The pre-trained (and now tuned) vision-language model can be used in any suitable manner. For example, the trained vision-language model can be stored, output, or used at step 718. This can include, for example, the processor 120 of server 106 storing the vision-language model and using the vision-language model during inference. This can also or alternatively include server 106 deploying the vision-language model to one or more other devices (such as electronic device 101) for inference.

[0110] Although Figure 7 illustrates an example of method 700 for training a machine learning model to support zero-shot compositional learning, various changes can be made to Figure 7 For example, although shown as a series of steps, Figure 7 the various steps in

[0111] Figure 8 can overlap, occur in parallel, occur in a different order, or occur any number of times (including zero times). Figure 1 illustrates an example method 800 for using a machine learning model trained to perform zero-shot compositional learning according to the present disclosure. For ease of explanation, method 800 is described as being performed by Figure 2 the electronic device 101 in the network configuration 100 shown in

[0112] using the architecture 200 shown in Figure 8As shown, an image of the object is obtained at step 802. This can include, for example, the processor 120 of the electronic device 101 obtaining at least one image 208 from at least one imaging sensor 180 or other suitable source. At step 804, a set of attribute tags and a set of object tags are obtained. This can include, for example, the processor 120 of the electronic device 101 obtaining an input 220 that includes or identifies the set of object tags 222 and an input 232 that includes or identifies the set of attribute tags 234.

[0113] At step 806, an object text feature associated with each object tag is generated using a first text encoder of the vision-language model, and at step 808, an attribute text feature associated with each attribute tag is generated using a second text encoder of the vision-language model. This can include, for example, the processor 120 of the electronic device 101 using the first text encoder implemented by the branch 204 of the architecture 200 to generate object text features 224 from the set of object tags 222. This can also include the processor 120 of the electronic device 101 using the second text encoder implemented by the branch 206 of the architecture 200 to generate attribute text features 236 from the set of attribute tags 234. At step 810, an image feature associated with the image is generated using a vision encoder of the vision-language model. This can include, for example, the processor 120 of the electronic device 101 using the vision encoder implemented by the branch 202 of the architecture 200 to generate visual features 212 from the image 208.

[0114] At step 812, the generated features are processed using a transformer encoder layer of the text and vision encoders. This can include, for example, the processor 120 of the electronic device 101 using the transformer encoder layers 214a - 214n of the branch 202 to process the visual features 212, using the transformer encoder layers 226a - 226n of the branch 204 to process the object text features 224, and using the transformer encoder layers 238a - 238n of the branch 206 to process the attribute text features 236. As part of the processing, at step 814, layer-specific multimodal shared prompts are concatenated to the input of the specified layers in the text and vision encoders. This can include, for example, the processor 120 of the electronic device 101 using the above techniques to generate shared prompts 218, 230, and 242, which can be provided to the transformer encoder layers 214a - 214k, 226a - 226k, 238a - 238k in each of the branches 202, 204, 206.

[0115] At step 816, an object score is generated based on the similarity between the object text feature and the image feature, and an attribute score is generated based on the similarity between the attribute text feature and the image feature. This may include, for example, the processor 120 of the electronic device 101 using the outputs of the final transformer encoder layers 214n and 226n to generate the object score 244 and using the outputs of the final transformer encoder layers 214n and 238n to generate the attribute score 246. At step 818, one of the attribute tags and one of the object tags are selected. This may include, for example, the processor 120 of the electronic device 101 selecting the object tag 222 and the attribute tag 234 whose associated object score 244 and attribute score 246 generate the largest product when multiplied.

[0116] The selected object and attribute tags can be used in any suitable manner. For example, the selected object and attribute tags can be stored, output, or used at step 820. This may include, for example, the processor 120 of the electronic device 101 using the selected object and attribute tags 222, 234 to identify and perform one or more actions (or initiate the execution of one or more actions) based on the selected object and attribute tags 222, 234. Examples of ways in which the selected object and attribute tags 222, 234 can be used are described above with reference to Figures 3 to 6 but the selected object and attribute tags 222, 234 can be used in any other suitable manner.

[0117] Although Figure 8 FIG. 8 shows an example of the method 800 for using a machine learning model trained to perform zero-shot compositional learning, various changes can be made to Figure 8 it. For example, although shown as a series of steps, Figure 8 the various steps in

[0118] can overlap, occur in parallel, occur in a different order, or occur any number of times (including zero times). Although the present disclosure has been described with reference to various example embodiments, various changes and modifications can be suggested to those skilled in the art. The present disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims.

Claims

1. A method, comprising: Obtaining an image, a set of attribute labels, and a set of object labels; And Performing prompt tuning of a pre-trained vision-language model, the pre-trained vision-language model including a first text encoder, a second text encoder, and a vision encoder, the pre-trained vision-language model being trained during the prompt tuning to select one of the attribute labels that matches the content included in the image and one of the object labels; Wherein performing the prompt tuning includes, for each of a plurality of attribute label-object label pairs: Using the first text encoder to generate object text features associated with the object label of the attribute label-object label pair; Using the second text encoder to generate attribute text features associated with the attribute label of the attribute label-object label pair; and Using the vision encoder to generate image features associated with the image; Wherein intermediate outputs from initial layers of the first text encoder, the second text encoder, and the vision encoder are integrated to generate layer-specific learnable prompt tokens, the layer-specific learnable prompt tokens being appended to the input of a specified layer in the first text encoder, the second text encoder, and the vision encoder during the prompt tuning.

2. The method according to claim 1, wherein: During the prompt tuning, a plurality of images including objects associated with a subset of the attribute label-object label pairs are used to train the pre-trained vision-language model; And The pre-trained vision-language model after the prompt tuning is configured to evaluate additional images in a zero-shot manner and identify objects within the additional images associated with attribute label-object label pairs not seen during the prompt tuning.

3. The method according to claim 1, wherein, Performing the prompt tuning further includes: Determining an object score based on the similarity between the object text features and the image features; Determining an attribute score based on the similarity between the attribute text features and the image features; and Tuning the prompts of the pre-trained vision-language model based on the object score and the attribute score.

4. The method according to claim 1, wherein: The vision encoder processes the image during the prompt tuning using learnable visual prompts added to the image; and The pre-trained vision-language model is trained during the prompt tuning to consider each integration of the attribute labels from the set of attribute labels and the object labels from the set of object labels when identifying the content included in the image.

5. The method according to claim 1, wherein: The pre-trained vision-language model includes a transformer-based machine learning model with three branches having transformer layers; The first branch in the branches of the transformer layer implements the first text encoder; The second branch in the branches of the transformer layer implements the second text encoder; and The third branch in the branches of the transformer layer implements the vision encoder.

6. The method according to claim 5, wherein: Generate one or more initial visual chunking cues at one or more random positions in the image; Generate one or more initial text cues using a set of known text categories; The initial visual cues and text cues are projected to at least one of the transformer layers in each of the branches in the branch to generate layer-specific initial cue tokens, where different branches are associated with different layer-specific initial cue tokens; And In each branch of the transformer-based machine learning model: Each of one or more initial layers in the branch receives one or more of the layer-specific initial cue tokens and uses the one or more layer-specific initial cue tokens to generate one or more embeddings; And Each of one or more subsequent layers in the branch receives one or more embeddings and one or more cues from a previous layer in the branch.

7. The method according to claim 1, wherein: The first text encoder and the second text encoder represent separate machine learning classifiers; The machine learning classifiers are jointly trained using a loss function that integrates losses associated with the machine learning classifiers; And The machine learning classifiers are used to select one of the attribute labels and one of the object labels by identifying the attribute label-object label integration with the highest score.

8. An apparatus, comprising: At least one processing device configured to: Obtain an image, a set of attribute labels, and a set of object labels; And Perform prompt tuning of a pre-trained vision-language model, the pre-trained vision-language model including a first text encoder, a second text encoder, and a vision encoder, the pre-trained vision-language model being trained during the prompt tuning to select one of the attribute labels and one of the object labels that match the content included in the image; Wherein, to perform the prompt tuning, the at least one processing device is configured, for each of a plurality of attribute label-object label pairs: Generate object text features associated with the object label of the attribute label-object label pair using the first text encoder; Generate attribute text features associated with the attribute label of the attribute label-object label pair using the second text encoder; and Generate image features associated with the image using the vision encoder; Wherein the at least one processing device is configured to integrate intermediate outputs from initial layers of the first text encoder, the second text encoder, and the vision encoder to generate layer-specific learnable cue tokens, the layer-specific learnable cue tokens being appended to the input of a specified layer in the first text encoder, the second text encoder, and the vision encoder during the prompt tuning.

9. The apparatus according to claim 8, wherein, To perform the prompt tuning, the at least one processing device is further configured to: Determine an object score based on the similarity between the object text features and the image features; Determine an attribute score based on the similarity between the attribute text features and the image features; And A prompt for tuning the pre-trained vision-language model based on the object score and the attribute score.

10. The apparatus according to claim 8, wherein: the vision encoder is configured to process the image by using learnable visual prompts added to the image during the prompt tuning; and during the prompt tuning, the pre-trained vision-language model is trained to consider each integration of the attribute labels from the set of attribute labels and the object labels from the set of object labels when identifying the content included in the image.

11. The apparatus according to claim 8, wherein: the pre-trained vision-language model includes a transformer-based machine learning model with three branches having transformer layers; the first branch in the branches of the transformer layer implements the first text encoder; the second branch in the branches of the transformer layer implements the second text encoder; and the third branch in the branches of the transformer layer implements the vision encoder.

12. The apparatus according to claim 11, wherein: the at least one processing device is further configured to: generate one or more initial visual chunk prompts at one or more random positions in the image; generate one or more initial text prompts by using a set of known text categories; project the initial visual prompts and text prompts for at least one transformer layer in the transformer layer in each of the branches to generate layer-specific initial prompt tokens, with different branches associated with different layer-specific initial prompt tokens; and in each branch of the transformer-based machine learning model: each of one or more initial layers in the branch is configured to receive one or more of the layer-specific initial prompt tokens and use the one or more layer-specific initial prompt tokens to generate one or more embeddings; and each of one or more subsequent layers in the branch is configured to receive one or more embeddings and one or more prompts from the previous layer in the branch.

13. The apparatus according to claim 8, wherein: the first text encoder and the second text encoder represent separate machine learning classifiers; the at least one processing device is configured to jointly train the machine learning classifiers by using a loss function that integrates the losses associated with the machine learning classifiers; and the machine learning classifiers are configured to select one of the attribute labels and one of the object labels by identifying the attribute label-object label integration with the highest score.

14. A method, comprising: obtaining an image of an object; obtaining a set of attribute labels and a set of object labels; and using a vision-language model including a first text encoder, a second text encoder, and a vision encoder to select one of the attribute labels and one of the object labels associated with the object; wherein selecting one of the attribute labels and one of the object labels associated with the object includes: using the first text encoder to generate object text features associated with one of the object labels; Generate an attribute text feature associated with one of the attribute labels using the second text encoder; and Generate an image feature associated with the image using the visual encoder; wherein one or more layers in the first text encoder, one or more layers in the second text encoder, and one or more layers in the visual encoder are each associated with a layer-specific multimodal shared prompt that is concatenated to the input of the layer.

15. The method according to claim 14, wherein, Selecting one of the attribute labels associated with the object and one of the object labels further includes:[[]] For each object label, determining an object score based on the similarity between the image feature generated by the final layer of the visual encoder and the object text feature associated with the object label generated by the final layer of the first text encoder; Selecting the object label associated with the highest object score; For each attribute label, determining an attribute score based on the similarity between the image feature generated by the final layer of the visual encoder and the attribute text feature associated with the attribute label generated by the final layer of the second text encoder; and Selecting the attribute label associated with the highest attribute score.