Systems and methods for multimodal conversational recommendation

By creating large, naturalistic datasets and using retrieval-augmented generation, the challenges of generating effective fashion recommendations are addressed, enhancing model controllability and fairness, and enabling high-quality multimodal fashion recommendations.

WO2026015602A1PCT designated stage Publication Date: 2026-01-15RGT UNIV OF CALIFORNIA

Patent Information

Application Number
PCT/US2025/036925
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2025-07-09
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Conversational models struggle to generate effective fashion recommendations due to challenges in handling long-tail domains, temporal dynamics, visual semantics, and multimodal data, and lack appropriate evaluation strategies, leading to issues like item bias and fairness concerns.

Method used

Develop large, naturalistic datasets that combine conversational and collaborative data, incorporate multimodal inputs, and use retrieval-augmented generation to enhance model controllability and fairness, employing LLMs and VLMs for personalized fashion recommendations.

Benefits of technology

The approach enables high-quality, multimodal fashion recommendations by addressing data and evaluation limitations, improving model controllability and fairness, and expanding the domains where conversational recommendations can be effectively made.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025036925_15012026_PF_FP_ABST
    Figure US2025036925_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Conventional recommender systems lack context awareness and semantic reasoning capabilities, while language models lack the ability to clearly delineate recommendations from generated text so as to correct biases. Methods and devices for providing multi-modal fashion recommendations are described. In some implementations, a dataset is collected that includes conversations between users seeking fashion recommendations using natural language text and images. Recommendation items are identified in the conversations and mapped to known entities. The dataset is used to train a multi-modal machine learning model that can interpret visual semantics of an image to generate fashion recommendations in response to natural language requests. The known entities can be extracted from the model output and biases can be mitigated before the final recommendation is presented.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR MULTIMODAL CONVERSATIONALRECOMMENDATIONCROSS-REFERENCE TO RELATED APPLICATION

[0001] This patent document claims priority to, and benefits of, U.S. Provisional Patent Application No. 63 / 669,166 entitled “SYSTEMS AND METHODS FOR MULTIMODAL CONVERSATIONAL RECOMMENDATION” filed on July 9, 2024. The contents of the aforementioned patent application are incorporated by reference in their entirety as part of the disclosure of this patent document.TECHNICAL FIELD

[0002] This document generally relates to machine learning applications and associated methods, and more particularly, to multimodal architectures for content recommendation.SUMMARY

[0003] Embodiments of the disclosed technology relate to systems, methods, and devices for a conversational Al shopping assistant. The disclosed embodiments can be used in many fields, such as dataset collection, multimodal model design, and conversational recommendation generation. For example, the embodiments described in the present document enable the creation of new data sets and foundation models to support fashion recommendation, especially via conversational interfaces.BRIEF DESCRIPTION OF THE DRAWINGS

[0001] Figure 1 illustrates an example of item bias in an LLM-based recommender.

[0002] Figure 2 illustrates relationships between the methodology, applications, and broader impacts of the disclosed technology.

[0003] Figure 3 illustrates example conversation quotations included in conversational recommendation datasets, in accordance with example embodiments.

[0004] Figure 4 illustrates a data construction pipeline for generating a conversational recommendation dataset that includes naturally occurring conversational recommendations, according to an example embodiment.

[0005] Figure 5 enumerates potential sources of conversational recommendation data from Reddit.

[0006] Figure 6 illustrates examples of multimodal inputs, in which media content is requested in part by using an image.

[0007] Figure 7 illustrates an example re-index framework (the Re-index-Then-Adapt, or RTA, framework), in accordance with an embodiment of the disclosed technology.

[0008] Figure 8 illustrates two example supervised fine-tuning strategies, according to some embodiments of the disclosed technology.

[0009] Figure 9 illustrates a neighborhood+LM recommender framework, in accordance with an embodiment of the disclosed technology.

[0010] Figure 10 illustrates a prompt-based retrieval-augmented conversational recommendation, according to an example embodiment.

[0011] Figure 11 illustrates results demonstrating fairness of recommendations across user demographics.

[0012] Figure 12 illustrates five tasks designed to evaluate a property that a synthetic user should exhibit, according to embodiments of the disclosed technology.

[0013] Figure 13 illustrates a comparison of item ratings between humans and a simulator prompted with demographic information and / or a “pickiness” personality.DETAILED DESCRIPTION

[0004] The disclosed technology relates to the development and use of foundation models for conversational fashion recommendation applications, and specifically for tasks involving a fusion of text, temporal, visual, and user dynamics. “Conversational recommendation,” as used herein, broadly refers to any type of recommender system that involves iterated interaction, usually featuring natural language. Examples that apply conversational paradigms include: free-formconversation with a goal of finding a relevant item; systems that produce recommendations with a natural language interface; systems that explain recommendations, where the user can provide feedback, such as natural language input and / or traditional feedback (e.g., clicking one of a plurality of user interface elements); and / or systems that augment dialogs between humans with recommendations. Augmenting dialogues between humans can include configuring systems to perform as conversational recommenders for settings in which users discuss fashion, especially in recommendation-oriented communities where users upload images, discuss potential purchases, or ask for fashion advice.

[0005] Conversational paradigms have recently seen an explosion of interest, following developments in general -purpose Large Language Models (LLMs). Such conversational models are permeating all aspects of computer science, including systems for clinical support, education, gaming, and the like. These models have a capacity to personalize their outputs, which they mostly accomplish simply by handling long context.

[0006] Conversational models can be used as general-purpose recommenders with little to no modification. For example, a model can provide a conversational interface to discuss products, movies, music, and other content of interest to a user. A user wishing for movie recommendations might describe their current circumstances, and the types of movies they normally enjoy, and in response a language model might suggest items semantically related to the query. The suggested items may be beyond what is possible with traditional collaborative approaches, which focus on user-item interaction data. Such conversational systems can be particularly effective for certain types of general-purpose recommendation that are well -represented in training corpora, such as movie recommendations, as the details and semantics of popular movies can consequently be easily understood by such systems.

[0007] However, because such well -represented domains dominate current datasets and evaluations, conversational models are ineffective at making recommendations in other domains. One such domain where general -purpose conversational models are ineffective is in generating fashion recommendations. Fashon presents several difficulties for conversational (e.g., pretrained) models: new, out-of-vocabulary items are constantly introduced; trends are dominated by shifting temporal dynamics; recommendation involves interpreting complex visual semantics; and users want interfaces that support relevant searching modalities, such as by allowing users toupload their own images.

[0008] Generating fashion recommendations in this context can be particularly challenging for existing language models for a number of reasons. Current recommendation systems based on machine learning (e.g., Al-powered chatbots) are limited by the quality of their dataset and often provide irrelevant or incorrect answers. Pre-trained models work poorly in long-tail domains (such as fashion), where there may have been few discussions about the specific items in question, and also work poorly in “cold-start” situations in which there is little to no prior context for an item, which are ubiquitous in fashion. Large pre-trained models are also inadequate when generating recommendations for domains where information and / or content (and therefore relevant recommendations) are temporally evolving.

[0009] Furthermore, existing conversational recommenders struggle to fuse different modalities of data, such as temporal, visual, and linguistic signals. A current state of the art chatbot, when asked what shoes would match an outfit, would not be able to process the subjective concept of “matching” and would return all products tagged in the “shoe” category. While these systems have demonstrated success in a few domains, they operate purely on text data, and struggle to handle recommendation modalities that incorporate other types of information. For example, existing systems (like ChatGPT) are effective at making recommendations in certain domains (like movies), but struggle with data like images, time-series, geographical information, and the like, which aren't as easily supported through chat-like interfaces.

[0010] Additionally, generating content recommendations is not often implemented by training conversational models for content recommendation, due to several existing limitations. For example, evaluating existing models for conversational recommendation is particularly challenging due to a lack of appropriate datasets. Relevant datasets tend to be small, unnatural, or focused on narrow domains. There is difficulty implementing knowledge grounding and controllability in conversational recommenders, especially for performing fairness and bias interventions. Fairness issues that have effective strategies when using traditional collaborative recommenders can re-emerge when using conversational approaches, where one has limited inference-time control over the model's output distribution. Additionally, existing evaluation strategies are generally limited to assessing whether a model can predict a withheld (e.g., predetermined) item, which fails to consider other desirable characteristics, including beyond-accuracy qualities such as persuasiveness, inspiration, (user-side) fairness, serendipity, and the like.

[0011] In the present disclosure, these problems are addressed and novel approaches to conversational recommendation are disclosed, addressing issues of data, methodology, and evaluation. For example, described herein are methodologies for collecting and using datasets that are large, naturalistic, highly contextual, and carefully annotated. These datasets can have broad coverage of long-tail categories, can be connected to user interaction histories, and can be linked to metadata and knowledge about items. The disclosed technology includes a controllable generation framework that can address issues related to both user- and item-fairness in conversational settings. In some embodiments, a system uses a retrieval -augmented generation that enables model controllability, facilitating the same types of interventions that are possible with traditional (collaborative) models. Additionally, controllable simulation approaches are described that allow conversational models to be evaluated offline (e.g., without requiring real-time human interaction) by mimicking the behavior of real users. This reduces evaluation effort, aids reproducibility, and facilitates controllable simulation of user subgroups.

[0012] The disclosed technology includes systems and methods for utilizing multimodal data to generate high-quality content recommendations. For example, systems and methods to develop new foundation models for conversational fashion tasks are described. In some embodiments, the kinds of task(s) targeted involve conversational datasets which include traditional recommendation objectives, such as predicting the next item a user will interact with, based on collaborative interactions in the dataset; temporal predictions, including temporal recommenders, but also prediction of long-term and periodic trends; visually-aware discriminative tasks, such as visual recommendation and outfit completion; pairwise tasks such as compatibility estimation; generative image tasks such as item synthesis; generative language tasks such as conversational recommendation; and fairness interventions such as calibrated recommendation, diversification, and the like.

[0013] In one aspect, the disclosed embodiments include a recommendation system utilizing visual data (e.g., for use in fashion recommendations), temporal data (e.g., data that describes changes in information, such as fashion trends, over time), geographical data (e.g., data that describes changes in information, such as fashion preferences, based on location and / or geographicposition), along with data from external knowledge bases. In one aspect, the disclosed embodiments include a visual recommendation system able to provide suggestions for complementary products based on their appearance using fashion-based matching. For example, the system can use fashion-based matching to suggest a product that is compatible with the information (e.g., conversation prompts and / or images) presented to the system. In another aspect, the disclosed embodiments include a shopping assistant that can process a user-uploaded image and a natural language question (e.g., “what shoes would go with this outfit?”) to provide high- quality suggestions based on a complementary product’s visual characteristics.

[0014] In another aspect, disclosed embodiments include a method for collecting a dataset from authentic human fashion recommendations. This dataset can be used to configure a system with fashion recommendation capabilities exceeding technologies that determine recommendations primarily by using text tags associated with a product (e.g., that label or describe a product, such as a book labelled as “mystery”, or a jacket labelled as “denim”). In another aspect, disclosed embodiments include systems for conversational recommendation, in which recommender systems (e.g., systems that recommend items to users based on user preferences and past interactions) are augmented with conversational interfaces (e.g., so that users can interact with these systems via natural language).

[0015] The disclosed technology includes combining recent advances in Large Language Models (LLMs) and Vision Language Models (VLMs) with recommender systems. In some embodiments, an example system includes one or more LLMs (e.g., to receive natural language questions or commands from users), one or more VLMs (e.g., to process images or examples received from users), and one or more recommender systems (e.g., to record and / or process users' interaction histories and / or semantic relationships among items). The example system can generate multimodal recommendations through natural language interfaces. In some embodiments, an example system can use one or more LLMs to make interactive, conversational recommendations to users. For example, an example system can combine state-of-the-art algorithms for personalized recommendations with LLMs (including visually-aware models) in order to expand the domains in which effective conversational recommendations can be made. In some implementations, an example system incorporates a VLM to perform visual discrimination tasks based on an iterated interaction (e.g., between the system and a user).

[0016] Outputs of the disclosed technology can include the following. First, a large dataset for conversational recommendation model training, including (a) collaborative data, i.e., interactions between users and items; (b) conversational threads, (e.g., sourced from Reddit posts describing Amazon products); (c) user and product images. Second, an evaluation suite to assess conversational recommendation models (and related predictive tasks) under a variety of scenarios. Third, foundation models for conversational fashion tasks.

[0017] In addition to training large models, the disclosed methods and systems also require significant computational support for (a) data processing and labeling (especially using language models to process conversational data into structured form); (b) synthetic evaluation of conversational recommenders (to simulate users); (c) model inference and user studies.

[0014] The disclosed technology extends the current knowledge and techniques in each of the areas in which it builds, including recommender systems, NLP, controllability, and fairness. Combining and integrating technologies in these areas enables a host of exciting new applications dealing with mixed datasets that combine language with user interaction data, ranging from standard recommendation problems, to personalized language understanding tasks in fashion, health, and therapy, among others. The disclosed technology can be used well outside of traditional content recommendation, and can apply, for example, to fields such as medicine to combine Electronic Health Record (EHR) data with mixed-modality patient attributes and imaging data.Introduction

[0015] The areas of conversational recommendation, and more broadly personalization of large language models have gained renewed interest recently, as general-purpose models (e.g., ChatGPT) have suddenly outperformed traditional, carefully crafted models.

[0016] Although this new class of models has quickly supplanted traditional methods, their performance and generalization characteristics are currently understood only superficially. For example, little is understood about (a) whether they perform as well as they are anecdotally believed to; (b) what mechanisms they use to achieve such excellent performance; (c) whether their excellent in-domain performance generalizes to longer-tail domains; and (d) what the hidden risks are of deploying such models, especially in terms of fairness and bias.

[0017] In the current disclosure, several advancements are described that address these questions and realize progress in this quickly-growing space:

[0018] First: better datasets are disclosed that are targeted toward conversational recommendation with LLMs. Although a few popular benchmarks exist (see, e.g., ReDial, INSPIRED), they tend to be small; artificially constructed (e.g., via crowdsourcing); and / or limited to a few specific domains (especially movies).

[0019] Second: models for conversational recommendation are disclosed that involve a combination of collaborative and semantic knowledge, i.e., they fuse components from recommender systems and language modeling. Ideal LLM-based recommenders should offer mechanisms to leverage external knowledge and facilitate fine-grained controllability.

[0020] Third: methods for evaluating LLM-based recommenders are disclosed. LLM-based recommenders can potentially have significant bias issues, partly because they lack the finegrained control mechanisms available to conventional recommenders. LLM-based recommenders also present new opportunities for evaluation, especially in terms of “beyond-accuracy” metrics (serendipity, persuasiveness, and the like), and user-side fairness.

[0021] The disclosed technology proposes solutions in three thrusts that address the above limitations. Each thrust shows how novel methods for personalized conversational agents can be implemented via new datasets, methods, and applications. In Thrust 1, we focus on new datasets, noting limitations of current data in terms of breadth, connections to collaborative filtering data, and multimodality. In Thrust 2, we focus on issues of knowledge and control in conversational recommendation, noting that many previously successful techniques are no longer applicable in conversational settings, and proposing corrections via controllable LLMs. In Thrust 3, we consider evaluation strategies, focusing on issues of user-side fairness, and offline evaluation via LLM- based user simulators.

[0022] Throughout, connections will be drawn to broader impacts, such as the development of foundation models for patient EHR data, conversational movie recommendations, and large- scale model evaluation.Conversational Recommender Systems

[0023] Broadly speaking, conversational recommender systems combine ideas from dialog generation, explainability, and interactive recommendation. Early work with the term is generally used to refer to any number of systems that collect feedback iteratively over several turns (whether or not they involve natural language). As such, early methods facilitate simple iterative feedback from users, and only more recently have methods adopted a paradigm of free-form conversation. Early approaches for conversational recommendations essentially treated conversation as a form of iterative query refinement. For example, users could be asked questions that attempt to determine their preferences or constraints toward a fixed set of attributes (e.g., cuisine type, price range, and the like), effectively implementing a user model that is a weighting over potential attribute values. Other early approaches to conversational recommendation are essentially forms of interactive recommendation, in which a system can query users or gather feedback from users during each round. The goal is generally to use a conversational strategy to quickly infer users' preferences toward item properties. Interactions consist of a one-sided form of conversation in which the system asks simple questions to the user about their preferences.

[0024] More recent approaches try to follow the conversational paradigm more literally. That is, conversations are free-form, in which both the system's question and the user's response take the form of text. A major milestone of this approach is to build initial ground-truth datasets for this task. In these (and most existing benchmarks), dialogs are constructed by crowd workers, who assume roles of a recommender or a seeker; conversations between the recommender and the seeker are tagged in terms of the items (e.g., movies) mentioned, as well as explicit feedback (e.g., has the seeker seen the movies mentioned and did they like them). Until recently, methods which carefully combined recommendation modules, conversational modules, and knowledge graphs (e.g., UniCRS) represented the state-of-the-art on conversational recommendation benchmarks.

[0025] General-purpose language models, despite not having been designed for conversational recommendation specifically, quickly matched or outperformed the strongest handcrafted models. New datasets built to support this hypothesis (based on movie discussions from Reddit) also revealed surprisingly different characteristics compared to existing benchmarks.

[0026] Figure 1 illustrates an example of item bias in an LLM-based recommender. While LLM-based recommenders can perform well at generating recommendations, they have significant fairness issues, especially in terms of item bias. Conversational models frequently recommend items that are rarely mentioned in real data, and conversely there are many popular items that are rarely recommended. Figure 1 illustrates results showing the frequency of recommendations given compared to the frequency of recommendations as they appeared in training data. For a fair model, points on the scatterplots in Figure 1 would approximately follow a positive line extending from the origin. The bias described above is essentially a form of item-side bias known as popularity bias. Such issues are relatively easy to correct for traditional recommenders, but are difficult to correct using LLMs, as one has relatively little control over item distributions during decoding. That is, while a traditional recommender can return an ordered list of entries, an LLM can return one or more recommendations in a natural language form, making it difficult to correct the distribution of recommendations using statistical techniques. Additionally, conversational models perform well due to their ability to handle semantic information, but struggle to leverage collaborative signals. Such collaborative signals can be derived from interactions between users and items, and can include prior knowledge of which items are often paired by users (e.g., what items “go together”). An ideal model should leverage both collaborative information (e.g., from a traditional recommendation framework), and semantic knowledge (e.g., from an LLM). The strongest pre-LLM models (e.g., UniCRS) generally followed this paradigm, though such systems are hard to combine with LLM frameworks.

[0027] A focus of the disclosed technology is to implement practical personalization in the context of large language models. Developing practical approaches to large-scale, personalized, conversational recommendation — especially involving LLMs — presents a range of challenges. Methods must be scalable, as many of the datasets that are considered span several terabytes of actions, text, and images. Moreover, such datasets are sparse and long-tailed when considering the problem of modeling individuals. There is thus a need for models that trade off complexity with parsimony: generative models of text and images already require billions of parameters and lengthy training schedules, and thus practical solutions should leverage a combination of modeling, fine-tuning, and retrieval-augmented generation (RAG).

[0028] Methodologically, the disclosed technology contributes to each of the research areas upon which it builds. In terms of recommender systems and conversational agents, the disclosed technology contributes a suite of techniques and approaches that combine ideas from traditional collaborative filtering methods with those from natural language generation. The datasets introduced below facilitate the development of new methods that are diverse, knowledge grounded, and multimodal. Such applications are ubiquitous within recommender systems and online dialogue, but also emerge in other settings, such as those related to clinical machine learning, behavioral therapy, or hiring.

[0029] Figure 2 illustrates relationships between the methodology, applications, and broader impacts of the disclosed technology. The disclosed technology enables new applications in recommender systems and personalized modeling. The disclosed technology advances the way systems are trained and evaluated, contributes to new domains (rather than common domains, such as movie recommendations), explores long-tail scenarios, considers beyond-accuracy metrics, and enables models that better represent human-to-human interaction patterns. The disclosed technology is directly related to the core areas of data acquisition / exploration / analysis, and ML / DM techniques applicable to large, high-velocity, complex, and heterogenous datasets. The disclosed technology also contributes to understanding knowledge in the context of ML models, and expands the breadth of recommender systems in terms of application domains.

[0030] The disclosed technology is applicable to clinical applications, especially those that deal with clinical NLP and Electronic Health Record (EHR) data. For applications like personalized health, it is critical to deal with structured, multimodal data, while simultaneously accounting for variance that arises due to individual behavioral patterns. Examples can include the use of personalized conversational models on applications in cognitive behavioral therapy, and in particular, the use of “reframing” strategies. Such ideas are closely related to personalized conversation.Thrust 1 : New Datasets for Conversational Recommendation

[0031] The first thrust includes building datasets that are more natural and realistic. Existing conversational datasets do not often collect real-world, naturally-occurring conversations. Rather, users (usually crowdsourced) are not genuinely seeking a recommendation, but are ratherinstructed on which item they should seek from another user. While useful for establishing benchmarks for evaluation, such datasets have several limitations. For example, such datasets cover few domains (see, e.g., ReDial and INSPIRED, which focus on movies). It is unclear to what extent similar data collection efforts could be applied in other settings, particularly ones not based on general knowledge, for which crowd workers would struggle to engage in synthetic conversations. Even within a common domain such as movies, it is hard to tell how closely conversations in current benchmarks represent organic conversations (e.g., whether constructed datasets are similar to how people would naturally discuss and recommend items). Furthermore, given the synthetic nature of data collection described above, constituent conversations tend to be shallow (e.g., by focusing on recommending popular and / or otherwise obvious items), users in such synthetic conversations generally have properties of a known item in mind, and as such, these approaches are not particularly useful for beyond-accuracy metrics including novelty, discovery, serendipity, user satisfaction, and the like.

[0032] In our first thrust, we disclose approaches for building datasets that mimic how humans would actually interact when making recommendations, versus current, more synthetic, settings. Some related efforts, such as INSPIRED, aim for a natural setting, but are (relatively) small; in contrast, our own previous efforts to synthesize conversational datasets (from product review text) were much larger but of low quality (e.g., not natural conversations). In contrast to both, we disclose an approach to acquire datasets that are simultaneously large and highly natural.

[0033] Discussed below are approaches for building (or harvesting) large conversational recommendation datasets that contain more natural data (e.g., closer to human-to-human recommendation dialogues), focusing on breadth (e.g., covering more item domains), fusing conversational and collaborative data , and building multimodal (especially vision-language) conversational recommendation datasets.

[0034] Figure 3 illustrates example conversation quotations included in conversational recommendation datasets, in accordance with example embodiments. In focusing on breadth, conversational recommendation datasets are built that have greater breadth in terms of conversational style and the types of items being considered. Generally, datasets for conversational recommendation focus on just a few application domains, especially movies. Movie datasets have proliferated partly because they have long been a focus of conventional collaborative filteringmodels; but more critically, it is relatively easy to collect data from the laity of users, e.g., crowdsourced workers, which may be much more difficult for other (higher-expertise, longer- tailed, and the like) domains. Language models also have significant exposure to data about movies (and conversations about movies), and are thus surprisingly effective conversational recommenders for movies, despite having never been trained for the task. However, there are many reasons to suspect that the performance of models on movie datasets may be non-representative, and will not generalize to other domains. Existing datasets do not capture long-tailed data, such as cold-start items or items that have little exposure in conversational datasets. Furthermore, existing datasets do not describe items that have subtle or difficult-to-describe semantics, or which require significant expertise. Compare for example the semantics of movies versus those of songs, power tools, or local restaurants.

[0035] Figure 3 contrasts simpler dataset examples with more complex conversation segments included in some embodiments of the present disclosure. While prior data could include only items (e.g., a traditional dataset), or a combination of items and a natural language verbal preference (e.g., a conversational dataset), data included in an advanced conversational recommendation model can include items along with complex, natural language, verbal preference or opinion concerning the items. The latter data can represent natural, organic conversations that contain more information and / or more complex information than previous data in which users often explicitly specify preferences.

[0036] Figure 4 illustrates a data construction pipeline 600 for generating a conversational recommendation dataset that includes naturally occurring conversational recommendations, according to an example embodiment. The above issues can be addressed by collecting large datasets of naturally occurring conversational recommendations (e.g., as opposed to synthetic conversational recommendations created specifically for model training). At 602, a system can receive raw data. This can include real conversations harvested from social media, such as Reddit (e.g., datasets from pushshift. io). At 604, the system can identify subsets of interest (e.g., those in which a recommendation was requested and / or offered). At 606, the system can extract posts that are relevant to recommendations. This can involve finding appropriate posts based on metadata (e.g., tags) or other content-related cues. The system can filter undesirable posts (e.g., posts that are low-quality and / or too short). At 608, the system can extract paths through posts to buildconversations. For example, the system can extract the posts by two specific users that address each other, and ignore intervening unrelated posts. This leaves the system with unstructured conversational data. At 610, the system can generate structured conversational data by processing the unstructured conversational data. The system can resolve entities in the datasets to known sets of items and item metadata (e.g., by using matching tools to map entity mentions to known items). The dataset described with respect to Figure 3 used such a pipeline to construct a conversational recommendation dataset about movies, using movie-focused subreddits including r / movies, r / bestofrietflix, r / moviesuggestions, r nelflixbeslof and r / truefihn.

[0037] Determining which entity mentions correspond to the ground-truth (e.g., recommended items rather than items discussed in passing) can be accomplished by techniques such as querying a language model or training a classifier. Correctly resolving mentioned entities to known items (e.g., matching a specific item to a phrase such as “I preferred the original”) can include using strategies such as (a) approximate string matching; (b) using an LLM to resolve entity mentions; (c) building restricted datasets with resolved entities (e.g., via links).

[0038] Figure 5 enumerates potential sources of conversational recommendation data from Reddit, though the current disclosure can be applied when collecting data from a number of sources. Following the above data collection and filtering efforts, these datasets are combined with collaborative information, as described below. The left side of Figure 5 describes which subreddits contain the most conversations linking to products on Amazon. The right side of Figure 5 describes which Amazon categories are most frequently recommended in Reddit conversations. Statistics are compiled simply by retrieving URLs in Reddit conversations and linking against a public Amazon dataset.

[0039] In focusing on fusing conversational and collaborative data, large datasets are built that fuse conversational data with traditional collaborative filtering signals. In one example embodiment of the present technology, a data construction method consists of harvesting large volumes of conversational data, and connecting the data to datasets of historical interactions. Each historical interaction can include a user that recommended one or more items that were accepted by another user. Generally, this approach leverages large, “natural” conversational datasets, in which users discuss and make recommendations, as opposed to synthetic conversational datasets in which users generate recommendations to be used to train models rather than to genuinelyreceive recommendations. Collecting data by itself is not enough, and data must be linked to known items (e.g., movie conversations must have movie mentions resolved to entities in IMDB or in interaction datasets). The main challenge in data collection is to combine this natural conversation data (used for language modeling) with interaction data (e.g., collaboration data, used to train / evaluate a recommender system).

[0040] Many of the datasets collected above with respect to Figures 5 and 6 are associated with items from traditional collaborative filtering datasets (e.g., connecting movie-oriented conversations with items from MovieLens). Making these types of connections will be critical in Thrust 2 when implementing methods that combine LLMs with collaborative filtering. By combining conversational data with interaction data, both semantic and collaborative item representations can be taught, which is critical for certain types of control (e.g., correcting bias).

[0041] Entity mentions can be extracted and NLP techniques used to match entities to known objects in external knowledge bases. However, for many categories, including those in Figure 5, conversations already link to specific items (e.g., via Amazon links). This gives us an immediate means of building several large and broad datasets combining language and interaction components.

[0042] Figure 6 illustrates examples of multimodal inputs, in which media content is requested in part by using an image. Users often convey requests for recommendations in a multimodal manner. For example, a user requesting a fashion recommendation (see categories in Figure 5) may upload images (e.g., of themselves) along with a recommendation request. Such multimodal recommendations can involve complex associations, such as between images and music or other media content. In other, more nebulous settings, users may request recommendations for media based on an image prompt. In focusing on multimodal conversational recommendation datasets, conversational recommendation datasets must be created that fuse language, images, and multimodal data.

[0043] In some embodiments, multimodal datasets of recommendation requests containing both text and images are collected. These multimodal datasets are used to create new models for multi-modal conversational recommendation, as well as benchmarks to evaluate such models. Figure 5 describes some well-suited data sources, e.g., fashion, beauty, DIY, and the like. LinkedAmazon products already include product images that can be used to build visual recommenders, though in many cases a user-uploaded image is included as part of the prompt (e.g., a user posts a selfie when asking for fashion advice, or their project when asking for DIY advice). This technology can also be implemented for recommendation-oriented conversations in other areas that frequently feature visual prompts (e.g., “music that feels like this image”).

[0044] Handling these types of queries is difficult for existing conversational recommenders, which do not handle combined vision-language modalities. Such tasks require nuanced understanding beyond the objects depicted, such as mood, themes, and aesthetics. Visionlanguage models (VLMs) exhibit impressive performance in understanding images, particularly through visually-conditioned language generation tasks. However, VLMs struggle in (visual) conversational recommendation. For example, GPT-4V achieves around 65% accuracy in selection (among 5 candidate items) and includes the ground-truth recommendations among the top-10 items in less than 40% of requests. Open-source VLMs, e g. LLaVA-13B, lag behind GPT- 4V in this task.

[0045] Such datasets can include large, natural conversation datasets, in which users discuss and make genuine recommendations. Data sources can include, for example, natural language conversations on Reddit and / or existing interaction datasets. These datasets can extend the conversation recommendation datasets described above with respect to Figures 5 and 6.

[0046] In summary, disclosed technology associated with Thrust 1 includes: New conversational datasets including a wide range of item categories, including (but not limited to) categories described with respect to Figure 5; datasets that fuse conversational and collaborative information, including data combining Reddit and Amazon crawls (as well as other sources); datasets that fuse vision and language, including conversational fashion recommendation; processing, filtering, cleaning, and annotation (via language rules, classification, or querying LLMs); evaluation of existing methods against the benchmarks developed above; dissemination in the form of public dataset releases, benchmarking and evaluation papers.Thrust 2: Knowledge, Control, and Fairness in Retrieval-Augmented ConversationalRecommendation

[0047] The second thrust addresses issues of knowledge and control in the context of conversational recommendation. We are interested in control mechanisms that will increase the fairness and trustworthiness of conversational recommendation models, while also improving their performance. Making recommender systems more trustworthy and fair has historically been a significant area of research on conventional recommender systems.

[0048] Modifying a recommender system to improve its fairness characteristics typically consists of modifying the training objective of the recommender to avoid or penalize some undesirable outcome. For instance, if a recommender system exhibits poor performance for some underrepresented group (e.g., female users), the training objective might be modified to penalize errors, as well as penalizing the difference of aggregate errors between males and females.

[0049] Unfortunately, a lot of recent progress on making traditional recommender systems more fair does not carry over to conversational paradigms. Conversational recommender systems are generally trained using language modeling objectives (rather than recommendation objectives). Even simple issues like concentration (overly focusing on a small subset of items) and popularity bias seem to emerge in generated conversational recommendations (as described with respect to Figure 1): Certain items are recommended extremely frequently (e.g., "The Hangover" and "John Wick" are significantly over-represented in predictions compared to their frequency in human recommendations); and conversely, several popular human recommendations are very rarely recommended by LLM agents.

[0050] There are known strategies to address these types of fairness issues context of non- conversational recommendation. For example, the training objective can be modified to minimize some undesired disparity, or the recommender system's outputs can be re-ranked (e.g., by a secondary system) to increase the diversity of the recommended items. Neither of these types of interventions are possible in a conversational setting, where (a) it may not be possible to access the original language model in order to modify its training objective (or retraining may simply be prohibitively expensive); and (b) there may be no way to straightforwardly alter the model's free-text outputs (e.g., to re-rank them) due to the recommendations being integrated into a natural language output.

[0051] Existing (general-purpose) LLMs are not trained with any recommendation objective in mind, and therefore may simultaneously have sub-optimal accuracy while also being biased. Thus, in this context, bias and accuracy are not at odds with each other, and alignment between conversational and recommendation metrics can lead to systems that are more fair while also improving their accuracy. Below, these questions of knowledge, control, and fairness metrics are discussed in the context of conversational recommenders.

[0052] Discussed below are approaches for implementing mechanisms in conversational recommenders that allow more control over recommended items, such as implementing control through knowledge augmentation, collaborative filtering / LLM fusion, and retrieval-augmented prompting.

[0053] Specifically, in Thrust 2 we and characterize and address the main controllability and bias issues in conversational recommendation. Partly this involves measuring the same types of outcomes as in existing work on recommendation bias / faimess. Within conversational recommendation settings, this may also involve addressing new issues (e.g., identifying which types of recommendations are most prone to hallucination). This can include measuring the extent to which existing conversational agents are useful in terms of novelty, discovery, user satisfaction, and other standard metrics, including in terms of subjective experience via user studies. This can also include the design of new fairness and beyond-accuracy metrics that are specifically designed for conversational and / or fashion settings (e.g. in issues of user experience when interacting with conversational recommenders, in terms of discovery, serendipity, persuasiveness, and ultimately determine which qualities make conversational recommender systems most trustworthy).

[0054] Described below are new intervention protocols using various modalities of retrieval- augmented generation, based on knowledge grounding, KNN-LMs, and prompting. These three methods differ in terms of what is retrieved (knowledge, training samples, or items, respectively). The relationship between these three approaches is described in Table 1.Table 1: Modalities of controllable, retrieval-augmented conversational recommendation to be explored in Thrust 2.

[0055] A drawback of existing conversational recommenders is that LLMs generate items autoregressively, making it challenging to obtain (or perturb) rankings over potential recommendations. This drawback means that existing techniques to implement fairness interventions (e.g., popularity bias correction, calibration, or other use cases that call for controllability in recommender systems) become inapplicable. In some embodiments, this is addressed by re-indexing the representation of each item into a single token embedding in an LLMs' semantic space, which in turn allows us to directly estimate the probability distribution over all items without multi-step decoding. From here, multi-token items (e.g., multi-word titles) can be converted into single tokens within LLMs, and then adjust the probability distributions (e.g., using an auxiliary recommender like FISM, or by simple bias correction) over these single-token item titles, in order to achieve desired fairness or control objectives.

[0056] Figure 7 illustrates an example re-index framework (the Re-index-Then-Adapt, or RTA, framework), in accordance with an embodiment of the disclosed technology. An LLM can generate a list of recommendations, including item titles, given conversation contexts. A re-index step includes re-indexing an item title (e.g., multi-word movie titles) in the LLMs as single tokens to obtain predicted logit vectors efficiently. The logit vectors can be an output of a layer of the LLM that has not been transformed into probabilities over tokens (e.g., via an activation function). An adapt step includes adapting the recommenders towards target data distributions effectively with multiple options on the logit vectors, such as adjusting bias terms and / or recombining RecSys models with gating mechanisms.

[0057] Table 2 shows results in applying an example re-indexing framework, such as described above with respect to Figure 7, on two conversational recommendation datasets (INSPIRED and Reddit-Movies). Reindexing was implemented in this example using a GRU- based aggregator on top of Llama2, which was used to re-index multi-token movie titles into single-token movie titles as recommendation candidates (Llama2-R in Figure 7). Following these steps (e.g., the re-index and adapt steps), popularity bias can be correctly accounted for, and significantly outperform baselines on the INSPIRED and Reddit-Movie datasets. These results show that simple control mechanisms can be used to fuse semantic knowledge from an LLM with collaborative knowledge from a recommender system (e.g., Llama2+FISM in the case of Table 2).ReDIAL Reddit- V 1.5Popularity .035 .003 .008 .001FISM .065 .004 .022 .001SASRec .068 .004 .022 .001MPT .072 .004 .026 .001Mistral .082 .004 .029 .001Llama2 .094 .004 042 001ReDIAL .067 .004 .029 .001UniCRS .085 .003 .028 .001SBERT .016 .002 .003 .000Instructor .025 .002 .009 .001Llama2-R .071 .004 .055 .001+Bias .083 .004 .059 .001+FISM .094 .004 .061 .001Table 2: Results (hit5) for controllable re-indexing, compared to (1) traditional recommendation models; (2) zero-shot LLMs; (3) traditional conversational recommendation models; and (4) zeroshot dense retrievers. Llama2-R denotes the Llama2-7b model after the re -index step.

[0058] Approaches for implementing control through knowledge augmentation include augmenting conversational recommenders with domain knowledge to enhance semantic reasoning and improve item contextual knowledge. The superior performance of LLMs as conversational recommenders compared to conventional recommender systems can largely be attributed to LLMs' semantic reasoning abilities and their large-scale memories of previous recommendation conversations. In conversational recommendations, collaborative (user-item) information may be scarce, so LLM-based recommenders must leverage their semantic understanding of user queries and item contexts for recommendation. Despite the competitive zero-shot performance of LLMs in general, LLM-based recommender systems are still far from optimal in terms of accuracy (e.g., as seen in Table 2), especially in domain-specific recommendation tasks. Therefore, it would be advantageous to incorporate collaborative information for use by an LLM in recommendation applications. Discussed below are methods to use fine-tuning to incorporate collaborative information for use by an LLM and additionally enhance LLMs' semantic reasoning abilities in conversational recommendation.

[0059] Figure 8 illustrates two example supervised fine-tuning strategies, namely metadatabased fine-tuning (top) and query-based fine tuning (bottom), according to some embodiments of the disclosed technology. Supervised fine-tuning methods can simulate realistic conversational recommendation via two types of retrieval -based data augmentation. Metadata-based fine-tuning includes extracting metadata associated with items in conversational datasets (e.g., item metadata from Amazon as in Thrust 1, or Wikipedia metadata of movies in Reddit-Movie) in order to obtain domain-specific, natural language descriptions of items. The combined language-collaborative datasets (e.g., datasets described above with respect to Figures 5 and 6) can be used to learn item representations and retrieve related items. By linking item contexts with relevant items (which are often liked by users sharing similar preferences), this can encourage the model's semanticreasoning to form an association with user queries, in order to make recommendations. Querybased fine tuning includes constructing training samples by pairing user queries with previous recommendations in a similar format. This can involve using datasets from Thrust 1.

[0060] Approaches for implementing control through knowledge augmentation include implementing control through collaborative filtering and / or LLM fusion. This can leverage developments (outside of recommender systems research) in joint language-neighborhood methods to improve conversational recommendation. Recently there has been an effort to combine LLMs with neighborhood-based approaches. One such example is the KNN-LM architecture, which interpolates a pre-trained language model (LM) with a k-nearest neighbors (KNN) model, and can outperform standard language models by directly querying training examples at test time. Another such example is neighborhood+LM recommenders, which follow a similar paradigm, interpolating a retrieval based component and a neural network component for predicting the probability distribution over candidate items. Compared to the RTA framework (described above with respect to Figure 7) that requires fine-tuning pre-trained LLMs, this approach can benefit from “wisdom of the crowd” by leveraging neighborhood signals already present in training data. This complementary learning paradigm can provide a plug-and-play approach for adapting conversational recommendation models to domain-specific data without the need for pre- training / fine-tuning LLMs.

[0061] Figure 9 illustrates a neighborhood+LM recommender framework, in accordance with an embodiment of the disclosed technology. This type of model has two main components: a recommender component that outputs a probability distribution over all candidate items (e.g., similarly to the example re-index framework, described above with respect to Figure 7), and a complementary retrieval component. The retrieval component works by retrieving a set of training contexts that are similar to the test-item context, and counts the occurrences of items associated with the retrieved training contexts. These item counts are then normalized into a probability distribution over items. The two probability distributions from the recommender component and the retrieval component can then be interpolated to determine the final probability distribution.

[0062] Formally, given a training dataset T> that consists of (q, v) pairs of dialogue context q and associated items v, the example neighborhood+LM recommender uses a feature extraction function fretrievai() to map each dialogue context into a vector representation. At inference time, thesimilarity of a training context and the test context can then be calculated via a vector similarity function sim(-, ■) (e g., cosine similarity). In this way, a system can retrieve a neighborhood A' of the most relevant training contexts and their associated items at inference time, and calculate the probability of items via normalized item counts in the neighborhood.

[0063] Figure 10 illustrates a prompt-based retrieval-augmented conversational recommendation, according to an example embodiment. Approaches for implementing control through retrieval-augmented prompting include incorporating collaborative recommendation knowledge directly into prompts. Illustrated in Figure 10 is a prompt template, which can compose retrieved information into a natural language query.

[0064] The recent development of LLMs has demonstrated their abilities in complex reasoning, including their ability to deduce user preferences based on very few previous interactions in (long-tail) recommendations. However, since most LLM-based systems rely on items' semantic meaning as the sole evidence for reasoning, the LLM's reasoning can be misaligned with task-specific collaborative information, which causes recommendations to be biased.

[0065] Despite LLMs' powerful reasoning abilities in aligning LLMs' semantic understanding with the implications of user queries, “wisdom of the crowd” also exists in useritem collaborative knowledge that is not explicitly contained in user queries. Given a user u and item i (about which we are trying to assess u's preference), retrieval policies can be implemented to collect collaborative information U^°11and Tc°11, which are sets of relevant users and items respectively. The retrieval policies can be configured to identify an optimal set of user-item interactions as the supporting evidence for an LLM's reasoning. We can then construct retrieval- augmented context with Uc°u, Tc°11and the user's previous interaction histories Tpp, to prompt the LLMs' predictions on the user's preference on the item. For example, the prediction likelihood ptof whether user u likes item i can be defined by:

[0066] where C is a prompt template, such as the prompt template described above with respect to Figure 10.

[0067] A relevant interaction set T™ppcan be determined through a sequential decisionmaking process. This type of retrieval framework can be used to (a) explore a wide range of retrieval and prompting strategies; (b) learn retrieval policies that fuse additional signals from collaborative data, especially including temporal and visual information; and (c) explore the use of RL approaches to learn retrieval policies that target specific bias / faimess interventions (as in Table 2), rather than held-out accuracy.

[0068] In summary, disclosed technology associated with Thrust 2 includes: new models and methods for knowledge-enhanced conversational recommendation; models and methods for KNN-LM-based retrieval-augmented conversational recommenders; models and methods for prompting-based retrieval-augmented conversational recommenders; evaluation against existing benchmarks and datasets from Thrust 1; dissemination of research papers; and open-source code releases.

[0069] The disclosed technology contributes generally to the topic of retrieval augmented generation, and particularly to personalization and recommendation. The disclosed technology can improve performance characteristics and implement bias / faimess interventions, closing the gap between LLMs and the types of interventions that have previously been possible for conventional recommenders. Furthermore, controllability of conversational recommenders is critical in order to realize their potential impact in viable commercial applications, where there is a need to avoid issues of fairness and bias, surface specific items (e.g., sponsored or new content), calibrate outputs, and the like.Thrust 3: Evaluation Beyond Accuracy, and User Simulation

[0070] The third thrust addresses challenges related to evaluation of conversational recommender systems. This can include addressing issues of item-fairness, which can include controlling item distributions. Addressing challenges related to evaluation can include implementing strategies to conduct offline evaluation of conversational methods (especially in beyond-accuracy contexts) without having to conduct expensive (and non-repeatable) user studies, by implementing LLM-based simulators.

[0071] Discussed below are approaches for implementing better evaluation schemes for conversational recommenders, such as designing protocols to more richly assess fairness characteristics (other than just item-side fairness) and developing rich and controllable offline evaluation protocols based on user simulators.

[0072] Approaches for designing protocols to more richly assess fairness characteristics include using LLMs to evaluate user-side fairness, especially when lacking explicit user demographics. Recent work has shown that LLMs provide different recommendations when a user's demographic information is provided explicitly in the prompt (controlling for other factors). Assessing fairness as a function of user demographics (so-called “user-side” fairness) is a fundamental question in traditional recommender systems research. In conversational recommendation scenarios, such information is generally not available. However, another line of recent work has demonstrated that even in the absence of explicit information, LLMs are able to infer demographic information from user-written text and display alignment biases with respect to perspectives of different demographic groups.

[0073] Figure 11 illustrates results demonstrating fairness of recommendations across user demographics. Generally speaking, differing recommendations across different groups does not inherently constitute unfairness (preferences would naturally be different for, e.g., clothing on Amazon); instead, users are treated unfairly when the recommendation quality is different across groups. Although there are many possible quantitative definitions for the extent to which conversational recommenders exhibit inequitable treatment across different groups, a simple working definition is as follows:

[0074] where R is the set of top k items recommended to user u, and Iuis the set of items that are relevant to user u. A denotes the sensitive attribute (e.g., gender, age, political view), with user belonging to group A = a or A = b. This fairness definition states that users should get equitable treatment in the effectiveness of recommendations, regardless of which group they belong to. Results under this definition are shown in Figure 11. These results correspond to user demographics on the INSPIRED dataset, such as gender (left graph), age (middle graph), andpolitical affiliation (right graph). It can be seen that models indeed exhibit unfairness and inequitable treatment in recommendations based on different demographic groups, such as gender, age, and political affiliation. Results for Vicuna are similar to those presented in Figure 11.

[0075] Approaches to develop evaluation schemes for conversational recommenders include developing rich and controllable offline evaluation protocols based on user simulators. In addition to fairness considerations, there has been little progress on “beyond accuracy” metrics for (e.g., LLM-based) conversational recommendation. Currently, there are two predominant strategies to evaluate (conversational) recommender systems. One can either cast the task as one of held-out item prediction (in which case, the recommender should correctly guess the withheld item, or otherwise assign it a high rank or probability); or they can conduct a user study (especially when it is necessary to evaluate qualitative metrics).

[0076] Both of these approaches have significant drawbacks. Held-out item prediction is possibly unrealistic in the context of conversational recommendation. In a real recommendation dialog between two humans (e.g., selecting a restaurant to visit or a movie to watch) there is no such ground-truth item to be the held-out item. Instead, users may interact with conversational recommenders precisely because they struggle to articulate their preferences, or because they need to be persuaded to select a particular item. Conversely, user studies are expensive, and generally non-reproducible. In general, academic user studies involve synthetic users that generate content for the purpose of creating data to train models with, rather than involving data from users who are naturally requesting recommendations. User studies are more suitable for general knowledge items and domains, but are unsuitable in cases where users requiring specific knowledge or expertise may be difficult to recruit.

[0077] Synthetic users can potentially act as cost-effective proxies for real users in the evaluation of conversational recommender systems. LLMs show promise in simulating humanlike behavior, implying their ability to represent a diverse population of users.

[0078] Figure 12 illustrates five tasks designed to evaluate a property that a synthetic user should exhibit, according to embodiments of the disclosed technology. The five items illustrated in Figure 12 include choosing which items to talk about, expressing binary preferences, expressing open-ended preferences, requesting recommendations, and giving feedback. Through evaluationof baseline simulators, it is demonstrated that these tasks effectively reveal deviations of language models from human behavior, and offer insights on how to reduce the deviations with model selection and prompting strategies.

[0079] The above protocol for evaluating user simulators can be used to build controllable simulators that align with groups of real users. It is observed that simple prompting methods enhance alignment, such as prompting with interaction history and personality types. For example, Figure 13 illustrates a comparison of item ratings between humans and a simulator (GPT-4) prompted with demographic information and / or a “pickiness” personality. As shown, the simulator successfully aligns with human preferences when supplied with demographic information and the pickiness personality. In some embodiments, soft-prompting techniques are used that use embeddings as direct input to LLMs instead of text.

[0080] Following the types of combined conversational -collaborative datasets described above, rich user and item embeddings can be obtained based on modeling interaction data r(u,i) (of user u on item z) as the dot product of embeddings (e.g., simple vector embeddings of the form r(u,i) — yu-yr). Then, approaches such as prefix-tuning can be used to map yuto a set of prefix vectors to be prepended to each layer of an LLM. The mapping function / t? can be learned by fitting a loss:

[0081] where T> is the set of user-item interactions, P, is the representation of a question about item i (e.g., “what are your thoughts on item / ?”) and Rwis the representation of the desired response obtained from user u.

[0082] Reliable user simulators for conversational tasks can facilitate reproducible evaluation of personalized conversational models, which will contribute to the development of standard benchmarks covering a wide range of metrics and tasks. Arguably, these simulated users are no more of a compromise than typical user studies, in which (crowdsourced) users are only fulfilling an assigned role. Successful user simulators may eventually lead to new training regimes for personalized conversational models, e.g., a sufficiently reliable simulator may be used toaugment datasets, or optimized against in a closed-loop setting via a reinforcement learning framework.

[0083] Rather than simply trying to mimic the characteristics of a training set, user simulators are also useful to evaluate hypothetical scenarios, and to assess situations well outside of those seen during training. Practitioners deploying conversational agents could use these techniques to stress-test their systems against various scenarios; such agents could also be used to train human support agents by simulating various types of user query in a controllable way.

[0084] In summary, disclosed technology associated with Thrust 3 includes: LLM-based evaluation conversational recommenders, including against datasets from Thrust 1 and models from Thrust 2; controllable offline evaluation conversational recommenders with LLM-based user simulators; dissemination in the form of research papers; and open-source code releases.

[0085] Recommender systems, and conversational recommender systems, such as those disclosed herein, may have direct applications in e-Commerce scenarios including fashion recommendation, or recommendation of other items that involve combinations of different data modalities (visual, temporal, geographical, and the like). Additional applications of the disclosed technology include developing systems to understand multimodal clinical data, e.g. radiology images involving a combination of visual data, natural language, and the like.

[0086] The disclosed embodiments include and device and method for generating multimodal fashion recommendations. In an aspect, an example method for providing multimodal recommendations related to fashion using a conversational interface includes receiving, via the conversational interface, one or more iterated interactions, wherein the one or more iterated interactions comprise a natural language recommendation instruction and an image. The example method includes generating a recommendation response at least in part by inputting the natural language recommendation instruction and the image into a multimodal fashion recommendation model. The multimodal fashion recommendation model has been trained by: receiving one or more conversational threads, wherein each of the one or more conversational threads includes a natural language conversation between two or more users, each of the one or more conversational threads includes one or more natural language requests for a fashion recommendation and one or more natural language responses including corresponding fashion recommendations, and the one ormore conversational threads include one or more images; determining, for each of the one or more conversational threads, one or more recommended items, wherein each of the one or more recommended items was accepted as a recommendation by a user who submitted a natural language request for a fashion recommendation; and training the multimodal fashion recommendation model, wherein the multimodal fashion recommendation model is trained at least in part by using a training dataset comprising training data and training labels, the training data comprising the one or more conversational threads, the training labels comprising the one or more recommended items. The method includes causing presentation of the recommendation response on the conversational interface.

[0087] In another aspect, a device includes a processor and a memory including instructions stored thereon. The instructions, upon execution by the processor, cause the processor to receive an iterated interaction via a conversational interface and determine a recommendation task to be performed based on the iterated interaction. The instructions cause the processor to obtain a dataset based on information included in the iterated interaction, the dataset comprising one or more of conversational data, visual data, and temporal data, wherein the dataset has been generated by: receiving a plurality of conversations between two or more users and identifying one or more recommended items in the plurality of conversations, wherein the one or more recommended items are suggested by a user in response to a recommendation request by another user. The instructions cause the processor to perform the recommendation task based on an assessment of the dataset by a trained machine learning (ML) model.

[0088] Various operations disclosed herein can be implemented using a processor / controller configured to include, or be coupled to, a memory that stores processor executable code that causes the processor / controller carry out various computations and processing of information. The processor / controller can further generate and transmit / receive suitable information to / from the various system components, as well as suitable input / output (IO) capabilities (e.g., wired or wireless) to transmit and receive commands and / or data. The processor / controller may, for example, provide signals to control the operation of various components such as excitation sources and detectors that are disclosed herein. The processor / controller may be further configured to perform various method steps and computations that are disclosed in this patent document.

[0089] Various information and data processing operations described herein may be implemented in one embodiment by a computer program product, embodied in a computer- readable medium, including computer-executable instructions, such as program code, executed by computers in networked environments. A computer-readable medium may include removable and non-removable storage devices including, but not limited to, Read Only Memory (ROM), Random Access Memory (RAM), compact discs (CDs), digital versatile discs (DVD), and the like. Therefore, the computer-readable media that is described in the present application comprises non- transitory storage media. Generally, program modules may include routines, programs, objects, components, data structures, and the like, that perform particular tasks or implement particular abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of program code for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps or processes.

[0090] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0091] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0092] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random-access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0093] While this patent document contains many specifics, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

Claims

CLAIMSWhat is claimed is:

1. A method for providing multimodal recommendations related to fashion using a conversational interface, the method comprising: receiving, via the conversational interface, one or more iterated interactions, wherein the one or more iterated interactions comprise a natural language recommendation instruction and an image; generating a recommendation response at least in part by inputting the natural language recommendation instruction and the image into a multimodal fashion recommendation model, wherein the multimodal fashion recommendation model has been trained by: receiving one or more conversational threads, wherein: each of the one or more conversational threads includes a natural language conversation between two or more users, each of the one or more conversational threads includes one or more natural language requests for a fashion recommendation and one or more natural language responses including corresponding fashion recommendations, and the one or more conversational threads include one or more images, determining, for each of the one or more conversational threads, one or more recommended items, wherein each of the one or more recommended items was accepted as a recommendation by a user who submitted a natural language request for a fashion recommendation, and training the multimodal fashion recommendation model, wherein the multimodal fashion recommendation model is trained at least in part by using a training dataset comprising training data and training labels, the training data comprising the one or more conversational threads, the training labels comprising the one or more recommended items; andcausing presentation of the recommendation response on the conversational interface.

2. The method of claim 1, wherein determining a recommended item for a conversational thread comprises one or more of: approximate string matching, processing at least a portion of the conversational thread using an LLM, or following a link to a product site.

3. The method of claim 1, wherein training the multimodal fashion recommendation model comprises: extracting metadata associated with a recommended item in one or more conversational threads, and generating, from the metadata, a natural language description of the recommended item.

4. The method of claim 1, wherein receiving the one or more conversational threads comprises: receiving a series of posts; extracting a conversational subset of posts, wherein the conversational subset of posts comprises posts between two users, the two users directly addressing each other, the conversational subset of posts leading to a recommended item; and including the conversational subset of posts as a conversational thread in the one or more conversational threads.

5. The method of claim 1, wherein training the multimodal fashion recommendation model comprises: receiving temporal data, wherein the temporal data describes changes in information over time, and training the multimodal fashion recommendation model at least in part by using the temporal data.

6. The method of claim 1, wherein generating the recommendation response comprises: mapping a multiple-word phrase generated by the multimodal fashion recommendation model into a single token; generating a preliminary response, wherein: the preliminary response comprises a logit vector of values that have not been input into one or more activation functions, and wherein the logit vector has a value corresponding the single token; and adjusting the values of the logit vector to reduce popularity bias.

7. The method of claim 1, wherein generating the recommendation response comprises: receiving, from the multimodal fashion recommendation model, a first probability distribution over recommended items of the training dataset; retrieving, from the training dataset, context-similar data that has a context similar to the natural language recommendation instruction; calculating, for each recommended item of the training dataset, a total count of a number of times that the recommended item is present in the retrieved context-similar data; generating, from the total counts for each recommended item, a second probability distribution over recommended items of the training dataset; and calculating, from the first and second probability distributions a total probability distribution over recommended items of the training dataset.

8. The method of claim 1, wherein generating the recommendation response comprises: retrieving information about users and items that are relevant to the natural language recommendation instruction, wherein the information about the users and items is extracted from the training dataset;generating a prompt comprising the natural language recommendation instruction and the information about the users and items; and inputting the prompt into the multimodal fashion recommendation model.

9. The method of claim 8, wherein retrieving information about the users and items comprises: training a ML model to retrieve information when provided with the natural language recommendation instruction, wherein the ML model is configured to implement a bias or fairness intervention.

10. The method of claim 1, wherein the multimodal fashion recommendation model comprises one or more of a large language model (LLM) or a vision language model (VLM).

11. A devi ce compri si ng : a processor; and a memory including instructions stored thereon, wherein the instructions, upon execution by the processor, cause the processor to: receive an iterated interaction via a conversational interface; determine a recommendation task to be performed based on the iterated interaction; obtain a dataset based on information included in the iterated interaction, the dataset comprising one or more of conversational data, visual data, and temporal data, wherein the dataset has been generated by: receiving a plurality of conversations between two or more users, and identifying one or more recommended items in the plurality of conversations, wherein the one or more recommended items are suggested by a user in response to a recommendation request by another user; and perform the recommendation task based on an assessment of the dataset by a trained machine learning (ML) model.

12. The device of claim 11, wherein the trained ML model is a recommendation system configured to provide fashion-informed recommendations based on the iterated interaction using one or more recommendation models.

13. The device of claim 11, wherein the trained ML model is trained at least in part by: receiving a training dataset; extract fashion-related metadata associated with one or more items identified in the training dataset; and generate domain-specific, natural language descriptions of the one or more items, wherein the natural language descriptions are incorporated into the training dataset.

14. The device of claim 11, wherein performing the recommendation task comprises: implementing one or more bias or fairness intervention.

15. The device of claim 11, wherein performing the recommendation task comprises: inputting at least a portion of the iterated interaction into a large language model(LLM), wherein the LLM has been configured at least in part by: identifying an item that is represented in a semantic space of the LLM as multiple tokens; and re-index a representation of the item in the semantic space such that the item is represented by a single token in the semantic space of the LLM.

16. The device of claim 11, wherein performing the recommendation task comprises: inputting at least a portion of the iterated interaction into a large language model(LLM) to generate a logit vector corresponding to one or more fashion recommendations, wherein : the logit vector comprises a plurality of numbers, each number corresponds to one of a plurality of tokens, andthe plurality of numbers correspond to a probability distribution over the plurality of tokens; updating the logit vector such that the plurality of numbers corresponds to an updated probability distribution, the updated probability distribution being closer to a target distribution; choosing one or more tokens in accordance with the updated probability distribution; and causing display on the conversational interface of the one or more tokens as part of a natural language output17. The device of claim 16, wherein the target distribution decreases a bias or improves a fairness in the probability distribution.

18. The device of claim 11, wherein performing the recommendation task involves suggesting a product that is compatible with the information included in the iterated interaction.

19. The device of claim 11, wherein the assessment involves evaluating the dataset against long-term or periodic trends in fashion.

20. The device of claim 11, wherein the iterated interaction comprises an image and a query related to the image.

Citation Information

Patent Citations

  • Computer vision based methods and systems of universal fashion ontology fashion rating and recommendation

    US20200143454A1

  • Conversational persuasion systems and methods

    US20230162261A1

  • Multi-modal machine learning model and system

    US20230368268A1

Cited By

  • Multimodal interactive guide recommendation method and system based on knowledge enhancement

    CN121980089A