Interactive user engagement systems and methods
The system synchronizes conversational output with media elements to enhance user engagement in customer-facing environments, addressing static information and staff availability issues, and reducing labor costs by providing a dynamic and immersive experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LOVEBITE AI LTD
- Filing Date
- 2025-11-27
- Publication Date
- 2026-06-04
AI Technical Summary
Existing customer-facing environments lack dynamic and immersive engagement, rely heavily on static information sources, and face challenges with inconsistent information delivery, staff availability, high labor costs, and restricted customer autonomy, leading to disconnected user experiences.
A system that synchronizes conversational user engagement with dynamically selected media elements, using a large language model to generate conversational output and align it with media elements in real-time, ensuring seamless integration of video and text/audio outputs.
Provides a dynamic and immersive user experience by ensuring media elements are temporally aligned with conversational output, reducing staff dependency, and enhancing customer autonomy, thus improving engagement and operational efficiency.
Smart Images

Figure GB2025052607_04062026_PF_FP_ABST
Abstract
Description
[0001] INTERACTIVE USER ENGAGEMENT SYSTEMS AND METHODS
[0002] FIELD OF THE INVENTION
[0003] This invention relates to video-based, interactive, user engagement systems or methods in customer facing environments, including hospitality, retail, transportation and healthcare environments.
[0004] DESCRIPTION OF PRIOR ART
[0005] In many customer facing environments, such as restaurants, retail stores, hotels, transport hubs, or healthcare facilities, staff interactions play a central role for delivering information, answering questions, processing payments, guiding users through services, and completing transactions, and assisting with various requests. Customers are typically presented with static information sources, such as printed menus, or simple digital list, which provide only limited detail and require staff intervention for clarification or additional context.
[0006] Customer engagement technologies have also evolved to incorporate Al for providing automated services. Existing systems typically rely on interactive voice response technologies that process user inputs and return predefined responses. However, they often provide disconnected experiences that fall short to maintain user engagement throughout the interaction process.
[0007] However, across both staff based and existing Al based systems, several limitations remain:
[0008] • Information Gaps: static menus, brochures, listings etc, often lack rich, demonstrative content. Staff may not always possess detailed knowledge of every product, service, menu item, dietary accommodations, or local area recommendation, leading to incomplete or inconsistent information delivery.
[0009] • Availability and Responsiveness: During peak times, human staff may be busy, unavailable, or unevenly distributed during peak times.
[0010] • Labor Costs: Recruiting, training, and retaining staff in fluctuating-demand environments can be expensive, particularly in sectors with high turnover or variable customer flow. Basic Al systems also often require manual updates, fixed scripts, or manager review, reducing their efficiency gains. • Restricted Customer Autonomy: Customers rely on staff or rigid interface for options exploration and ordering, which can lead to longer wait times and limited flexibility.
[0011] • Lack of Interactive or immersive Engagement: traditional information sources provide no dynamic interaction or multimedia explanation.
[0012] To address these challenges, various digital solutions have attempted to bridge the gap between human service and automation.
[0013] US4547851A discloses an “Integrated interactive restaurant communication method for food and entertainment processing” in which a restaurant tabletop device is used for food ordering and entertainment. While this removes the need for staff during ordering, it relies on conventional software for menu presentation and does not support deeper interaction with individual items or dynamically generated responses.
[0014] US20200192984Aldiscloses a “System and Method for Interactive Table Top Ordering in Multiple Languages and Restaurant Management”. However, this system uses traditional management logic and requires a restaurant manager to review translated customer inputs, limiting scalability.
[0015] CA3073727A1 discloses an “Electronic Menu, Ordering, and Payment System and Method” enabling menu viewing, ordering and payment on a tabletop electronic device, but it relies on standard digital menu software and lacks the ability to provide dynamic, context dependent media responses or deeper interrogation of content.
[0016] INDEX
[0017] 100 device
[0018] 101 onb oar ding vi deo
[0019] 102 video of menu item
[0020] 103 text of conversation between the user and the system
[0021] 104 language selection feature 105 microphone activation feature
[0022] 106 text input feature
[0023] 107 software application of device, such as a web browser
[0024] 108 menu category navigation feature 109 video menu
[0025] 110 user / customer
[0026] 111 management system
[0027] 112 Al subsystem
[0028] 113 video server 114 staff
[0029] 115 payment system
[0030] 116 point of sale (POS) system
[0031] 200 kiosk
[0032] 201 video of hospitality service 202 facility informational video
[0033] SUMMARY OF THE INVENTION
[0034] The invention relates to systems that deliver synchronised conversational user engagement, in which Al generated conversation is presented together with dynamically selected media elements, so that the displayed media is temporally aligned with the conversational output, as the conversational output unfolds.
[0035] In one aspect, the invention provides a system for synchronised conversational user engagement, the system comprising one or more processing units configured to:
[0036] (a) receive user inputs including voice, text, and / or selection-based inputs;
[0037] (b) generate conversational output using a large language model (LLM) or Al model in response to the user inputs;
[0038] (c) dynamically select or generate one or more media elements relevant to the conversational output, the system having access to an operator defined media library comprising media assets;(d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output.
[0039] The term media assets as used herein refers to any resources used by the system to select, or generate, or form part of one or more media elements. Media assets may include, but are not limited to: images, videos, animations, icons, templates, audio snippets, structured data tables, or other content resources that the system may use to generate or form part of a media element. Media assets may further include media generation inputs such as prompts, templates, constraints, style guides, metadata descriptors, to enable controlled Al based media generation. Media assets are preferably associated with metadata relevant to the conversations for which the system of the invention is intended, such as price and / or calorie count for food items, as well as length of video clip, for example. In a preferred embodiment, all media elements are selected from the library designated by the operator. This assists in preventing Al hallucinations, as well as in providing uniformity of output. Preferably, priority is given to media assets in the operator-defined media library, particularly where those assets are designated as pre-approved for use in customer-facing conversational experiences.
[0040] It will be understood that the term ‘conversational output’, as used herein, refers to the spoken or textual output of the LLM or Al in a manner that a human assistant, such a server in a restaurant, might respond when answering a query, or making a recommendation, for example.
[0041] Temporal alignment, and the associated term, ‘temporally aligned’, as used herein, refers to any coordinated presentation in which a portion of one or more media elements is displayed, transitioned, initiated or updated at an appropriate time relative to a portion of the conversational output. Temporal alignment may apply whether the conversational output is visual (text) or audible (speech), or a combination thereof. Temporal alignment may be achieved using synchronisation data, which includes any information, metadata usable to determine, control, or adjust the relative timing of conversational text, conversational speech, or any portion of media playback. Synchronisation data may be generated before, during or after the LLM output is generated, and may be embedded within or associated with the conversational output or generated separately. Synchronisation data may be harvested from any aspect of the overall process involved in generating the conversational user engagement, and may be processed in a manner suitable to enable the desired synchronisation of media content the conversational output. It will be appreciated that the synchronisation will preferably be in a manner that the user might expect to see in an instruction presentation, for example.
[0042] In some implementations, temporal alignment is achieved without explicit synchronisation data, using approximate or implicit triggers such as conversational start time, keyword detection, entity recognition, semantic grouping, heuristic timing rules, or predicted audio or text duration.
[0043] In various embodiments, the system further includes a latency resolution mechanism that enables provisional playback of a media element while synchronisation data is still pending, and subsequently overrides, replaces or adjusts the provisional playback once synchronisation data becomes available, thereby maintaining a seamless presentation despite variable system latency. For example, the system may establish the first media element to be shown before sufficient data has been collected to determine the total and sequence of all media elements, and the system may commence showing the first media element while establishing the final sequence. The system may then commence the remainder of the sequence in a timely fashion, once determined.
[0044] The invention may further incorporate additional functionalities, such as intent management modules, contextual retrieval modules, presence detection, personalisation, and user interface features that enable multimodal interaction (voice, text, touch) for controlling the conversational session or performing transactions.
[0045] The system may be implemented in general purpose devices, websites, kiosks, customer-facing service terminals, mobile devices, embedded displays, hospitality ordering systems, retail environments, entertainment platforms, or itinerary based experiences in which media elements may represent locations, events, or attractions.
[0046] BRIEF DESCRIPTION OF THE FIGURES
[0047] Aspects of an implementation of the invention will now be described, by way of example(s), with reference to the following Figures, which each show features of an implementation of the invention:
[0048] Figure 1 shows a video synchronisation service coordinating an LLM service, a text to speech engine, a video database, and a synchronised front end presentation layer.
[0049] Figure 2 shows a device with the user interface of the system displaying a prompt to engage the microphone.
[0050] Figure 3 shows a device with the user interface of the system displaying a video of a menu item and text of the conversation between the user and the system.
[0051] Figure 4 shows a device with the user interface of the system displaying a video of a menu item taking up a significant portion of the screen.
[0052] Figure 5 shows a device with the user interface of the system displaying a prompt to swipe to view videos of menu items.
[0053] Figure 6 shows a device with the user interface of the system displaying a video of a menu item with alternative menu items to either side of the displayed menu item and text of the conversation between the user and the system displayed below.
[0054] Figure 7 shows an overview of the user interface.
[0055] Figure 8 shows a user interface in which a media element associated with a menu item is displayed.
[0056] Figure 9 shows a user interface in which the conversational Al generates natural language responses that appear beneath the video presentation of an item being discussed.
[0057] Figure 10 shows a user interface in which menu categories can be displayed not only as a video carousel but also as a list representation.
[0058] Figure 11 shows a user interface where a user may place an order by interacting with the Al persona through voice or text, or by using touch controls.
[0059] Figure 12 shows a user interface illustrating an option specification panel in which a user can specify item customisations such as sauces, side dishes, cooking preferences, or other defined choices. Figure 13 shows a user interface where a user may review items added to the basket, modify existing selections, or add new items by interacting with the video waiter through voice, text, or touch.
[0060] Figure 14 shows the relationship between the user, the device and the various subsystems of the system including: the management system, the Al-based voice processor, the video server, the staff, the payment system, and the POS system.
[0061] Figure 15 shows the relationship between the user and the device on which the system is accessed
[0062] Figure 16 shows the relationship between the user, the device on which the system is accessed and the management system.
[0063] Figure 17 shows the relationship between the user, the device on which the system is accessed, the management system, the Al-based voice processing system, the video server, and the staff.
[0064] Figure 18 shows the relationship between the user, the device on which the system is accessed, the management system, the Al-based voice processing system, the video server, the staff, the payment system, and the Point of Sale (POS) system wherein the payment system may be either accessed by the management system or by the
[0065] POS system.
[0066] Figure 19 shows a kiosk with the user interface of the system displaying a video of a menu item and text of the conversation between the user and the system.
[0067] Figure 20 shows a device with the user interface of the system displaying a video related to a hospitality service and the text conversation between the user and the system.
[0068] Figure 21 shows a device with the user interface of the system displaying a video related to a hospitality product and the text conversation between the user and the system.
[0069] Figure 22 shows a device with the user interface of the itinerary linked concierge video system.
[0070] Figure 23 shows a device with the user interface of the itinerary linked concierge video system.
[0071] Figure 24 shows a device with the user interface of the itinerary linked concierge video system. Figure 25 shows a device with the user interface of the itinerary linked concierge video system.
[0072] DETAILED DESCRIPTION OF THE INVENTION
[0073] This Detailed Description describes various implementations of the invention and is divided into the following sections:
[0074] 1. Conversational Al with synchronised media
[0075] 2. Example implementation: restaurant, kiosk, and hotel concierge examples
[0076] 3. Itinerary linked concierge video system
[0077] 4. Dynamic kiosk interaction interface
[0078] 5. Al driven social media content module
[0079] 1. Conversational Al with synchronised media
[0080] 1.1 Overview
[0081] A conversational experience that synchronises Al responses with dynamically selected short-form video content is provided. An integrated conversational Al and video-synchronisation architecture is now presented. This technical architecture forms the core of the system and governs how speech recognition, natural -language generation, video retrieval, and front-end playback coordination occur in real time. The following section provides an overview of these technical components before describing their use in restaurant, kiosk, hotel, and other customer-facing environments.
[0082] The system includes an input module configured to receive user interaction signals, including, but not limited to: voice audio, text entry via physical or virtual keyboards, and selection-based inputs through touchscreens, graphical user interface elements such as on-screen prompt buttons or soft keys, physical buttons, or gesture recognition systems. The input module connects directly to the processing units, such as to be capable of providing a continuous stream of user interaction data the processing units.
[0083] The input module may also incorporate external sensor data from a variety of modality. Visionbased sensing may include: computer vision-based tracking of hand or body gestures, eye gaze estimation, and head pose estimation. Depth sensing may also be supported and may be achieved using computer vision techniques, time of flight sensors, LiDAR, or other depth sensing techniques. Non vision sensor data may also be processed, such as accelerometer data, , or environmental sensors. The input module may also be configured to detect voice activity.
[0084] User inputs capable of being interpreted as language, or which generate a language response or instruction, are analysed using natural language processing algorithms and / or semantic analysis techniques. This analysis may include intent recognition, sentiment analysis, and / or contextual understanding to determine the user's specific needs regarding products, services, or general enquiries. This analysis is performed in real-time, allowing for immediate system responses to user queries. As used herein, an ‘immediate response’ is one that is delivered with preferably no more than 3 seconds’ delay after the system determines that the user has finished the present input and is awaiting a response. The delay is more preferably no more than 2 seconds, and more preferably one second, or less.
[0085] 1.2 Synchronisation pipeline
[0086] Figure 1 shows a diagram illustrating a video synchronisation service for integrating short videos with Al-generated text and / or voice response in a conversational system. The figure illustrates the interactions among the main LLM service, the video synchronisation service, a text to speech system, a video database, and a synchronised front end presentation layer or output module.
[0087] The main LLM (Large Language Model) service receives a user's input, such as voice, text, interactive prompt, such as prompt bubble, button press, or gesture etc, and determines or generates an output message, which forms the basis of the conversational content. The video synchronization service analyses the LLM output and other information available (including details of menus, full conversation history, items stored in an order basket or wish list, or other contextual elements). Using this information, the synchronisation service queries the video database and injects timestamps or markers into the text output indicating which videos should be played at specific points and for how long during the conversation. The video database contains details of available videos like ID, location, description, length, and format. The text to speech system converts the timestamped enhanced text output into a corresponding voice output. The user interface receives the synchronized text, voice, and associated video instructions, and provides a seamless experience in which videos elements are aligned with the conversational flow defined by the LLM’s output. In preferred embodiments, the “video database” illustrated in Figure 1 corresponds to an operator defined media library comprising media assets. The term media assets as used herein refers to any resources used by the system to select, or generate, or form part of one or more media elements. Media assets may include, but are not limited to: images, videos, animations, icons, templates, audio snippets, structured data tables, or other content resources that the system may use to generate or form part of a media element. The media library may additionally contain media generation inputs, such as still images, textual descriptions, templates, brand assets, or style parameters, which may be used by an Al based media generation model to create new media elements. In such embodiments, any dynamically generated media element is produced by operator-provided inputs and is therefore treated as a media element “retrieved” from the operator- defined media library. This configuration reduces latency by restricting retrieval to a controlled set of pre-indexed or locally cached assets, and ensures consistent branding, style, across all displayed media elements. Preferably, priority is given to media assets in the operator-defined media library, particularly where those assets are designated as pre-approved for use in customerfacing conversational experiences.
[0088] Where a conversational output requires a media element that is not directly available in the operator-defined media library, the system may generate a suitable media element by transforming, adapting, or using only a portion of one or more media assets stored in the library. The system may additionally combine two or more media assets to produce a composite media element suitable for synchronised presentation. If no suitable media element can be created using assets from the media library, the system may, in certain embodiments, obtain media elements from one or more external sources.
[0089] In alternative embodiments, the system may permit conditional retrieval of media from external sources, such as online repositories or remote content services, for example when no suitable media element or media generation input exists within the operator-defined media library. In further embodiments, externally retrieved media may be used directly, or may be combined with operator defined assets to produce hybrid or composite media elements. In yet further embodiments, the system may employ an Al model to generate entirely new media elements based on operator defined inputs, external data sources, or a combination thereof. The synchronisation pipeline of Figure 1 may operate on pre-existing media assets, dynamically generated media elements, fallback external media elements, or any combination of these.
[0090] The LLM output used to drive synchronisation may be in the form of text, tokens, structured data, vectors, or any representation suitable for downstream analysis.
[0091] Any reference to a “large language model”, or “LLM” is also intended to encompass any suitable Al or machine learning model capable of generating output, or interpreting conversational input, including but not limited to: transformer models, recurrent neural networks, encoder / decoder architecture, multi-agent systems, or any other hybrid or equivalent Al architecture.
[0092] The rendering of the video can be performed entirely on the client device, on one or more server side components, on an edge computing node, or through a distributed or hybrid arrangement.
[0093] In some implementations, timestamps or markers may also be omitted, approximate, inferred, probabilistic, or generated by an alternative scheme, rather than precise timestamps, and may be triggered by token patterns, or other detected conversational markers.
[0094] In some embodiments, where timestamp generation may create latency, the system may incorporate a latency-resolution mechanism. The latency-resolution mechanism enables the rendering module to immediately begin rendering text and voice output while initiating playback of a provisional video, typically the first referenced item in the LLM output, before timestamped instructions arrive from the video synchronization service. When the video synchronization service transmits the finalized synchronization schedule, the rendering module dynamically overrides the provisional playback and inserts a precise video sequence at the correct temporal positions within the ongoing audio / text narration.
[0095] In alternative implementations, the system may initiate provisional playback using a non-item specific media element while awaiting synchronisation data. Such a media element may include for example a stock video, brand themed animation, introductory clip, or any other default media asset. This approach ‘packs’ the gap created by synchronisation latency and maintains visual continuity until an appropriate item-specific media element is determined or selected by the video synchronisation service.
[0096] There may be some latency, as the conversational output may be sent to a text to speech service, which may send back synchronised voice and text data files in incremental or chunked segments which then need to be processed by the video synchronisation service as they become available.
[0097] To reduce the effects of latency, the video synchronisation service may include or more methods to determine and enable the immediate display of an initial or provisional media element while waiting for full video synchronisation data. This may be achieved by several methods, including the front end or back end recognising the first item name for which a video is available, or by using an auxiliary Al process within the video synchronisation service that analyses the conversational output as soon as it becomes available for purposes of identifying a suitable a provisional media element.
[0098] Once the synchronisation service timestamps arrive at the front-end application, they may override the initial or provisional media and present the new synchronised media sequence.
[0099] In various embodiments, the system may support various display configurations, including singlescreen setups where video and text are displayed concurrently on the same display surface, whether in separate layout regions or with text visually overlaid on the video. Alternatively, the system may also employ multi-screen arrangements where media elements and conversational output are distributed across separate displays.
[0100] In various embodiments, the presentation layer may include smooth transitions between media elements, ensuring a fluid viewing experience without jarring cuts or pauses. The system continuously monitors playback performance and can dynamically adjust synchronization parameters to compensate for any processing delays or network latency.
[0101] In various embodiments, the user interface may provide an interaction mechanism for expressing feedback, such as “like” “dislike” control or favourite indicator. The control may be activated through touch input, gesture input, or voice commands (e.g. “I like this” or “add this to my favourite”). The visual presentation of like control may take any suitable form, such as a heart icon, star icon, or animated indicator that updates when activated. User likes may also be recorded as preference signals and may be used to influence future conversational output or prioritise media selection within the system. These features may be configured or disabled based on operator preferences.
[0102] Additionally, the system may include a feedback mechanism that evaluates user engagement with the presented media elements. This mechanism may track such aspects as eye movement, interaction time, and explicit feedback to determine which media elements were most effective in addressing a user's query. This information can then be used to refine the selection of future media elements, for example.
[0103] Personalization features may also be included that adapt the selection of media elements based on user history and preferences. The system may optionally maintain user profiles that record past interactions and preferred information formats, allowing for increasingly tailored responses over time.
[0104] As an example, the user submits the input query: “What vegetarian mains do you have?”, the LLM processes the input and generates the natural language response: “Our vegetarian mains include Cauliflower Steak, with smoky cashew red pepper dip, and Verdura Pizza, topped with torn mozzarella and baby plum tomatoes”. Immediately upon receiving the LLM output, and prior to receiving timestamp instructions from the video synchronization service, the front end initiates its latency-resolution mechanism and identifies the first referenced menu item associated with an available video (in this case, the Cauliflower Steak) and begins playback of the corresponding video when the spoken or displayed output reaches the word “Our”. Once the video synchronisation service completes its semantic analysis and timestamp generation, it transmits the precise synchronization data to the front-end application. The front-end then overrides the provisional playback schedule in real time based on the received timestamps: the Cauliflower Steak video continues until its designated termination point, after which the Verdura Pizza video begins on the word “Verdura,” as specified by the timestamped synchronization markers embedded or associated with the Al-generated response. This process results in a seamlessly coordinated presentation in which the system’s text, voice, and video outputs remain fully synchronized. It should be noted that the synchronisation of video with the output language may be adjusted in any fashion desired, and will normally be such as to give the appearance of the video matching the first word of the relevant description. Given loading times, the video may be sent to the display such that the screen switches to the video as the first word of the relevant description is output, and may be phased to be slightly ahead or behind the beginning of the word, if desired.
[0105] In alternative examples, the user may also interrupt playback with a follow-up query, causing the system to dynamically cancel, pause, or reprioritize videos. In another scenario, multiple simultaneous videos may be displayed using overlay formats. The system may also dynamically replace a video with a higher-relevance media element based on real-time user attention signals.
[0106] In certain embodiments, the system is provided as a website or web-based application accessible through a standard web browser. The web-based implementation may present the same conversational interface, synchronised media elements, and Al-driven guidance described herein, and may enable users to browse information, explore menu items or services, place orders, make reservations, or otherwise engage with the operator’s offerings. In some embodiments, the webbased interface is accessed via a QR code displayed within a restaurant or venue, causing a user’s mobile device to load a browser-based session that connects to the conversational Al platform. In further embodiments, the website may be accessed directly by users outside the venue, including through links presented in search results, mapping services, or third-party discovery platforms. The web-based implementation may therefore function as an adjunct to the operator’ s conventional website, providing conversational interaction, dynamically selected or Al-generated media, and synchronised audiovisual guidance within a browser environment.
[0107] In certain implementations, the web-based platform operates as an information interface rather than an ordering interface. In such embodiments, the conversational Al may answer questions and may provide recommendations accompanied by synchronised media elements without necessarily initiating a transaction. Optional prompts for bookings, takeaway orders, or pre-orders may be presented through conversational responses, Al-generated interactive prompt, or selectable buttons on the media cards, but such transactional features are not required. In this configuration, the system functions as a general-purpose conversational information assistant capable of supporting users who are browsing an operator’ s offerings remotely, including users who are not presently inside the venue.
[0108] The term “interactive prompts” as used herein includes any system-generated selectable interface elements such as suggested questions, recommended actions, or follow-up options. These may be presented as bubbles, cards, tiles, banners, icons, or any other UI format.
[0109] In some implementations, the system may present Al-generated interactive prompts that suggest follow-up questions, recommended actions, or next conversational steps. These interactive prompts are configured to increase user engagement. As an example, when integrated with booking or reservation systems, the conversational Al may use such interactive prompts to guide users toward reservation actions, optionally checking availability or presenting alternative times or options in response to user queries. This configuration enables the system to assist users in asking for recommendations, completing bookings or reservations more efficiently, providing a smoother and more responsive user experience within the browser environment.
[0110] 1.3 Conversational Al system optimised for a specific application
[0111] The conversational engine supporting the synchronisation pipeline may also be tailored for a specific application, such as food and hospitality data. In some implementations, the LLM may be configured with a context defining a persona or set of attributes associated with a waiter, server, concierge, or equivalent hospitality role, and interprets the user’s inputs, whether spoken, typed, gesture-based, or otherwise, as prompts that drive the generation of real-time conversational output. The contextual data supplied to the LLM may be dynamically updated using operational information such as current menu data, live promotional offerings, stock or availability changes, user preferences, and behavioural feedback derived from the ongoing session or prior interactions. The term ‘avatar’ is used interchangeably herein with the term ‘persona’, and generally indicates an animated figure, or a still representation of an animated figure, that accords with preferences of, typically, the establishment offering the present service. In a preferred aspect, the avatar is a close facsimile of a brand character owned by the establishment, although any representation may be used.
[0112] In some implementations, the contextual updates may be approximate, inferred, probabilistic, or partially derived from prior conversation history, or a combination thereof. Context may be delivered to the LLM in the form of structured metadata, tokens, embeddings, or any suitable representation. This dynamic context adaptation allows the LLM to adjust its persona, tone, recommendations, and response structure in real time as underlying operational conditions evolve.
[0113] In some implementations, the LLM may also employ retrieval augmented generation (RAG) techniques, in which one or more operator’s approved sources of truth are used to constrain or enrich the model’s output. Such sources of truth may include, for example: menu descriptions, allergen lists, pricing data, stock or availability information, promotional documents, or other curated documents, including PDF based resources. The LLM may be instructed to prioritise or exclusively rely upon these sources of truth for certain classes of requests to minimise hallucination and ensure factual accuracy.
[0114] RAG may also reduce token usage by allowing the model to generate responses based on precise retrieved fragments rather than maintaining large volumes of descriptive text within a prompt. The use of RAG therefore provides both computational efficiency and improved factual reliability.
[0115] The term “contextual data”, as used herein, may refer to both these static or semi-static data sets provided by the one or more operators, such as menu information, pricing, allergen data, item availability, configuration parameters, or branding information. Contextual data may additionally include user-generated data arising from interaction within the system, such as conversation history, basket contents, inferred preferences, and behavioural feedback.
[0116] In a further aspect, the system may include a secondary Al component or verification engine configured to evaluate the primary LLM output against contextual data and to request regeneration or correction where inconsistencies or errors are detected. This conversational framework may operate in conjunction with, and / or independently of, the synchronisation pipeline described above. For example, the LLM may refine its recommendations, re-query updated menu information, or modify its service persona even while the synchronised presentation layer is rendering videos, managing provisional playback, or applying timestampbased overrides. The result is a responsive, context-aware conversational experience in which the system adapts both its language output and its associated media selections according to the realtime operational state of the restaurant, bar, hotel, or related hospitality environment.
[0117] 1.4 Brand specific voice agent
[0118] A further component of the conversational engine supporting the synchronisation pipeline may include a brand specific voice agent configured to deliver consistent, recognisable audio output across various interactions. The Large Language Model (LLM) generates outputs which are rendered through a synthesised voice corresponding to one of the defined brand-specific source profiles. The system may deploy the synthesised voice across a variety of channels, for example kiosks, telephone-based systems, web interfaces, mobile applications, in-restaurant devices, and in-hotel room devices, ensuring consistent cross channel behaviour. In a preferred embodiment, the output is delivered by an avatar that may also be brand specific, and is preferably Al-driven.
[0119] In some implementations, the LLM may receive brand-specific marketing language or stylistic guidance as contextual input, enabling the generation of responses that remain consistent in tone, phrasing, and persona. The brand voice may be selected, adapted, or updated dynamically based on user preferences, operational context, or real-time service requirements, allowing the conversational experience to maintain a coherent and recognisable identity aligned with the associated brand.
[0120] The Al voice synthesis module may be trained on voice samples originating from multiple distinct sources, each sample set being characterised by user-recognisable vocal attributes associated with a particular brand, restaurant chain, product line, or defined brand personality.
[0121] The following sections illustrate how the system is implemented in various customer-facing environments. 1.5 Automated onboarding
[0122] In some implementations, the system may further include an automated onboarding module configured to streamline initial setup, configuration, or deployment of the conversational Al environment.
[0123] As an example, the automated onboarding module may guide an operator or organisation through a sequence of configuration steps, such as including providing menu data, uploading media assets, defining brand preferences, linking social media accounts, or other operational parameters.
[0124] 2. Example implementations: restaurant, self-serving kiosk, and hotel concierge examples Illustrative examples of how the synchronised conversational Al architecture described in Section 1 can be deployed in real customer-facing environments are now described. These implementations demonstrate how the same underlying technical framework adapts to restaurants, kiosks, hotels, and other service contexts, while preserving the core synchronisation, media selection, and conversational capabilities.
[0125] 2.1 Restaurant implementation
[0126] In this restaurant implementation, a synchronisation pipeline as described in Section 1 governs how conversational output, media selection, and video playback are generated, aligned and presented during customer interactions.
[0127] The system provides a video-driven and Al-linked ordering and information solution with a video menu, to be accessed via a device 100 by a customer 110. The customer can converse through speech, typed text, and / or interactive prompt, for example, with the system providing an intelligent answer to queries and orders with relevant video content and order assistance provided by a background Al. In a preferred embodiment, the system integrates with a Point of Sale (POS) 116 system, allowing customer orders to be accepted and fulfilled by the establishment. The POS 116 system may be provided as part of the system of the invention, or the system of the invention may be adapted to operate with a suitable POS. The system simultaneously reduces customer-staff interactions and facilitates an increase in sales- per-cover by providing videos of menu items directly to the customer. The system can include data on such items as the ingredients of the menu items as well as their sales performance and popularity in the restaurant, allowing it to provide suggestions to the customer. These suggestions may be unsolicited by the user and hence an opportunity for the system to upsell items. Alternatively, the suggestions can be provided by the system at the user’ s request for a more personalised experience.
[0128] The system is not limited to food and / or drink ordering and can be adapted for other uses in the hospitality sector, or other sectors altogether. The system uses natural voice, text, and one or more videos to communicate with the customer.
[0129] The following figures depict example implementations of the system components and illustrate how the synchronised conversational engine controls user interaction, video presentation, and AI- generated responses within practical deployment contexts.
[0130] Figure 2 shows the customer ordering system running on a device 100. In Figure 2 the centre of the screen of the device 100 displays an onboarding video 101 that provides an overview of the functionalities of the system.
[0131] The display may be a user’s phone or tabletop tablet, or other suitable display means. The artificial intelligence may operate as a local standalone system, such as DeepSeek Rl, or as a remote service such as a ChatGPT-based platform or equivalent LLM / speech synthesis service.
[0132] In an alternative aspect, the present invention provides an interactive video-based ordering and information system adapted to allow customers to view a video menu, receive video as an integral part of the response to questions or requests from the customer, optionally place orders, whether for dining-in or for collection or for delivery, optionally make table reservations, optionally pay bills, and optionally request assistance via a mobile device, a tablet presented to the customer, a website or a kiosk. The system preferably serves a video-based interactive guide, either complementing or replacing static menus and descriptions with contextually relevant videos that adapt in real time. Figure 3 shows the device during conversation between the customer and the system. The device displays a video of a menu item 102, below which there is text of the conversation between the user and the system shown at 103, with user input provided either through the microphone activation feature 105, or via the text input feature 106. Any Al-generated response to the user’s input is also shown at 103, and can also be spoken aloud if the system has audio. The dynamically generated text-based dialogue (103), when present, can serve to confirm that any spoken input has been interpreted correctly by the Al, and may replace or appear in parallel with spoken output generated by the system. The text and / or spoken response may incorporate meaningful contextual data such as menu details, availability, or personal suggestions based on the media elements presented and / or any previous interactions or as determined suitable by the Al. The Al-generated response can also provide actionable insights, such as alternative recommendations, special offers, or ingredient details. The Al may also dynamically generate interactive output while users are interacting with the media content.
[0133] Figure 4 shows the user interface without the text of the conversation between the user and the system at the bottom of the screen. The dialogue box 103 may be suitably designed to disappear after a period of inaction, if the user asks the Al to hide the dialogue box 103 or if an onscreen button (not shown) is activated. When the dialogue box 103 is hidden, this permits the video of the menu item to fill a greater portion of the screen for easier viewing.
[0134] Figure 5 shows informative text explaining to the user that the video menu can be accessed by swiping the screen, allowing for intuitive and seamless browsing. The same effect may also be achieved by incorporating arrow buttons to allow toggling through the available videos.
[0135] Figure 6 shows text of a conversation between the user and the system as well as a video of the recommended menu item as a result of the conversation. Swiping sideways allows the user to access other videos of menu items. The video response is also synchronised with text responses, allowing customers to visually assess their options in real time, such as asking the system to display certain items. Here, the system has identified key menu items, such as Pumpkin Flan Brulee or chocolate Brownie, which are then highlighted in the generated text. The text description can also take into account user’s past interactions, and browsing history, ensuring that the displayed text is relevant to the user. Hence first-time users may have simplified explanations while returning users may be shown detailed ingredient breakdowns, such as allergens that is of interest to the user.
[0136] Figure 7 displays the user interface features of the system. An onboarding video is represented at 101, a video of a menu item 102, text of a conversation between the user and the system 103, a language selection feature 104, a microphone activation feature 105, a text input feature 106, a software application for the device 107 (in the case of Figure 7 a web browser is illustrated), and a menu category navigation feature 108.
[0137] The timing of the displayed videos and the conversational text or audio follows the synchronisation and latency -resolution mechanisms described in Section 1, ensuring that media elements remain aligned with the LLM’s conversational flow even when user interactions interrupt or modify the dialogue.
[0138] Additional examples of the user interface are also illustrated in Figures 8 to 13.
[0139] Figure 8 illustrates an example of the video waiter interface in which a media element associated with a menu item is displayed. A collapsible menu icon (hamburger menu) may be provided for category selection, language selection, or other navigation functions. A category navigation element enables the user to swipe or scroll through available categories or through videos associated with items within a category. The interface may further include a “like” or “save” control for adding the item to a personalised list, a display area for the name, description, and pricing of the item, and a playback region in which the corresponding video is rendered. Certain embodiments may also present prompt bubbles that suggest conversational prompts a user may select, and touch, text, or voice input controls enabling the user to communicate with the Al persona. Figure 9 shows an interface in which the conversational Al generates responses that appear beneath the video presentation of the item being discussed. As the Al speaks or generates text, relevant media elements may be highlighted or surfaced within the interface, and videos may be synchronised with the spoken words or textual output as described in Section 1. Portions of the text may be emphasised or visually linked to specific media elements to reinforce the synchronised presentation.
[0140] Figure 10 shows an interface which menu categories can be displayed not only as a video carousel but also as a list. The list may include still images, thumbnail videos, previews, or other simplified media elements. By selecting an item from the list, the corresponding media element (e.g., video) may be displayed in the main presentation area. This allows for flexible browsing of menu items and improved accessibility.
[0141] Figure 11 shows an interface demonstrating how a user may place an order by interacting with the Al persona through voice or text, or by using touch controls. The media element corresponding to the selected item may continue to play or remain displayed during the ordering interaction, and the Al persona may confirm the order or provide additional details as needed.
[0142] Figure 12 shows an interface with an options panel in which the user can specify item customisations such as sauces, side dishes, cooking preferences, or other choices. These selections may be made via voice, text, or touch input, and the Al persona may guide the user through the customisation workflow.
[0143] Figure 13 shows an example of an order basket interface. The user may review items added to the basket, modify existing selections, or add new items by interacting with the video waiter through voice, text, or touch. The order basket may be displayed visually on screen or read aloud by the Al persona. The user may confirm or adjust the order through any of the supported input modalities, and the interface may reflect updates in real time.
[0144] As seen, the user interface may employ any suitable arrangement of media displays, list views, carousel views, navigation menus, collapsible menu elements (including hamburger style menus) for accessing categories or settings, conversational text panels, feedback indicators, or order- management features. The appearance, positioning, and behaviour of interface elements may vary according to device capabilities, operator configuration, accessibility requirements, or branding preferences.
[0145] In alternative embodiments, the user interface may also present media elements using a grid layout, tabbed layout, multi-panel layout, split screens, or augmented-reality overlays, and may support interaction via touch, stylus, voice, gesture, controller input, or remote device.
[0146] Figure 14 shows the interaction between different elements of the system. The user 110 interacts with the device 100. The inputs provided via the software on the device are passed to the management system 111 which interacts with various subsystems as appropriate. The subsystems include an Al-based voice processing subsystem 112, a video server 113, staff 114, a payment system 115 and a point of sale (POS) system 116.
[0147] Figure 15 illustrates the interaction between the user 110 and the device 100 on which the system is accessed.
[0148] The user can interact and communicate with the video waiter in various ways, for example:
[0149] • The user can manually navigate through categories and swipe dish videos, in order to view videos dishes of each dish on the menu, and to view information on each dish.
[0150] • The user can speak to the video waiter by pressing the microphone button.
[0151] • The user can also type a message to the video waiter
[0152] • The video waiter may display prompts that the user can either press or respond to. For example, “Would you like to see our cocktails?” with Yes please / No Thanks prompts.
[0153] Video Waiter communicates with the user in various ways, for example:
[0154] • The Video Waiter will speak to the user in natural voice using the loudspeaker on the phone or other device,
[0155] • The Video Waiter displays messages as text on the screen for the user to read,
[0156] • The Video Waiter displays videos that are synchronized with the messages from the Video Waiter and other information relevant to the conversation. The Video Waiter may display or speak prompts to the user.
[0157] Figure 16 illustrates that the device 100 interacts with the management system 111 in a two-way exchange of information. Information provided from the device 100 to the management system I l l is related to the user inputs, while information provided by the management system 111 to the device 100 is relayed to the user 110.
[0158] The management system controls all aspects of the video waiter’s configuration, including which restaurant, menu and dish details, branding details, videos to be served, table numbers supported, which languages are supported.
[0159] The management system includes all relevant instructions for the Al system, including the Al’s role and persona, its voice accent, tone of voice, the limitations of what questions it should attempt to answer, instructions on recommendations, what to do in various circumstances (i.e. guest has allergies, guest wants to make a complaint, etc)
[0160] As an option, the management system can synchronise with the POS system and holds relevant details of the POS configuration including menus, products, prices, service charges, opening hours, and hours that each menu or item is available.
[0161] As an option, the management system can integrate with payment systems, so users can pay their bill with assistance from the video waiter.
[0162] Figure 17 expands on Figure 16, highlighting that the management system 111 then interacts with an Al-based voice processing system 112, a video server 113, and staff 114. Interaction with the Al-based voice processing system 112 allows the text or speech input provided by the user 110 via the device 100 to be interpreted appropriately. Once the user input has been interpreted, the management system 111 can then call on the video server 113 to provide relevant videos that satisfy the user’s request or the management system 111 could alert the staff 114 to the user input, for example. This could be useful in a situation where the user 110 would like to make a complaint or requires a staff member’s assistance at the table.
[0163] The management system links to an Al system, that provides the natural language communication and other Al services that are relevant to the service provided by the main conversation. The management system provides the Al with all relevant information for it to act as required. The management system links to a video server. As an option, the management system can also initiate voice or video calls with staff and send messages to staff.
[0164] Figure 18 further builds on Figure 17, highlighting the interaction between the management system 111, the payment system 115, and the POS system 116. Upon request of the user, the management system 111 can provide the user 110 access to the payment system 115 via the device 100 - allowing payment to be completed without the need for staff. Furthermore, when the user 110 requests a menu item to be ordered, the management system 111 passes this order to the POS system 116 of the establishment for fulfilment by the establishment. It is possible for payment system integration to be made via the POS system 116 as opposed to direct integration with the management system 111.
[0165] As an option, the management system can synchronize with a POS system and hold relevant details of the POS configuration, including menus, products, prices, service charges, opening hours, and hours that each menu or item is available. The management system passes order details to the POS system via an integration, and receives relevant back information from the POS. The system may also work without integration with a POS system, with orders sent to a tablet, or sent via email or other communication method. The system may also work by providing links to third-party systems, including ordering, payment, table booking and ticket ordering systems, via buttons or by the Al instigating a link.
[0166] The invention could be implemented without one or more of the subsystems. For example, a lack of a payment system 115 could be overcome by the establishment offering a manual payment system upon the customer leaving. The payment system may be linked to the POS system, or may be linked only to the Management System. As an option, when the user requests the bill, they may be presented with the option to pay at the table on their mobile. If they wish to pay on their mobile they are redirected by the video waiter to the payment page for their table.
[0167] System Components and Structure
[0168] User Interface: The system, also referred to herein as a ‘video waiter’ can be made accessible on customers' mobile devices via QR codes, NFC, or direct links, or may be accessed by tablets presented to the customers at the table, for example. It may also be integrated into mobile apps, websites, or physical kiosks. The user interface can include the following:
[0169] • A language selection feature.
[0170] • An onboarding video for new users or users who would like the system explained.
[0171] • A category navigation feature from which the user can swipe to navigate through categories. Selection of a category displays menu items relevant only to that category. Example categories could include: starters, mains, specials, desserts, soft drinks, and cocktails, and may be further divided by such considerations as dietary preference, including vegan and kosher, for example.
[0172] • A video box that automatically resizes to fill the maximum available area on the screen at any one time. A manual sideways swipe of the video box summons the next video box.
[0173] • A speech box in which user speech or text input is displayed, confirming what is understood by the video waiter. The response of the video waiter can also be displayed in the speech box, as well as in audio form. There is a visual indication, such as underlining, of which text in the speech box is being played in audio form.
[0174] • A voice input / microphone button.
[0175] • A text input button.
[0176] • Prompt bubbles, which may optionally be scrollable.
[0177] In some embodiments, the user interface may allow the user to switch between a video representation of a menu item and an alternative list-based representation. The list view may include still images, reduced-resolution videos, thumbnails, or other simplified media elements associated with each item. A user may select an item within the list, by touch, voice command, or other input, to view the full media presentation, obtain more detailed information, or add the item directly to an order. These interaction controls are optional and may be adapted or extended depending on operator requirements or device capabilities.
[0178] It is preferred that the Al employed by the system has multilingual capabilities that extend beyond basic language selection to include support for example for regional dialects and / or localized language variations. A preferred Al enables the system to recognize and adapt to diverse accents, colloquialisms, and culturally specific phrasing, ensuring an intuitive and inclusive experience for users across different geographic regions. For example, a preferred Al, such as ChatGPT, can accommodate both formal and conversational tones in multiple languages, tailoring its responses to the user’s preferences or regional norms. This feature is particularly valuable in international hospitality settings, where seamless communication is critical. It will be appreciated that this Al feature has applicability across all aspects of the present invention.
[0179] In some embodiments, the system may be deployed using an edge-based architecture to provide faster responses and to reduce the computing cost normally associated with large language models. In such configurations, part of the language model may run directly on a device located at or near the restaurant, such as a tablet, kiosk, router, or local edge server, while the remainder of the model or supporting services run in the cloud. The system may divide the processing between the edge and the cloud in a way that minimises network usage and latency, for example by generating early or partial results on the edge device and sending only the remaining work to the cloud. The system may also store commonly used data locally so that repeated requests can be answered more quickly without requiring full cloud processing. This hybrid approach allows the system to scale to many devices and locations with lower operating costs, while still providing the seamless synchronised media experience.
[0180] Video Sub-system: A video-based menu allows diners to view detailed videos of dishes, showcasing each dish with options for browsing by categories (e.g., vegetarian, desserts) and specific inquiries. This interactive video menu empowers customers to make better-informed dining choices. The video sub-system includes a catalogue of videos, each video associated with a particular menu item which in turn are associated with relevant data which may include, but not necessarily limited to, the ingredients in the menu item, allergens, kCal and the popularity of the menu item.
[0181] Alternatively, the video sub-system may also include a generative Al model trained to generate media content tailored to user inputs.
[0182] A dynamic video selection algorithm is used to enhance user engagement and optimize decisionmaking. The algorithm evaluates multiple contextual factors, such as the user's spoken or textual queries, menu item popularity metrics, dietary restrictions, real-time availability of items, and promotional offers. Particularly preferred is to take account of the user's known preferences, their reason for visiting, such as a celebration made known at the time of booking, and prior usage, such as what sort of dishes and drinks the user may prefer. By combining these data points, the system ranks and selects the most relevant videos to display that are both relevant to the user’s interaction and strategically beneficial for a restaurant. For example, the algorithm can prioritize videos of frequently ordered items or items currently on special promotions, which may ensure the user receives engaging and informative content. This allows the system to adapt to customer preferences and operational changes seamlessly.
[0183] AI-Based Voice processing: A processing module including: a) Natural Language Interaction: The Al interprets natural language input, allowing customers to ask questions about dishes, request recommendations, place orders, and process payments. Where the system is accessed remotely from the establishment operating the system, then the Al may also provide any further information and / or services the operator may deem appropriate, including locations and table availability, for example. To manage noise challenges in restaurant settings, it is currently preferred that users activate voice input via a microphone button. Customers may also prefer to use text input, which may include the use of prompt bubbles, for queries or instructions. b) Communication Outputs:
[0184] • Audio Response: The system responds verbally using the device’s speaker, simulating traditional waiter interactions and assisting visually impaired customers. This option may be muted, if desired. • Text Response: Complementary text displays provide a point of reference, and can help customers with hearing difficulties and reduce issues related to ambient noise.
[0185] • Video Selection: In response to the customer’s query, the system selects relevant videos from the video catalogue to be displayed to the user, and optionally
[0186] • Visual Prompts and Confirmations: The system displays prompts (e.g., “Would you like to see our cocktails?”) and confirms command understanding on-screen.
[0187] The Al-based voice processing module can retrieve data relevant to each menu item. This data can influence how the user receives one or more videos of menu items. For example, data that the system can retrieve might include the ingredients of each menu item. This allows the system to provide one or more videos of menu items that satisfy the user’s ingredient constraints. In this case, all menu items that do not meet the ingredient restraints will not receive a displayed video. Alternatively, the data may be used to sort the videos such that they are provided to the user in a particular order. An example of data that may be used to contribute towards this order is a metric capturing the frequency with which each menu item is ordered. This could contribute towards identifying a “best-seller” that is displayed first. Equally, the menu items may be subject to dynamic pricing scheme such as a flash-sale wherein less frequently ordered menu items are displayed first either with or without an associated discounted price.
[0188] The system may also incorporate advanced noise-reduction techniques tailored specifically for hospitality environments, including for example adaptive background noise filtering or hardware enhancements like beamforming microphones to isolate user voices in noisy surroundings. This subsystem is also capable of dynamic adjustment based on ambient noise levels, ensuring accurate processing of user commands regardless of environmental conditions. These features enhance user experience and operational reliability.
[0189] The voice input path employs beamforming with voice-activity detection (VAD) and adaptive noise suppression tuned for restaurant noise spectra; the resulting conditioned waveform reduces audio speech recognition (ASR) error and lowers buffering jitter for the TTS / video pipeline, thereby increasing ASR accuracy and reducing CPU / GPU cycles spent on false positives in noisy environments. Management System: This backend controls all settings, including restaurant-specific configurations, menu details, language support, and Al parameters. The management system integrates with an Al system. The management system provides all relevant instruction for the Al system including the Al’s role and persona, its voice accent, tone of voice, the limitations of what questions it should attempt to answer, instructions on recommendations, and what to do in various circumstances. The management system is synchronised with any POS and payment systems and synchronizes product details, processes orders, tracks item availability. While it is not necessary, a preferred system can manage orders and / or payments, if desired. It is not necessary to have synchronisation with a POS system. Instead, it is sufficient for the device being used by the user to send ordered items directly to a device associated with the establishment via email, Bluetooth, or another communication method, for example. The management system links to a video server that provides up-to-date video content for display on the device being used by the customer. Further, the management system can initiate voice or video calls with staff.
[0190] The management subsystem provides seamless integration with Point-of-Sale (POS) systems and payment gateways, where present, to enhance operational efficiency. Orders placed through the system are transmitted to the POS in real-time, allowing immediate processing and tracking. The system also supports dynamic pricing strategies, such as flash sales or promotional discounts, by updating menu items in response to real-time inventory data or sales trends. In cases where direct POS integration is unavailable, the system utilizes alternative communication protocols, such as Bluetooth or email, to ensure compatibility across a wide range of operational setups. These features collectively ensure that the system is adaptable and capable of meeting diverse industry requirements.
[0191] POS and Payment System Integration: Orders placed through the video waiter are sent to the POS, with an option for customers to pay directly via their device. In cases where direct POS integration is unavailable, payment can be conducted independently by the establishment.
[0192] Operational Flow and Functionality
[0193] 1. Menu Navigation and Ordering: Users can manually browse the video-based menu or verbally request recommendations. Orders are placed through speech or text, with the system confirming each order and notifying users of any item unavailability, providing alternative suggestions where appropriate.
[0194] 2. Information and Assistance Requests:
[0195] Customers can enquire about dishes (e.g., ingredients, dietary options) and receive relevant video or text responses. Beyond menu enquiries, the system may also offer local area information, for example, and may call for a taxi, or connect users with staff via voice or video calls. Enquiries can be made through speech or text.
[0196] 3. Multilingual Support:
[0197] The video waiter preferably adjusts to the language settings of the user’s device or the language spoken to the video waiter or allows manual language selection, enhancing accessibility for international guests.
[0198] 4. Payment and Table Management:
[0199] Customers can request the bill through the system, which redirects them to an integrated payment portal for completing transactions. The management system monitors table numbers and service charges, facilitating efficient and accurate billing.
[0200] 5. Data Collection and Usage:
[0201] The system is advantageously adapted to gather usage data, providing restaurants with insights on customer preferences, high-demand dishes, and operational efficiency. This data enables ongoing improvements in menu offerings, pricing, and service strategies.
[0202] Unique Features and Benefits
[0203] The following illustrates features and benefits of preferred aspects and embodiments of the present invention. Each feature and / or benefit may be present in combination with any one or more other feature and / or benefit unless otherwise apparent or specifically excluded.
[0204] For the restaurant
[0205] • It enhances the dining experience for their customers, with the videos empowering consumers to make better decisions. They are more likely to return due to a better dining experience.
[0206] • Staff training is simplified, and less time is spent explaining dishes and cocktails • Solve language issues — helps international guests
[0207] • Increases sales per cover - the videos naturally upsell optional dishes, such as starters and desserts, and help upsell cocktails.
[0208] • Self-ordering by customers lowers staff costs as fewer waiting staff required
[0209] • It helps solve staff availability issues - it can help keep the whole restaurant open even if short of waiting staff
[0210] • Quicker table turnaround — customers order quicker and pay the bill quicker
[0211] • It can inform diners of unavailable items and promote items near their expiration date or with a higher profit margin.
[0212] • The usage data can provide valuable insights, enabling a better understanding of a restaurant’s customers and helping make menu and dish improvements.
[0213] For the diners
[0214] • Empowered decisions - by seeing videos of the dishes on offer, the diners are empowered to make better-informed choices.
[0215] • Diners can get quality recommendations and information on dishes available.
[0216] • Convenience - they do not have to wait for a waiter to become available
[0217] • Better service - the video waiter is always available, always knowledgeable
[0218] • Pay at table - saves time and frustration
[0219] • Multi-lingual - service and description of menu items in any language
[0220] 2.2 Self-serving kiosk example
[0221] Use of the system does not have to be restricted to the restaurant table. As shown in Figure 19, the system can be accessed by customers in the form of a kiosk 200. Kiosks could be positioned in restaurants or hotel receptions.
[0222] Kiosks of the art, currently used in fast food and quick service restaurants, require the user to navigate manually through selections by touching screens to place an order. The video waiter transforms the food and drink kiosk ordering experience, by using videos of dishes, natural voice and text in combination, and can eliminate the user having to touch a screen. The video waiter as a self-ordering kiosk provides a similar experience to placing an order with a human, with the advantage of being shown relevant videos as part of the communication and showing a clear text record of the communication. The video waiter is knowledgeable about the menu and can converse in multiple languages in natural voice and text.
[0223] In a preferred embodiment, the kiosk provides an animated avatar that activates when a camera associated with the kiosk detects a customer, and wherein the avatar interacts with the customer in a manner deemed appropriate by the operator of the kiosk, and typically to mimic counter staff. The avatar may ask what the customer would like to order and offer prompts and recommendations, as well as repeating back the order, once complete.
[0224] The video waiter as a kiosk is more effective at upselling than a human as enticing videos are displayed as well as verbal and text recommendations and suggestions.
[0225] The video waiter as a kiosk provides a more natural alternative to touch screen solutions, and helps users make more informed choices.
[0226] 2.3 Hotel concierge example
[0227] Additionally, use of the system does not need to be restricted to food outlets. The system can be applied to the wider hospitality sector, including use in hotels, for example. Each customer (or each hotel room) can have access to the system. Figure 20 shows a use case in which the system provides answers to general enquiries and bookable services. The system still provides one or more videos to the user, however in this use case they may be of bookable services 201, such as spa appointments, or facility informational videos 202, such as buffet breakfast timings, as shown in Figure 21. Experiences outside of the hotel may also be offered.
[0228] The unique aspect of the service is that relevant videos are displayed as part of the communication with the guest, along with natural voice and text communication. For example:
[0229] • An enquiry about the hotel buffet breakfast would include a video of the breakfast serving room and the food on offer as part of the response. • An enquiry about local activities and attractions would include videos of the relevant activity or attraction as part of the response.
[0230] • An enquiry about late check out would include a video of a guest relaxing in their room as part of the response.
[0231] • An enquiry about a spa massage would include an enticing video of the massages on offer as part of the response.
[0232] • An enquiry about room upgrades would include a video of the premium rooms and suites as part of the response.
[0233] By including videos as part of the communication, guests are better informed and more likely to buy additional services, such as spa treatments, room upgrades, excursions etc.
[0234] Hence in a hotel concierge scenario, the system enhances the guest experiences by integrating with the hotel’s point of sale as well as external booking systems. The system can dynamically present video previews of services on the guest’s in-room tablet, lobby kiosk, mobile concierge app or directly on hotel websites. Each video can also be synchronised with generated text detailing the benefits, duration or pricing of each service. Guests are then able to secure their bookings via a simple voice command or user input. This simplifies the booking process and provides a personalised guest experience.
[0235] These examples highlight how the same synchronised conversational Al engine can be adapted to distinct operational environments.
[0236] We now describe further configurations of the same underlying architecture, focusing on variants optimised for different operational needs. The following systems or modules may be implemented independently or in combination with any other embodiment described in this specification. In general, it will be understood that any feature or aspect described herein may be used in any other embodiment of the invention, where appropriate.
[0237] 3. Itinerary-Linked Concierge Video system
[0238] The conversational Al architecture described previously can also be configured as a Video Concierge Service (VCS), suitable for applications including travel planning, activity scheduling, service selection, or any context requiring ordered recommendations. In this configuration, the system interprets conversational dialogue to derive structured activities, locations, services, events, or content items (“sequence items”).
[0239] The system may present a dynamic wishlist, favourites list, or interest queue, and generate a personalised, temporally-structured sequence (such as an itinerary). Each sequence item is associated with one or more media elements selected or generated by the synchronisation pipeline. The sequence may be updated continuously as the conversational session evolves.
[0240] As shown in Figure 22, the interface may display a video preview of a selected tour (e.g., a tour bus experience) along with structured metadata such as item title, description, and booking options. Swipe based navigation may also be supported such as across horizontally arranged video cards, allowing users to browse multiple recommended items while maintaining alignment between the presented media elements and the LLM’s textual or spoken narrative.
[0241] As shown in Figure 23, users may add items to a personalised “wish list” either manually or through natural -language interaction. This list may populate automatically based on conversational cues detected by the LLM, historical preferences, implicit interest signals (e.g., prolonged viewing of a video), or other contextual factors. In some implementations, the wish list may be synchronised across devices or stored temporarily without identifying the user, depending on the privacy configuration.
[0242] As shown in Figure 24, the LLM may adopt a concierge-style persona, optionally represented by an avatar, such as a well-dressed hotel concierge.
[0243] During conversation, the LLM may also retrieve location-specific metadata, real-time availability, operating hours, estimated travel times, promotions, or dynamic pricing information, enabling the system to generate a contextual itinerary. As shown in Figure 25, the system may respond to user queries by playing the associated video of the attraction while simultaneously rendering the corresponding conversational response in synchronised speech and text. The itinerary generation process may rely on structured prompts (e.g., “Plan my day tomorrow”), inferred contextual cues (e.g., items recently viewed), or any other hybrid schemes combining user preferences, contextual inference, as well as external data sources.
[0244] The system can automatically construct a personalised itinerary including, but not limited to, recommended visit times, travel guidance between items, predicted durations, availability information, booking links, and alternative options. The itinerary may be presented as a scrollable or card-based interface, a downloadable file, an email summary, or a shareable link.
[0245] The system may also generate a customised video itinerary, comprising edited or sequenced clips of each selected item, optionally with narration generated by the LLM’s text-to-speech system. These itinerary videos may be exported for social media sharing or distributed as private links for friends and family.
[0246] The VCS can be integrated into hotel websites, in-room tablets, mobile apps, airline applications, airport lounges or kiosks, allowing businesses to present venue-specific videos and recommendations to their guests. For example, a hotel may surface videos of its spa, gym, breakfast buffet or meeting facilities directly within the conversational experience. Tourist operators or airlines may also integrate their own promotional videos or booking systems.
[0247] Videos may be rendered from a pre-recorded catalogue, generated on demand using generative media models, or composited across multiple sources. The itinerary logic may rely on one or more of the following: the LLM, a symbolic planner, a rule-based engine, a graph-based recommendation algorithm.
[0248] The itinerary can provide a range of information, including full details of each item, suggested transport between each item, suggested times for experiencing each item, booking links or details, advertisements, offers and alternative items to consider. As an option, the VCS can also generate a personalized video of the itinerary, including edited videos of each item within the itinerary, in formats suitable for the user to share with friends, family or social networks. Customised versions of VCS can be configured which also include videos and information about businesses that have tourist customers, such as hotels, tour operators or airlines, so that their users can benefit from the VCS while accessing information and advice on services provided by the business.
[0249] Alternatively, where the system is used in connection with itinerary planning, travel recommendations, or venue discovery, the presentation layer may display a media element together with an interactive prompt bubble or equivalent UI control. The prompt bubble may provide a link or call-to-action allowing the user to access an external site or operator service, such as for booking a visit, making a reservation, purchasing a ticket, or otherwise engaging with the location or feature depicted in the media element. The interactive control may be activated by touch input, voice command, or other user interaction, and may be tailored according to operator preferences or contextual data.
[0250] 5. Dynamic kiosk interaction interface
[0251] A self-service kiosk device implements the synchronised conversational Al architecture, including the LLM-based conversational engine, the video synchronisation service, timestamped or approximate marker-based media alignment, and the rendering pipeline. The kiosk operates as a video-driven interaction terminal capable of executing ordering, booking and general service transactions using natural speech, synchronised video, optional text, and avatar-based presentation.
[0252] The kiosk may operate as a fully autonomous standalone system or may be integrated with any of the other embodiments described. Users simply speak in their natural language, and the kiosk responds with Al-generated text and speech output, while presenting relevant synchronised videos and graphical overlays as part of the conversation. The system is preferably fully contactless, providing improved hygiene compared to touch-based kiosks and reducing staff burden. Videos, images and graphic elements presented in synchrony with the conversation help to inform, guide, upsell or educate (for example, a children’s avatar may guide a child through a children’s menu with entertaining animated content).
[0253] The kiosk supports accessibility and inclusivity. For visually impaired or blind users, speech-based interaction and avatar-led audio narration offer an accessible alternative to visual menus. For deaf or hard-of-hearing users, optional touch-screen navigation and text captions are available. The kiosk may further adjust its avatar persona, tone, and behaviour dynamically based on user type, context or service environment.
[0254] Example in food restaurant:
[0255] • The kiosk has an image of the server, which can be an avatar or realistic human image, that is providing service via the kiosk
[0256] • A camera or other sensor recognises when a human is in front of the kiosk, and the avatar automatically can greet the customer “Hi, welcome to QuickFood - my name is XXX, how can I help you today”
[0257] • The Video Avatar Kiosk (VAK) recognises the language spoken by the user, and automatically adapts the conversation and written words displayed to the user.
[0258] • As the conversation with the VAK progresses, graphics and videos are shown in sync with the conversation, and an order basket is visible which has items added as per the conversation.
[0259] • The VAK is integrated with the restaurant POS for orders, and can upsell various food and drink offerings, as well as promoting sign up to loyalty programs, using synchronised videos at part of the communication. It may also be integrated with the restaurant loyalty system, with the conversation, videos and graphics adapting appropriately.
[0260] • The system is integrated with standard payment terminals, enabling payment without the need to touch the screen.
[0261] Example in hotels:
[0262] • The Video Avatar Kiosk has applications in many other sectors, such as hotel reception, providing multilingual service 24 / 7.
[0263] • A camera or other sensor recognises when a human is in front of the kiosk, and the avatar automatically can greet the customer “Hi, my name is XXX, how can I help you today”
[0264] • The Video Avatar Kiosk (VAK) recognises the language spoken by the user, and automatically adapts the conversation and written words displayed to the user.
[0265] • As the conversation with the VAK progresses, graphics and videos are shown in sync with the conversation, helping inform the guest and sell services. • The VAK may be integrated with hotel systems such as the hotel PMS, Spa booking system or loyalty system, enabling completion of bookings and service sales, with the conversation, videos and graphics adapting appropriately.
[0266] • The VAK can assist with any a range of hotel services, including reservations, check-in, checkout, and concierge services.
[0267] The kiosk comprises a computing unit, a display module, an output module (e.g., speakers), optional microphones for audio capture, and various presence-detection sensors such as RGB / depth cameras, ToF sensors, PIR sensors or mmWave radar. Presence detection enables a low-power idle state and transitions to active mode when a detection threshold is met. These sensors may be substituted with any alternative modality, such as LIDAR, ultrasonic detectors or infrared arrays.
[0268] The kiosk presents short-form video elements alongside conversational dialogue. The avatar may be 2D or 3D, static or animated, cartoon-style or realistic, or even a video-rendered persona. Avatar personality, tone and prompt conditioning are configurable according to brand, business type or service context. Media elements may be pre-recorded, curated clips, dynamically composed media, or generated by Al models. The welcome speech and filler clips are typically pre-recorded, for example, as this can reduce Al token usage, while providing consistency of the experience.
[0269] Upon activation, the kiosk is configured to establish a session with the conversational Al engine, configured with domain-specific datasets such as menu descriptions, PMS data, spa schedules, loyalty parameters, allergens, pricing rules or tax data. Dynamic context conditioning reduces hallucinations and enables persona adaptation in real time. The LLM generates the naturallanguage response, which is forwarded to both the text to speech and the video synchronisation service. The synchronisation service identifies referenced entities or events (e.g., “cheeseburger”, “massage”, “room upgrade”) and returns timestamped or approximate markers together with video asset identifiers. The video rendering module initiates provisional playback and then overrides it with precise or approximate synchronisation data when received, ensuring coherent multimodal presentation. The workflow may extend beyond food and hospitality to any domain requiring self-service interaction, including transport booking, ticketing, retail checkout, pharmacy or medical consultations, gym / spa management, airline check-ins or business reception service.
[0270] 6. AI-Driven Social Media Content module
[0271] A module for automatically generating shareable social media optimised video content based on a user’s interaction with the conversational Al system, the video waiter, the video avatar system, the concierge system, or any other embodiment described in this specification is also provided. This module, forms a standalone subsystem that may operate independently or in combination with the conversational Al architecture previously described. The module has various use cases across sectors such as tourism, retail, spas / gyms, entertainment venues, museums, subscription experiences, and general commercial services.
[0272] The module includes an LLM-based engine that performs natural -language parsing of text or speech derived reviews, conversational transcripts, and other user-generated content. The user may instruct the Al which aspects, if any, of their interaction with the system should be used in generating social media output. The engine may perform sentiment classification using an LLM- based score, a classical ML classifier, a hybrid model, or any equivalent technique. Positivesentiment phrases or meaningful experiential references (e.g., dish names, attractions, spa treatments, product descriptions, location references) are extracted and structured into a content schema suitable for downstream media generation. The system may use embeddings, semantic tagging, probabilistic filters, or rule-based extraction. Accurate automated sentiment classification and information extraction enables the system to generate consistent, high-quality promotional videos without manual review.
[0273] The components execute on one or more processors coupled to memory storing:
[0274] • a video database containing short-form clips of dishes, hotel amenities, attractions, spa services, products, or other goods and services;
[0275] • user interaction history derived from any conversational Al embodiment;
[0276] • full conversational transcripts; item selections, order baskets and wish-listed items; and optional third-party system data including POS transactions, PMS bookings or loyaltysystem events.
[0277] The module receives user reviews in text, voice, image or video form, or implicit signals such as likes, ratings or favourites. The system may solicit sentiment input through conversational prompts (“Did you enjoy your visit today?”), post-transaction POS / PMS events, or more general heuristic triggers.
[0278] Based on the analysed content schema, the system retrieves relevant media assets, which may include:
[0279] • short clips of menu items, hotel services, spa treatments or attractions;
[0280] • still images, icons, labels or itinerary elements;
[0281] • user-provided images or clips;
[0282] • Al-generated or diffusion-model-generated imagery;
[0283] • cloned-voice narration or TTS-synthesised narration;
[0284] • soundtrack or background music selected using metadata or LLM-derived heuristics.
[0285] Asset retrieval may be based on embedding-similarity search, metadata lookup (ID, location, length, format), rule-based matching, template-based models, or any equivalent process. Automated retrieval improves relevance and permits scalable, personalised output.
[0286] The video composition engine performs multi-stage synthesis including:
[0287] • mapping textual elements and media assets to a defined temporal structure tailored to target platforms (e.g., short-form Instagram Reels, TikTok vertical clips, WhatsApp-friendly compressed formats);
[0288] • selecting appropriate segments based on feature prominence, narrative flow, or user emphasis;
[0289] • adding text overlays, headings, captions, itineraries or descriptive labels;
[0290] • generating optional narration using cloned voices or TTS outputs; and
[0291] • generating a final media file with platform-optimised aspect ratio, compression, frame rate, bit- rate and rendering profile. This automated timeline-composition process ensures consistent, reproducible output without manual video editing and supports deployment across multiple platforms with minimal adjustments.
[0292] In some embodiments, the system may format the generated output in visually stylised forms commonly associated such as Al assisted video or micro-story templates. These include but are not limited to: dynamic text overlays, animated captions, scene-based transitions, or other visually engaging elements commonly associated with social media creation tools.
[0293] The output module prepares a downloadable video asset, an optional pre-populated caption or text snippet, and a preview screen that includes key frames or a playable version of the full clip. Platform-specific optimisations are applied (e.g., vertical portrait mode for TikTok / Instagram, square 1 : 1 for legacy formats, compressed variants for low-bandwidth messaging services). The module may support automatic posting, scheduled posting (subject to platform policies), or export to email, messaging apps, review sites or social networks.
[0294] It should be noted that it is not necessary, in any embodiment of the present invention, that all visual representations should be videos. Static images and other representations may be incorporated as and when desired or appropriate, as will be recognised by those skilled in the art. However, it is an advantage of the present invention that videos form a part of the invention, as this has been found to enhance customer engagement and can now be presented substantially instantaneously from an appropriate internet server on request by a system of the invention, either by manual interaction with a screen, such as by swiping, or by an Al with which the customer interacts. It is, thus, preferred that the maj ority of images presented to the customer are video based media assets, including, but not limited to, MP4 (H.264 / H.265) or WebM (VP9 / AV1) format with AAC or Opus audio, or some other suitable format. Dynamically generated media, streaming video, animated sequences, or any alternative encoding scheme may also be used. The system is optimised for video-first responses, and the client employs hardware video decoding where available to minimise CPU load. APPENDIX A: KEY FEATURES
[0295] We list high level features, each with a number of optional features.
[0296] Note that any of the high-level features can be combined with one or more of the other features, and / or any of the optional features. Any of the optional features can be combined with one or more of the other optional features. Features are implementable across hospitality, retail, transportation, healthcare, and other customer facing environments.
[0297] Key feature A: Synchronised Al conversation, and video in real time
[0298] A computer implemented pipeline is provided that includes an LLM that generates conversational output that forms the basis of the interaction. A video synchronisation service selects media elements and may inject timestamp or marker, and a text to speech system renders speech. A presentation layer or output module then displays the synchronised text, voice, and / or video. Rendering may be performed on a client, server, an edge note, or other hybrid arrangement. Temporal alignment refers to the coordinated presentation of any portion of one or more media elements with any portion of the conversational output. Temporal alignment may apply whether the conversational output is visual (text) or audible (speech), or a combination thereof. Temporal alignment may be achieved using synchronisation data. Hence the media elements appear or are displayed at an appropriate time relative to the text or speech output, whether based on semantic logic, timing metadata, predicted durations, or other alignment indicators. Synchronisation data, as used herein, refers to any information or metadata usable for controlling, or adjusting the temporal alignment between any portion of the conversational output (text or speech) and any portion of one or more media elements.
[0299] In some embodiments, temporal alignment may be achieved without the use of explicit synchronisation data. In such implementations, the system may initiate, adjust, or select media playback based on implicit or approximate timing cues, including the start of the conversational output, keyword or entity detection within the user input, semantic or grammatical analysis of the conversational output, classification or topic recognition, LLM probability or scoring metrics, heuristic timing rules, or predicted timing derived from incremental segments of the conversational output. These implicit or approximate mechanisms allow the system to display relevant media elements in a timely manner even when explicit synchronisation metadata is unavailable, delayed, or not generated.
[0300] We can generalise as:
[0301] A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:
[0302] (a) receive user inputs including voice, text, and / or selection-based inputs;
[0303] (b) generate conversational output using a large language model (LLM) or Al model based on the user inputs;
[0304] (c) dynamically select or generate one or more media elements relevant to the conversational output;
[0305] (d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output.
[0306] Optional features:
[0307] Synchronisation data
[0308] • The temporal alignment is based on synchronisation data.
[0309] • The synchronisation data comprises any information or metadata used to determine, control, or adjust the temporal alignment between any portion of the conversational output and any portion of one or more media elements.
[0310] • Synchronisation data comprises timestamps or markers, playback instructions, ordering rules, semantic rules, or any metadata indicating when a media element should start, stop, or transition.
[0311] • Synchronisation data further specifies timing, sequencing, overlap, or transition rules for a plurality of video elements synchronised with different segments of the conversational output.
[0312] • The synchronisation data is generated in advance, during, or after generation of the conversational output or text to speech output.
[0313] • The synchronisation data is embedded within, associated with, or generated separately from the conversational output. • The synchronisation data is represented as tokens, structured metadata, vectors, or semantic annotation.
[0314] Without explicit synchronisation data.
[0315] • Temporal alignment is achieved without explicit synchronisation data.
[0316] • Temporal alignment is based on implicit or approximate triggers including conversational start time, keyword detection, entity recognition, semantic grouping, or heuristic timing rules, predicted timing based on the conversational output.
[0317] • Temporal alignment is based on detection of keywords or entities within the user input prior to the generation of the conversational output.
[0318] • Temporal alignment is based on classification of the conversational output.
[0319] • Temporal alignment is based on probability, or scoring metrics produced by the LLM.
[0320] Contextual data
[0321] • LLM or Al model is configured with domain-specific contextual data.
[0322] • contextual data comprises information relevant to the system’s specific domain of use.
[0323] • LLM or Al model is configured to retrieve contextual data generated or obtained during operation of the system, such as: user data, popularity metrics of menu items; user demographic or localisation data; environmental conditions including ambient noise level; and temporal conditions such as time-of-day or promotional period.
[0324] General features
[0325] • The system is configured to present multiple media elements simultaneously such as using overlay, split screen, or picture in picture format.
[0326] • Synchronisation data specify relationships between the multiple media elements.
[0327] • The LLM or Al model is configured to modify its conversational strategy based on detected system conditions, such as latency, available bandwidth, device capabilities, or sensor data.
[0328] • The system collects analytics data relating to user interactions with synchronised media, and to adapt future media selections using reinforcement-leaming-based ranking. • The system is configured to select or replace media elements in real time based on inferred or detected user attention or feedback.
[0329] • The system is configured to pre-cache or pre-fetch predicted media elements based on contextual data.
[0330] • The system uses presence detection to activate or deactivate the output module using a camera, depth sensor, radar, PIR sensor, mmWave, or other sensor to activate or deactivate the output module.
[0331] • Initiation of a video element is performed substantially contemporaneously with presentation of a first word or other portion of the conversational output associated with that video, with an allowable phase margin including a predefined time duration, such as ± 0.2s.
[0332] • Output modules include screens, speakers, projectors, kiosks, tablets, smartphones, wearables, AR interfaces, or multi-screen display.
[0333] • The system detects timing drift between predicted duration of the conversational output and an actual duration of the conversational output, and to dynamically adjust playback of the one or more media elements to maintain temporal synchronisation.
[0334] • The large language model is executed using a partitioned inference pipeline in which a first portion of the model is deployed on an edge-computing node near the output module and a second portion of the model is executed on a remote server.
[0335] • The system being configured to merge intermediate inference results generated by the respective portions of the model to reduce network latency and computational load.
[0336] • The system comprises a speech recognition or audio analysis component deployed on a remote server, on the access device, on an edge-based system, or across a distributed architecture.
[0337] Media library
[0338] • the system having access to an operator defined media library comprising media assets.
[0339] • priority is given to an operator-defined media library comprising media assets when selecting or generating said media elements.
[0340] • media assets are stored with associated metadata.
[0341] • Metadata associated with each of the one or more media element comprises timing data, semantic vectors, operator approval flags, or media-type identifiers. • Media library further comprises media-generation inputs including one or more of: templates, style parameters, seed vectors, or textual description.
[0342] • Retrieval of media elements from an external source is permitted, such as when no relevant media element is available in the operator-defined media library.
[0343] • The system generates new media elements using a generative Al model, and stores the generated media elements with associated metadata in the operator-defined media library.
[0344] • External or Al-generated media elements undergoes a validation pipeline comprising quality checks, compliance filtering, and / or metadata tagging.
[0345] • The media library is implemented using a cloud-hosted repository, on device cache, an edge node, or a hybrid architecture thereof.
[0346] Interactive prompts
[0347] • the system is configured to generate and present one or more Al-generated interactive prompts being dynamically generated based on at least one of: the user input, the conversational output, contextual data, or metadata associated with a media element.
[0348] • The interactive prompts suggest follow-up queries, recommended actions, or conversational options, the interactive prompts.
[0349] • The interactive prompts comprise one or more interface elements including: prompt bubbles, tiles, cards, or other selectable UI components.
[0350] • The interactive prompts are temporally aligned with the conversational output or with the presentation of the one or more media elements
[0351] • The interactive prompts include selectable actions relating to information browsing, recommendations, bookings, reservations, or orders.
[0352] Key Feature B: Latency resolution
[0353] A latency resolution mechanism is implemented that starts provisional video playback upon receiving the LLM output and subsequently adjusts or overrides that provisional playback when synchronisation data received from the video synchronisation service becomes available, while allowing the presentation of the text and / or voice to continue without interruptions.
[0354] We can generalise as: A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:
[0355] (a) receive user inputs including voice, text, and / or selection-based inputs;
[0356] (b) generate conversational output using a large language model (LLM) or Al model based on the user inputs;
[0357] (c) dynamically select or generate one or more media elements relevant to the conversational output;
[0358] (d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output; and wherein the system is configured to initiate provisional playback of a media element prior to receiving synchronisation data and to override or adjust the provisional playback when synchronisation data becomes available.
[0359] Optional features:
[0360] • the one or more processing units are configured to update or modify the playback timing of the media elements dynamically during generation of the text or audio output to compensate for latency or content changes.
[0361] • The system is configured to adjust the playback position based on a detected timing drift, lag, or other computed discrepancies between predicted and / or text or speech duration.
[0362] • The provisional playback is based on a predicted or preliminary media selection, a first ranked media element, or a previously cached element.
[0363] • The system is configured to update, correct, interrupt, or replace provisional playback upon receiving synchronisation data without interrupting or restarting the conversational session.
[0364] • The system includes a fallback mode when synchronisation data is delayed or unavailable.
[0365] • During fallback mode, the system initiates a provisional playback using a non-item specific media element, such as locally cached previews, stock video, introductory clip, static imagery, graphics, textual description, or any other default media asset.
[0366] • The system includes an auxiliary Al model configured to analyse early conversational segments to identify or predict a suitable provisional media element when synchronisation data is delayed or unavailable. Key feature C: Incremental conversational segments
[0367] The conversational output generated by the LLM or Al model may be provided in incremental or “chunked” segments, rather than as a single complete response. These conversational segments may include, but are not limited to: individual tokens, group of tokens, partial text, partial speech frames, partial sentences, or any other unit of conversational output. Segments may be produced in irregular intervals and may differ in length, structure, or modality. Temporal alignment may include sequencing or overlapping of a plurality of media elements associated with the different segments of the conversational output. Because the conversational output becomes available progressively, the system may update, refine, or regenerate synchronisation data in real time as new segments are received.
[0368] We can generalise as:
[0369] A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:
[0370] (a) receive user inputs including voice, text, and / or selection-based inputs;
[0371] (b) generate conversational output using a large language model (LLM) or Al model based on the user inputs;
[0372] (c) dynamically select or generate one or more media elements relevant to the conversational output;
[0373] (d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output; and wherein the conversational output is generated or provided in segments.
[0374] Conversational segments
[0375] • the system is configured to override or replace a provisional media element upon receipt of later conversational segments that identify a more relevant media element
[0376] • system detects keywords, entities, or semantic cues within individual conversational output chunks and initiates media selection or playback in response to those cues.
[0377] • the system estimates speech duration or predicted audio timing for each conversational segment and uses the estimate to align or adjust the timing of associated media elements. • the system computes timing deviations between predicted timing (based on earlier segments) and updated timing (based on later segments) and adjusts media playback to correct for drift.
[0378] • conversational segments correspond to different media elements, and the system schedules or re-schedules the media elements according to the sequence in which the segments are generated.
[0379] • Conversational segments are generated from a streaming system.
[0380] • If no relevant media is identified in an early conversational segment, the system initiates a fallback or default media element until a later segment provides sufficient content to determine a specific media element.
[0381] • The system adjusts the synchronisation updates based on the rate, latency, or size of conversational output segments.
[0382] • The system is configured to update or refine the synchronisation data as additional portions of the conversational output are generated in incremental segments.
[0383] • The system is configured to adjust playback of one or more media elements based on the updated synchronisation data.
[0384] Key Feature D: Intent management with contextual prompting
[0385] The system can also include an intent management module that processes either the user inputs and / or the generated LLM output to identify actionable intent with an ongoing interaction. Based on the intent, the system can then create follow put questions, or multimedia cues in order to guide the user towards completing an action. For example: “would you like to book a table for 6pm?” or “would you like to see the dessert menu?”. This can be expressed either in natural language, via text, via synchronised video or graphics, or any other multimedia form supported by the output module. This intent management module provides a structured way for the system to combine the LLM’s output with domain specific service logic, enabling the system to provide meaningful transactions while providing a natural conversational flow.
[0386] We can generalise as:
[0387] A system for synchronised conversational user engagement, the system comprising one or more processing units configured to: (a) receive user inputs including voice, text, and / or selection-based inputs;
[0388] (b) generate conversational output using a large language model (LLM) or Al model based on the user inputs;
[0389] (c) dynamically select or generate one or more media elements relevant to the conversational output;
[0390] (d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output; and wherein the system includes an intent management module configured to: identify, from the user input or the conversational output, an actionable intent relating to a transaction or service event, and to generate one or more action prompts in natural language and / or multimedia form.
[0391] Optional feature:
[0392] • intent management module is implemented using the LLM, an auxiliary Al model, a symbolic planner, a rule-based engine, or any hybrid thereof.
[0393] • upon user confirmation, the intent management module trigger execution of the identified intent through integration with one or more external systems or APIs, including booking, reservation, transport, payment or delivery services; and dynamically update the conversation state and synchronised media presentation based on the status returned from the external systems.
[0394] Key feature E: Dynamic trip generation
[0395] Another aspect of the conversational Al system is a dynamic trip generation module that enables the system to interpret the conversational dialogue relating to events, services, or other point of interest, and to construct a personalised sequence of items (such as itinerary) throughout the conversational session.
[0396] We can generalise as:
[0397] A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:
[0398] (a) receive user inputs including voice, text, and / or selection-based inputs; (b) generate conversational output using a large language model (LLM) or Al model based on the user inputs;
[0399] (c) dynamically select or generate one or more media elements relevant to the conversational output;
[0400] (d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output; and wherein the system includes a dynamic trip generation module configured to generate a sequence of activities, locations, items, or service based on the conversational output, and to associate each element of the sequence of activities with one or more corresponding media elements.
[0401] Specific Itinerary-Linked Concierge Video system
[0402] A system for video selection, comprising:
[0403] (i) a video store containing media tagged with location or point of interest metadata; and
[0404] (ii) a LLM that is configured to (a) receive user inputs defining a journey or an itinerary, including location or point-of-interest requirements and to (b) process those inputs as prompts, and (c) to generate an itinerary and to automatically provide video previews associated with locations or point-of-interest; (d) to respond conversationally to further prompts from the user to refine the itinerary; and
[0405] (iii) a display interface that presents an interactive itinerary alongside the associated video previews.
[0406] Optional features:
[0407] • the LLM is configured to guide the user through each item of the itinerary, such as for booking, purchasing, reservation, or providing transport.
[0408] • The system provides conversational guidance between itinerary items, including time-based reminders and follow-up prompts for related services (e.g. “Your dinner is booked for 7 pm: would you like to arrange transport for 6 pm?”).
[0409] • The system dynamically updates or regenerates the itinerary based on user inputs, contextual inference, schedule changes, real-time availability, or data received from external booking, scheduling, or transport APIs (for example, "you reserved dinner at 7pm, would you like to book a taxi for 6pm?").
[0410] • The system automatically adjusts timing between itinerary items using estimated travel times, venue hours, or other operational constraints.
[0411] • The itinerary comprises a single destination or multiple destinations, or represent a variablelength chain of events or activities.
[0412] • The system generates and updates a corresponding set of media previews or explainer videos as itinerary content changes.
[0413] Key feature F: Conversational Al system optimised for a specific application
[0414] The LLM may be dedicated to a specific application domain, with a dynamic context updating mechanism that refines the LLM behaviour based on real time data.
[0415] We can generalise as:
[0416] A LLM based conversational Al system optimised for an application-specific context, the system being configured to:
[0417] (i) receive user inputs including voice, text, and / or selection-based inputs;
[0418] (ii) operate a LLM configured with an application-specific context that defines a persona and / or behavioural attributes associated with the relevant service domain;
[0419] And in which the system is configured to dynamically update the context based on real-time data, including, but not limited to: service data, product data, availability information, promotions, user preferences, or feedback.
[0420] Key feature G: Brand specific voice agent
[0421] The conversational output may also be rendered in a brand aligned synthesized voice.
[0422] We can generalise as:
[0423] A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:
[0424] (a) receive user inputs including voice, text, and / or selection-based inputs; (b) generate conversational output using a large language model (LLM) or Al model based on the user inputs;
[0425] (c) dynamically select or generate one or more media elements relevant to the conversational output;
[0426] (d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output; and wherein the system further comprises a brand specific voice agent module configured to:
[0427] (i) operate an Al-powered voice synthesis module trained on voice models or samples from multiple different sources, in which the voice models or samples from each source are defined by user-recognisable characteristics that are specifically associated with each source (such as a brand, brand persona, or service provider identity);
[0428] (ii) enable the LLM to generate conversational output rendered in a synthesized voice specific to one of the sources; and
[0429] (iii) deploy the synthesised voice across multiple different types of channels, including any two or more of the following: kiosks, telephone-based services, web-based services, mobile applications, in-restaurant devices, and in-hotel room devices.
[0430] A system for interactive user engagement, comprising:
[0431] (i) an Al-powered voice synthesis module trained on voice models or samples from multiple different sources, in which the voice models or samples from each source are defined by user- recognisable characteristics that are specifically associated with each source (e.g. are from a brandspecific source, such as a specific restaurant chain, or a specific brand personality);
[0432] (ii) a LLM configured to interpret user inputs and to generate conversational output rendered in a synthesized voice specific to one of the sources;
[0433] (iii) a module configured to deploy the synthesised voice across multiple different types of channels, including any one or more of the following: kiosks, telephone-based services, web-based services, mobile applications, in-restaurant devices, and in-hotel room devices.
[0434] Optional features:
[0435] • the LLM is configured to generate conversational output based on brand-specific marketing language. conversational output maintain tone and / or consistency across all chanels.
[0436] A brand-specific avatar synchronised to said Al-powered voice output, all or part of the time.
[0437] Key feature H: Customer facing device (tablet / mobile / kiosk)
[0438] The methods and systems described can be implemented using a customer-facing device, such as a restaurant tablet, tabletop terminal, kiosk, mobile phone, or other personal mobile device. Such a device may include a display, one or more processors, a microphone, speakers, and an input / output module.
[0439] We can generalise as:
[0440] A customer-facing device for interactive food or drink ordering, the device comprising a display, one or more processors, and input / output module, the device being configured to:
[0441] (i) receive user inputs including voice, text, touch, and / or selection based inputs;
[0442] (ii) activate or interact with a conversational LLM-based Al system that is configured with a context that defines features or characteristics of a food or hospitality service provider;
[0443] (iii) present synchronised conversational output and one or more media elements relevant to the conversational output, wherein the media elements are retrieved from an operator defined media library comprising pre-approved media assets stored with associated metadata;
[0444] (iv) display the one or more media elements and the conversational output such that the media elements are temporally aligned with the conversational output;
[0445] (v) interpret user input and / or conversational output to extract a food and / or drink order; and
[0446] (vi) transmit the extracted order to a point of sale system or kitchen display system for fulfilment.
[0447] Optional features:
[0448] • The device is a tablet, a user’s mobile phone, or a personal mobile device.
[0449] • The device is configured to access the conversational Al system by scanning a QR code, and wherein the QR code loads a browser based interface, mobile web application, or native application.
[0450] • The device is configured to present a bill or itemised cost to the customer and receive payment through the device via a card reader, QR code, contactless payment, or mobile payment interface. • Where a response, or conversational output, requires a media element not present in the library, the Al may seek to generate a suitable media element from a part of a media element in the library, or combine two or more elements from the library. If there is no suitable element in the library, the system may be permitted to source elements from external sources, which have preferably been pre-designated by the operator.
[0451] Key feature I: Website based application (tablet / mobile / kiosk)
[0452] A website based application for synchronised conversational user engagement, the website being delivered to a user device via a web browser and comprising executable instructions configured to:
[0453] (i) receive user inputs including voice, text, touch, and / or selection based inputs;
[0454] (ii) activate or interact with a conversational LLM-based Al system that is configured with a context that defines features or characteristics of a food or hospitality service provider;
[0455] (iii) generate conversational output based on the user inputs via the conversational Al system;
[0456] (iii) dynamically select or generate one or more media elements relevant to the conversational output, wherein the one or more media elements are retrieved from an operator-defined media library comprising pre-approved media assets and / or media-generation inputs stored with associated metadata / sEp / e) present the conversational output together with the one or more media elements within the web browser such that the media elements are temporally aligned with the conversational output; andisEpj(f) enable the user to browse items or services.
[0457] Optional features:
[0458] • the website is accessed by scanning a QR code displayed within a restaurant or hospitality venue.
[0459] • the website is accessible via a link provided in an online search result, mapping service interface, or third-party discovery platform.
[0460] • the website is configured to operate in a browsing-only mode providing information, recommendations, or guidance without accepting orders. the website is configured to accept food and / or drink orders and transmit such orders to a point- of-sale system, kitchen display system, or bar system.
[0461] Key feature J: Dynamic kiosk interaction interface
[0462] An Al-powered food or drink ordering kiosk device that is configured to:
[0463] (i) activate automatically upon detecting a customer’s presence, such as by using a camera or a presence detection sensor;
[0464] (ii) activate or interact with a conversational LLM-based Al system that is configured with a context that defines features or characteristics of a food or hospitality service provider;
[0465] (iii) automatically initiate or continue a conversation with the customer via voice and / or relevant media content, in response to a customer prompt;
[0466] (iv) automatically analyse the conversation with the customer to extract a food and / or drink order from the conversation;
[0467] (v) generate a bill or cost for that order and present that orally and / or visually to the customer;
[0468] (vi) automatically receive and process payment for that order; and
[0469] (vii) then send that food and / or drink order to a bar or kitchen.
[0470] Optional features:
[0471] • the kiosk is configured to provide a natural language response to a query from the customer
[0472] • A brand-specific avatar synchronised to said Al-powered voice output, all or part of the time.
[0473] Key feature K: AI-Driven Social Media Content module
[0474] A system for automatically generating social media content comprising:
[0475] (i) an LLM-based Al engine that is configured to (a) receive interaction data relating to a user experience, the interaction data including one or more of: user provided text, selected menu items, interaction history, and user reviews, (b) automatically analyse the interaction data; And wherein the system is configured to automatically generate social media content, based on said analysis, said social media content being in a format optimised for distribution on social media platforms, such as Instagram, TikTok, or WhatsApp.
[0476] Optional features:
[0477] • Analysing the interaction data includes performing sentiment analysis on user reviews, user provided text, or other subjective content to identify positive sentiment for inclusion in the generated social media content.
[0478] • The Al engine is configured to retrieve one or more media elements associated with the selected menu items or user interaction history and incorporate the one or more media elements into the social media content.
[0479] • The system provides a user interface enabling a user to select, approve, exclude, or modify specific aspects of the user experience to be incorporated into the social media content.
[0480] • The system is configured to automatically schedule distribution of the generated social media content to one or more social media platforms.
[0481] • Scheduling of distribution is determined based on user preferences, historical posting behaviour, or predicted engagement metrics.
[0482] • the system is configured to automatically request a user to leave a review after an interaction or experience.
[0483] • the system includes an output module that automatically publishes or provides a preview of the social media content for the user or brand to review and / or share.
[0484] Key feature L: Video driven interactive customer engagement
[0485] We can generalise as:
[0486] A system for interactive customer engagement, the system comprising one or more processing units configured to operate or communicate with an onboard or remotely hosted artificial intelligence or machine learning model, wherein the one or more processing units are configured to:
[0487] (i) receive and analyse user inputs, including voice and / or text, relating to products, services, or general enquiries and to dynamically select or generate one or more media elements relevant to the user inputs, based on contextual data such as user preferences, historical behavior, real-time availability and promotional data; and
[0488] (ii) retrieve and present via an output module the selected or generated media elements, including videos, images, or other interactive content.
[0489] The system can provide video or media elements for each menu or product or service item. These videos or media elements can for example include details regarding one or more of the following: preparation process or ingredient details, offering a level of insight waiters may not be able to provide consistently. By providing contextual and visual interaction experience, users can make decisions faster. Media elements can include text, videos, images, 3D models, augmented reality (AR) content, holograms, or other interactive multimedia content.
[0490] Key feature M: Video driven interactive customer engagement system with synchronised text descriptions
[0491] A system for interactive customer engagement, the system comprising one or more processing units, configured to operate or communicate with an onboard or remotely hosted artificial intelligence or machine learning model, wherein the one or more processing units are configured to:
[0492] (i) receive and process user inputs, including voice and / or text and / or gesture interactions, relating to products, services, or general inquiries and to dynamically select or generate one or more media elements relevant to the user inputs, based on contextual data such as user preferences, historical behavior, real-time availability and promotional data;
[0493] (ii) generate text descriptions dynamically based on the selected or generated media element; and
[0494] (iii) retrieve and present via an output module the selected or generated media element, including videos, images, or other interactive content, wherein the output module presents the generated text alongside the media element in real time.
[0495] Key feature N: Multilingual and dialect processing engine
[0496] We can generalise as: A system for interactive customer engagement, the system comprising one or more processing units configured to operate or communicate with an onboard or remotely hosted artificial intelligence or machine learning model, wherein the one or more processing units are configured to: (i) receive and process multilingual user inputs, including voice and / or text, relating to products, services, or general inquiries, to interpret the user inputs across multiple languages using natural language processing (NLP) models trained on diverse linguistic datasets, including hospitalityspecific terminology, and dynamically select or generate one or more media elements relevant to the interpreted user inputs based on contextual data, such as user preferences, historical behavior, real-time availability and promotional data, and (ii) retrieve and present the selected or generated media elements, including videos, images, or other interactive content, on a suitable display.
[0497] Optional features:
[0498] Processing units
[0499] • the one or more processing units are configured to process voice inputs in noisy environments using adaptive noise reduction techniques, beamforming, or other acoustic conditioning techniques.
[0500] • the one or more processing units are configured to rank the media elements based on one or more of the following: user preferences and past interactions, current product or service availability, popularity metrics based on user engagement, real time inventory levels, promotional factors or external conditions such as seasonal trends.
[0501] • the LLM or Al model is configured to interpret user queries using natural language processing (NLP) models and to recommend media elements based on Al-driven content selection models.
[0502] • the LLM or Al model incorporates reinforcement learning-based ranking models to adjust realtime video prioritization based on user interaction.
[0503] • the LLM or Al model tailored for multilingual capabilities and support regional dialects, variations in language usage.
[0504] • NLP models are dynamically updated with feedback loops, ensuring continuous improvement in understanding diverse user inputs. • language models are pre-trained using diverse linguistic datasets specifically tuned for specific service domains.
[0505] • system further facilitates natural language conversations by generating verbal or textual responses to user queries in addition to delivering media content.
[0506] • the system is configured to monitor real-time user interaction and to store interaction history for future personalisation.
[0507] • the system is configured to dynamically prioritise media elements based on user input, contextual data such as dietary preferences, popularity metrics, promotional data, and real-time availability.
[0508] • the system is configured to sort media elements based on real-time user behaviour patterns or historical order data.
[0509] • the system is configured to support gesture-based navigation, allowing users to browse the media elements without physical interaction with the device.
[0510] • the system is configured to provide analytics on user engagement with medial elements to optimise future recommendations.
[0511] • the system is configured to predict user intent from incomplete or ambiguous user inputs.
[0512] • system is tailored for one or more of the following customer-facing environments: hospitality, retail, transportation healthcare, or entertainment environment.
[0513] • contextual data further includes environmental data such geolocation or ambient conditions.
[0514] • contextual data further includes biometric indicators, and device specific parameters.
[0515] • the system further includes a generative Al model trained to generate media content tailored to user inputs.
[0516] Conversational output
[0517] • LLM or Al model output is dynamically adjusted based on user preferences, historical behavior, real-time availability or promotional data.
[0518] • system generates text descriptions using the LLM or Al model that highlight key information from the media element such as ingredient detail, dietary information, allergen warnings or nutritional facts.
[0519] • system enables interactive text expansion, allowing users to tap or hover to reveal further insights beyond the default text displayed, such as additional details or ingredient breakdowns. • LLM or Al model output describe ingredient details and dietary information.
[0520] • LLM or Al model output implements updates in real time, such as based on chances in menu availability.
[0521] • the system automatically optimises text formatting for readability.
[0522] • system is configured to highlight critical details such as allergens or promotions.
[0523] • system is configured to adjust font size and contrast based on the display settings.
[0524] • the user input relates to any one or more of the following: ordering food and / or drink; ordering a hospitality service; a question about a food and / or drink item; a non-food and / or drink related question; a question regarding hospitality services; a non-food and / or drink, such as a unique identifier for the user’s location (e.g. table number); requesting a member of staff.
[0525] • the user input is multilingual, supporting multiple languages and regional dialects.
[0526] • user inputs also include touch, gaze tracking, or gesture-based inputs.
[0527] Media
[0528] • Media elements can include text, videos, images, 3D models, augmented reality (AR) content, holograms, or other interactive multimedia content.
[0529] • media elements are stored or generated by the video synchronisation service comprising: (a) a video-catalogue interface configured to receive metadata from the management subsystem; and (b) a timing engine configured to generate the synchronisation data based on a semantic analysis of the artificial-intelligence model output.
[0530] • the media elements depict one or more of the following: the food and / or drink item; a hospitality service; a promotional offer; directions to locations of interest within a restaurant space such as toilets, self-service bar, car park etc; directions to locations of interest external to a restaurant space such as local public transport connections, petrol stations etc; non-food and / or drink related information such as a weather forecast or news.
[0531] • the media elements are selected based on data linked to the user’s profile including, but not limited to, birthdays or frequently ordered menu items.
[0532] • the media elements are dynamically updated based on real time menu changes or ‘specials’.
[0533] • the video sub-system can be independently viewed by the user without the requirement of a user input. Output module
[0534] • The output module comprises one or more of: a display, a touchscreen, a speaker, a kiosk display, a mobile device screen, a mobile device screen, or a multi-screen environment.
[0535] • The output module is implemented using XR smartglasses, augmented reality (AR) devices, virtual reality (VR) headsets, mobile devices and other wearable or fixed computing platforms.
[0536] • The output module provides a user interface accessible via a display device.
[0537] • the user interface is executed via an application or a web-based platform accessible through the suitable display.
[0538] • the user interface includes a language selection feature.
[0539] • the user interface includes an onboarding video.
[0540] • the user interface includes categories into which t menu items are grouped.
[0541] • the user interface includes a feature that displays the speech or text input.
[0542] • the user interface includes a feature that clearly accentuates the word on the screen that is being conveyed via audio output.
[0543] • the user interface includes a feature that, when selected, allows text input on the device.
[0544] • the user interface includes a feature that, when selected, activates the microphone.
[0545] • the user interface displays a video related to a menu item on the screen of the device.
[0546] • the video related to a menu item that is displayed on the screen is maximised at any one time.
[0547] • the video related to a menu item that is displayed on the screen takes up the entirety of the screen.
[0548] • the video related to a menu item that is displayed on the screen fills the space that remains on the screen once other user interface features have been included on the screen.
[0549] • Output module supports gesture recognition or augmented reality features.
[0550] Management Subsystem
[0551] • the system includes a management subsystem configured to synchronise operational data including item data, service data, availability information, and / or payment parameters, and support transaction workflows. • the management subsystem is configured to synchronise with one or more enterprise systems including: POS systems, payment gateways, or video servers, allowing dynamic updates to menu items, pricing data, or promotional data in real time.
[0552] • Management subsystem allows enables configuration of prioritisation rules based on inventory levels, margin considerations, demographic insights, user behaviour, promotional strategy, or other business factors.
[0553] • prioritisation rules can also be time-dependent, ensuring operators can offer time-based promotions such as happy hour media elements.
[0554] • the management subsystem includes configuration instructions for the LLM or Al model, including persona definitions, tone of voice, recommendation boundaries, conversational constraints, and escalation rules.
[0555] • the management subsystem can hold or retrieve the POS or enterprise configuration data.
[0556] • the management subsystem can retrieve relevant data of the POS configuration.
[0557] • the management subsystem can operate without synchronisation with a POS system, including via fallback modes such as cached configuration, messaging, email delivery, Bluetooth transmission, or offline queuing.
[0558] • the management subsystem can hold media assets or retrieve media elements from a video catalogue or remote server.
[0559] • the management subsystem can initiate voice or video calls or other staff notification actions, including automated alerts, escalations, or service requests.
[0560] System
[0561] • the system includes a user profile storage component configured to store preferences, interaction history, and personalisation data.
[0562] • the system includes a video or media catalogue containing videos, images, previews, or other multimedia assets retrievable by the system.
[0563] • the system includes a data store containing item or service data, including ingredients or components, availability or inventory levels, popularity metrics, pricing, or other contextual metadata.
[0564] • the system can be accessed via QR code.
[0565] • the system can be accessed via NFC or message link. the system can be optimised for ordering hospitality services, retail services, travel services, entertainment services, or other customer facing service environments.
[0566] General features
[0567] • the system can be accessed by an untethered device with a screen, such as a mobile phone, tablet, smartwatch, or XR / AR / VR headset.
[0568] • the system can be accessed by a tethered device with a screen, such as a kiosk, tabletop terminal, desktop computer, or in-room system.
[0569] • the system retrieves media elements from a media catalogue stored on: a remote server, the access device, an edge-computing node, or any hybrid arrangement.
[0570] • the system operates with a speech recognition or audio analysis component located on a server or remote system, the device itself, an edge-based system, or any distributed configuration.
[0571] Method
[0572] A computer implemented method for synchronised conversational user engagement, the method comprising:
[0573] (a) receiving user inputs including voice, text, and / or selection-based inputs;
[0574] (b) generating conversational output using a large language model (LLM) or Al model based on the user inputs;
[0575] (c) dynamically selecting or generating one or more media elements relevant to the conversational output;
[0576] (d) presenting or displaying the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output.
[0577] Computer readable medium
[0578] A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the processor(s) to:
[0579] (a) receive user inputs including voice, text, and / or selection-based inputs;
[0580] (b) generate conversational output using a large language model (LLM) or Al model based on the user inputs; (c) dynamically select or generate one or more media elements relevant to the conversational output;
[0581] (d) present or display the conversational output together with the media elements via an output module such that the media elements are temporally aligned with the conversational output.
[0582] Example of use case applications include, but are not limited to:
[0583] The system is a food and / or drink ordering system.
[0584] The system is a travel advice or guidance system.
[0585] The system is a concierge advice or guidance system.
[0586] The system is a tourism advice or guidance system.
[0587] The system is a social media system.
[0588] The system is a retail product browsing or shopping assistance system.
[0589] The system is a booking or reservation management system.
[0590] The system is customer support or helpdesk assistance system.
[0591] The system is a navigation or wayfinding assistant system.
[0592] The system is social media content generation or publishing system.
[0593] The system is an educational or training assistance system.
[0594] The system is a healthcare guidance or patient assistance system.
[0595] The system is a general purpose Al-based conversational system.
[0596] Note
[0597] It is to be understood that the above-referenced arrangements are only illustrative of the application for the principles of the present invention. Numerous modifications and alternative arrangements can be devised without departing from the spirit and scope of the present invention. While the present invention has been shown in the drawings and fully described above with particularity and detail in connection with what is presently deemed to be the most practical and preferred example(s) of the invention, it will be apparent to those of ordinary skill in the art that numerous modifications can be made without departing from the principles and concepts of the invention as set forth herein.
Claims
1. CLAIMS1. A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:(a) receive user inputs including voice, text, and / or selection-based inputs;(b) generate conversational output using a large language model (LLM) or Al model in response to the user inputs;(c) dynamically select or generate one or more media elements relevant to the conversational output, the system having access to an operator defined media library comprising media assets; and(d) present or display the conversational output together with the one or more media elements via an output module such that the one or more media elements are temporally aligned with the conversational output.
2. The system of claim 1, wherein the temporal alignment is based on synchronisation data, wherein synchronisation data comprises any information or metadata used to determine, control, or adjust the temporal alignment between any portion of the conversational output and any portion of the one or more media elements.
3. The system of claim 2, wherein synchronisation data comprises timestamps or markers, playback instructions, ordering rules, semantic rules, or any metadata indicating when a media element should start, stop, or transition.
4. The system of any of claim 2-3, wherein synchronisation data further specifies timing, sequencing, overlap, or transition rules for a plurality of video elements synchronised with different segments of the conversational output.
5. The system of any of claim 2-4, wherein synchronisation data is generated in advance, during, or after generation of the conversational output or text to speech output.
6. The system of any of claim 2-5, wherein synchronisation data is embedded within, associated with, or generated separately from the conversational output.
7. The system of any of claim 2-6, wherein synchronisation data is represented as tokens, structured metadata, vectors, or semantic annotations.
8. The system of claim 1, wherein temporal alignment is achieved without explicit synchronisation data.
9. The system of claim 8, wherein temporal alignment is based on implicit or approximate triggers including one or more of: conversational start time, keyword detection, entity recognition, semantic grouping, heuristic timing rules, or predicted timing based on the conversational output.
10. The system of claim 8, wherein temporal alignment is based on detection of keywords or entities within the user input prior to the generation of the conversational output.
11. The system of claim 8, wherein temporal alignment is based on classification of the conversational output.
12. The system of claim 8, wherein temporal alignment is based on probability, or scoring metrics produced by the LLM.
13. The system of any preceding claim, wherein the LLM or Al model is configured with contextual data comprising information relevant to the system’s specific domain of use.
14. The system of any preceding claim, wherein the LLM or Al model is configured to retrieve data generated or obtained during operation of the system, such as: user data, popularity metrics of menu items; user demographic or localisation data; environmental conditions including ambient noise level; and temporal conditions such as time-of-day or promotional period.
15. The system of any preceding claim, wherein the system is configured to present multiple media elements simultaneously such as using overlay, split screen, or picture in picture format.
16. The system of any preceding claim, wherein the LLM or Al model is configured to modify its conversational strategy based on detected system conditions, such as latency, available bandwidth, device capabilities, or sensor data.
17. The system of any preceding claim, wherein the system is configured to collect analytics data relating to user interactions with synchronised media, and to adapt future media selections using reinforcement-learning-based ranking.
18. The system of any preceding claim, wherein the system is configured to select or replace media elements in real time based on inferred or detected user attention or feedback.
19. The system of any preceding claim, wherein the system is configured to pre-cache or prefetch predicted media elements based on contextual data.
20. The system of any preceding claim, wherein the system uses presence detection to activate or deactivate the output module using a camera, depth sensor, radar, PIR sensor, mmWave, or other sensor to activate or deactivate the output module.
21. The system of any preceding claim, wherein initiation of a video element is performed substantially contemporaneously with presentation of a first word or other portion of the conversational output associated with that video, with an allowable phase margin including a predefined time duration, such as ± 0.2s.
22. The system of any preceding claim, wherein the system is configured to initiate provisional playback of a media element prior to receiving synchronisation data and to override or adjust the provisional playback when synchronisation data becomes available.
23. The system of any preceding claim, wherein the one or more processing units are configured to update or modify the playback timing of the media elements dynamically during generation of the text or audio output to compensate for latency or content changes.
24. The system of any preceding claim, wherein the system is configured to adjust the playback position based on a detected timing drift, lag, or other computed discrepancies between predicted and / or text or speech duration.
25. The system of any preceding claim, wherein the provisional playback is based on a predicted or preliminary media selection, a first ranked media element, or a previously cached element.
26. The system of any preceding claim, wherein the system is configured to update, correct, interrupt, or replace provisional playback upon receiving synchronisation data without interrupting or restarting the conversational session.
27. The system of any preceding claim, wherein the system includes a fallback mode when synchronisation data is delayed or unavailable or not generated.
28. The system of any preceding claim, wherein during fallback mode, the system initiates a provisional playback using a non-item specific media element, such as locally cached previews, stock video, introductory clip, static imagery, graphics, textual description, or any other default media asset.
29. The system of any preceding claim, wherein the system includes an auxiliary Al model configured to analyse early conversational segments to identify or predict a suitable provisional media element when synchronisation data is delayed or unavailable.
30. The system of any preceding claim, wherein the conversational output is generated or provided in segments.
31. The system of any preceding claim, wherein the system is configured to override or replace a provisional media element upon receipt of later conversational segments that identify a more relevant media element32. The system of any preceding claim, wherein the system detects keywords, entities, or semantic cues within individual conversational output chunks and initiates media selection or playback in response to those cues.
33. The system of any preceding claim, wherein the system estimates speech duration or predicted audio timing for each conversational segment and uses the estimate to align or adjust the timing of associated media elements.
34. The system of any preceding claim, wherein the system computes timing deviations between predicted timing (based on earlier segments) and updated timing (based on later segments) and adjusts media playback to correct for drift.
35. The system of any preceding claim, wherein the conversational segments correspond to different media elements, and the system schedules or re-schedules the media elements according to the sequence in which the segments are generated.
36. The system of any preceding claim, wherein the conversational segments are generated from a streaming system.
37. The system of any preceding claim, wherein if no relevant media is identified in an early conversational segment, the system initiates a fallback or default media element until a later segment provides sufficient content to determine a specific media element.
38. The system of any preceding claim, wherein the system adjusts the synchronisation updates based on the rate, latency, or size of conversational output segments.
39. The system of any preceding claim, wherein the system is configured to update or refine the synchronisation data as additional portions of the conversational output are generated in incremental segments.
40. The system of any preceding claim, wherein the system is configured to adjust playback of one or more media elements based on the updated synchronisation data.
41. The system of any preceding claim, wherein the system includes an intent management module configured to: identify, from the user input or the conversational output, an actionable intent relating to a transaction or service event, and to generate one or more action prompts in natural language and / or multimedia form.
42. The system of preceding claim 41, wherein the intent management module is implemented using the LLM, an auxiliary Al model, a symbolic planner, a rule-based engine, or any hybrid thereof.
43. The system of any preceding claim 41 - 42, wherein upon user confirmation, the intent management module trigger execution of the identified intent through integration with one or more external systems or APIs, including booking, reservation, transport, payment or delivery services; and dynamically update the conversation state and synchronised media presentation based on the status returned from the external systems.
44. The system of any preceding claim, wherein the system includes a dynamic trip generation module configured to generate a sequence of activities, locations, items, or service based on the conversational output, and to associate each element of the sequence of activities with one or more corresponding media elements.
45. The system of any preceding claim, wherein the system further comprises a brand specific voice agent module configured to:(i) operate an Al-powered voice synthesis module trained on voice models or samples from multiple different sources, in which the voice models or samples from each source are defined by user-recognisable characteristics that are specifically associated with each source (such as a brand, brand persona, or service provider identity);(ii) enable the LLM to generate conversational output rendered in a synthesized voice specific to one of the sources; and(iii) deploy the synthesised voice across multiple different types of channels, including any one or more of the following: kiosks, telephone-based services, web-based services, mobile applications, in-restaurant devices, and in-hotel room devices.
46. The system of any preceding claim, wherein the system is configured to automatically generate, edit or convert analysed interaction data into a social media content format optimised for distribution on social media platforms, such as Instagram, TikTok, or WhatsApp.
47. The system of any preceding claim 46, wherein analysing the interaction data includes performing sentiment analysis on user reviews, user provided text, or other subjective content to identify positive sentiment for inclusion in the generated social media content.
48. The system of any preceding claim, wherein the system is configured to retrieve one or more media elements associated with the selected menu items or user interaction history and incorporate the one or more media elements into the social media content.
49. The system of any preceding claim, wherein the system provides a user interface enabling a user to select, approve, exclude, or modify specific aspects of the user experience to be incorporated into social media content.
50. The system of any preceding claim, wherein the system is configured to automatically schedule distribution of generated social media content to one or more social media platforms.
51. The system of preceding claim 50, wherein scheduling of distribution is determined based on user preferences, historical posting behaviour, or predicted engagement metrics.
52. The system of any preceding claim, wherein the system is configured to automatically request a user to leave a review after an interaction or experience.
53. The system of any preceding claim, wherein the system includes an output module that automatically publishes or provides a preview of the social media content for the user or brand to review and / or share.
54. The system of any preceding claim, wherein priority is given to the operator-defined media library comprising media assets when selecting or generating said media elements.
55. The system of any preceding claim, wherein the media assets are stored with associated metadata, and wherein the metadata associated with each of the one or more media element comprises timing data, semantic vectors, operator approval flags, or media-type identifiers.
56. The system of any preceding claim, wherein the media library further comprises mediageneration inputs including one or more of: templates, style parameters, seed vectors, or textual description.
57. The system of any preceding claim, wherein the system generates new media elements using a generative Al model, and stores the generated media elements with associated metadata in the operator-defined media library.
58. The system of any preceding claim, wherein the system is configured to generate and present one or more Al-generated interactive prompts being dynamically generated based on at least one of: the user input, the conversational output, contextual data, or metadata associated with a media element.
59. The system of claim 58, wherein the interactive prompts suggest follow-up queries, recommended actions, or conversational options, the interactive prompts.
60. The system of any of claim 58-59, wherein the interactive prompts comprise one or more interface elements including: prompt bubbles, tiles, cards, or other selectable UI components.
61. The system of any of claim 58-60, wherein the interactive prompts are temporally aligned with the conversational output or with the presentation of the one or more media elements.
62. The system of any of claim 58-61, wherein the interactive prompts include selectable actions relating to information browsing, recommendations, bookings, reservations, or orders.
63. The system of any preceding claim, wherein the system detects timing drift between predicted duration of the conversational output and an actual duration of the conversational output, and to dynamically adjust playback of the one or more media elements to maintain temporal synchronisation.
64. The system of any preceding claim, wherein the large language model is executed using a partitioned inference pipeline in which a first portion of the model is deployed on an edgecomputing node near the output module and a second portion of the model is executed on a remote server.
65. The system of any preceding claim, the system being configured to merge intermediate inference results generated by the respective portions of the model to reduce network latency and computational load.
66. The system of any preceding claim, wherein the system further comprises a speech recognition or audio analysis component deployed on a remote server, on the access device, on an edge-based system, or across a distributed architecture.
67. The system of any preceding claim, wherein the output modules include any one or more of the following: screens, speakers, projectors, kiosks, tablets, smartphones, wearables, AR interfaces, VR headset, mixed reality interface, or multi-screen display.
68. The system of any preceding system claim, wherein the system is a food and / or drink ordering system.
69. The system of any preceding system claim, wherein the system is a travel advice or guidance system.
70. The system of any preceding system claim, wherein the system is a concierge advice or guidance system.
71. The system of any preceding system claim, wherein the system is a tourism advice or guidance system.
72. The system of any preceding system claim, wherein the system is a social media content generation or publishing system.
73. The system of any preceding system claim, wherein the system is a retail product browsing or shopping assistance system.
74. The system of any preceding system claim, wherein the system is a booking or reservation management system.
75. The system of any preceding system claim, wherein the system is a customer support or helpdesk assistance system.
76. The system of any preceding system claim, wherein the system is a navigation or wayfinding assistant system.
77. The system of any preceding system claim, wherein the system is an educational or training assistance system.
78. The system of any preceding system claim, wherein the system is a healthcare guidance or patient assistance system.
79. The system of any preceding system claim, wherein the system is a general-purpose AI- based conversational system.
80. A computer implemented method for synchronised conversational user engagement, the method comprising:(a) receiving user inputs including voice, text, and / or selection-based inputs;(b) generating conversational output using a large language model (LLM) or Al model in response to the user inputs;(c) dynamically selecting or generating one or more media elements relevant to the conversational output, the system having access to an operator defined media library comprising media assets;(d) presenting or displaying the conversational output together with the one or more media elements via an output module such that the media elements are temporally aligned with the conversational output.
81. A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the processor(s) to:(a) receive user inputs including voice, text, and / or selection-based inputs;(b) generate conversational output using a large language model (LLM) or Al model in response to the user inputs;(c) dynamically select or generate one or more media elements relevant to the conversational output, the system having access to an operator defined media library comprising media assets;(d) present or display the conversational output together with the one or more media elements via an output module such that the media elements are temporally aligned with the conversational output.
82. A customer facing device for interactive food or drink ordering, the device comprising a display, one or more processors, and input / output module, the device being configured to:(i) receive user inputs including voice, text, touch, and / or selection based inputs;(ii) activate or interact with a conversational LLM-based Al system that is configured with a context that defines features or characteristics of a food or hospitality service provider;(iii) present synchronised conversational output and one or more media elements relevant to the conversational output, the device having access to an operator defined media library comprising media assets;(iv) display the one or more media elements and the conversational output such that the media elements are temporally aligned with the conversational output;(v) interpret user input and / or conversational output to extract a food and / or drink order; and(vi) transmit the extracted order to a point of sale system or kitchen display system for fulfilment.
83. The device of claim 82, wherein the device is a tablet, a user’ s mobile phone, or a personal mobile device.
84. The device of any of claim 82-83, wherein the device is configured to access the conversational Al system by scanning a QR code, and wherein the QR code loads a browser based interface, mobile web application, or native application.
85. The device of any of claim 82-84, wherein the device is configured to present a bill or itemised cost to the customer and receive payment through the device via a card reader, QR code, contactless payment, or mobile payment interface.
86. A website-based application for synchronised conversational user engagement, the website being delivered to a user device via a web browser and comprising executable instructions configured to:(i) receive user inputs including voice, text, touch, and / or selection based inputs;(ii) activate or interact with a conversational LLM-based Al system that is configured with a context that defines features or characteristics of a food or hospitality service provider;(iii) generate conversational output comprising text and / or speech based on the user inputs via the conversational Al system;(iv) dynamically select or generate one or more media elements relevant to the conversational output, the application having access to an operator defined media library comprising media assets;isEpj(v) present the conversational output together with the one or more media elements within the web browser such that the media elements are temporally aligned with the conversational output; andisEpj(vi) enable the user to browse items or services.
87. The website-based application of claim 86, wherein the website is accessed by scanning a QR code displayed within a restaurant or hospitality venue.
88. The website-based application of any of claim 86-87, wherein the website is accessible via a link provided in an online search result, mapping service interface, or third-party discovery platform.
89. The website-based application of any of claim 86-88, wherein the website is configured to operate in a browsing-only mode providing information, recommendations, or guidance without accepting orders.
90. The website-based application of any of claim 86-89, wherein the website is configured to accept food and / or drink orders and transmit such orders to a point-of-sale system, kitchen display system, or bar system.
91. A system for video selection, comprising:(i) a video store containing media tagged with location or point of interest metadata; and(ii) a LLM that is configured to (a) receive user inputs defining a journey or an itinerary, including location or point-of-interest requirements and to (b) process those inputs as prompts, and (c) to generate an itinerary and to automatically provide video previews associated with locations or point-of-interest; (d) to respond conversationally to further prompts from the user to refine the itinerary; and(iii) a display interface that presents an interactive itinerary alongside the associated video previews.
92. An Al-powered food or drink ordering kiosk device that is configured to:(i) activate automatically upon detecting a customer’s presence, such as by using a camera or a presence detection sensor;(ii) activate or interact with a conversational LLM-based Al system that is configured with a context that defines features or characteristics of a food or hospitality service provider;(iii) automatically initiate or continue a conversation with the customer via voice and / or relevant media content, in response to a customer prompt;(iv) automatically analyse the conversation with the customer to extract a food and / or drink order from the conversation;(v) generate a bill or cost for that order and present that orally and / or visually to the customer;(vi) automatically receive and process payment for that order; and(vii) then send that food and / or drink order to a bar or kitchen.
93. A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:(a) receive user inputs including voice, text, and / or selection-based inputs;(b) generate conversational output using a large language model (LLM) or Al model in response to the user inputs;(c) dynamically select or generate one or more media elements relevant to the conversational output, the system having access to an operator defined media library comprising media assets;(d) present or display the conversational output together with the one or more media elements via an output module such that the one or more media elements are temporally aligned with the conversational output, wherein the temporal alignment is based on synchronisation data, the synchronisation data comprising any information or metadata usable to determine, control, or adjust the temporal alignment between any portion of the conversational output and any portion of the one or more media elements; and wherein the system is further configured to initiate provisional playback of a media element prior to receiving the synchronisation data and to override or adjust the provisional playback when the synchronisation data becomes available.
94. A system for synchronised conversational user engagement, the system comprising one or more processing units configured to:(a) receive user inputs including voice, text, and / or selection-based inputs;(b) generate conversational output in incremental segments using a large language model (LLM) or Al model in response to the user inputs;(c)for each conversational segment, identify one or more candidate media elements relevant to a content portion of the segment based on metadata associated with media assets stored in an operator-defined media library; and(d) present the conversational output together with at least one of the candidate media elements via an output module, wherein the one or more processing units are further configured to update the media element selection when later conversational segments provide additional information.