Bar interactive ordering management system based on augmented reality

By using multimodal data acquisition and fusion recognition technology, combined with AI recommendation and environmental adaptive adjustment, the problems of recognition accuracy and blind recommendation in AR ordering systems in bar scenarios have been solved, achieving high-precision recognition and personalized recommendations, thus improving user experience and efficiency.

CN121685200APending Publication Date: 2026-03-17BEIJING HOLOGRAPHIC JULANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511907328.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing AR ordering systems suffer from decreased recognition accuracy and stability in bar scenarios with complex lighting, noisy backgrounds, and dynamically changing environments. They struggle to distinguish subtle gestures, lack personalized recommendations and environmental adaptability, resulting in a poor user experience.

Method used

The system employs a multimodal data acquisition module to acquire visual, auditory, and 3D spatial information. Through data preprocessing and a multimodal fusion recognition module, it achieves high-precision recognition. Combined with an AI recommendation engine, it provides personalized recommendations. Furthermore, through an environment adaptation and user interaction module, it adjusts display parameters and interaction sensitivity. Finally, it integrates a backend management and AR display module to provide intuitive feedback.

Benefits of technology

Achieve high-precision target recognition and personalized recommendations in complex environments, improve user experience and consumption efficiency, and enhance the system's adaptability and smoothness of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685200A_ABST
    Figure CN121685200A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of catering management systems, in particular to a bar interactive ordering management system based on augmented reality, comprising: A1, user front-end equipment integrated with a multi-modal data acquisition module for displaying augmented reality content and acquiring visual information, auditory information and three-dimensional space information in an environment in real time; and A2, a data preprocessing module which is connected with the multi-mode data acquisition module. Visual, auditory and three-dimensional space information is acquired through the multi-modal data acquisition module, and is subjected to deep fusion processing through the multi-modal fusion recognition module after being optimized through the data preprocessing module, so that the defect that single visual recognition is insufficient in precision in a bar weak light and shielding scene is avoided, and high-precision recognition of a target object, user gestures and environmental characteristics is realized; by means of an AI recommendation engine in combination with user historical consumption records, real-time preferences, inventory information and environmental factors, personalized drink or package recommendation is provided, and the problem of blind recommendation of a traditional system is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of restaurant management system technology, specifically to an interactive bar ordering management system based on augmented reality. Background Technology

[0002] With the rapid development of augmented reality (AR) technology, its applications in commerce and entertainment are becoming increasingly widespread, especially in the catering industry, where AR ordering systems are beginning to emerge. Traditional AR ordering systems typically use visual recognition technology (such as camera-captured images) to identify table signs, QR codes, or physical drinks, and then display virtual menus or product information on the user's device. However, these systems face many challenges in practical applications, especially in scenarios like bars with complex lighting (dim, vibrant colors), noisy backgrounds, and dynamically changing environments.

[0003] Existing AR systems primarily rely on single-vision recognition. In low light or obstructed conditions, recognition accuracy and stability significantly decrease, leading to a clunky user experience or even system malfunction. Furthermore, single-vision recognition struggles to effectively distinguish subtle gestures or understand user intent. Current ordering systems often lack sufficient intelligence, failing to provide personalized recommendations based on real-time user preferences, purchase history, inventory status, or even the environment. This results in unpredictable recommendations, reducing user satisfaction and efficiency. Moreover, these systems lack adaptability to complex environmental changes, failing to dynamically adjust display effects and interaction sensitivity based on ambient light, noise, and other factors, further impacting usability. Therefore, developing an AR ordering management system with high-precision recognition capabilities in complex environments, capable of providing personalized intelligent recommendations and environmental adaptability, is a pressing technical challenge. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an augmented reality-based interactive bar ordering management system to solve the problems mentioned in the background section.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an augmented reality-based interactive bar ordering management system, comprising: A1. User front-end device: The user front-end device integrates a multimodal data acquisition module, which is used to display augmented reality content and collect visual, auditory and three-dimensional spatial information in the environment in real time. A2. Data preprocessing module, connected to the multimodal data acquisition module, is used to perform noise reduction, enhancement and calibration processing on the acquired raw visual information, raw auditory information and raw three-dimensional spatial information; A3. Multimodal fusion recognition module, connected to the data preprocessing module, is used to receive preprocessed visual information, auditory information and three-dimensional spatial information, and to perform deep fusion processing on the visual information, auditory information and three-dimensional spatial information to achieve high-precision recognition of target objects, user gestures and environmental features. A4. Environmental Adaptation and User Interaction Module, connected to the Multimodal Fusion Recognition Module, is used to adaptively adjust the display parameters of AR content and the sensitivity of user interaction based on the environmental status information output by the Multimodal Fusion Recognition Module, and to parse the commands issued by the user through gestures or voice. A5, an AI recommendation engine, is connected to the environment adaptation and user interaction module, as well as to user preferences, historical databases, and beverage and food databases. It is used to intelligently and personally recommend beverages or set meals based on users' historical consumption records, real-time preferences, inventory information, and environmental factors. A6. The back-end management and service module, connected to the environment adaptation and user interaction module and the AI ​​recommendation engine, is used to receive ordering instructions, update real-time inventory, generate orders and distribute them to the waiter's terminal, as well as process payment requests and connect with third-party payment interfaces. A7, the AR display and feedback module, connects with the environment adaptation and user interaction module and the AI ​​recommendation engine. It is used to overlay virtual information onto the real world in real time through the user's front-end device based on the instructions of the environment adaptation and user interaction module and the recommendation results of the AI ​​recommendation engine.

[0006] Furthermore, the multimodal data acquisition module includes: The camera unit is used to capture video streams and image information from the environment in real time; The microphone unit is used to collect ambient sound information and user voice commands; The depth sensor unit is used to acquire three-dimensional depth information of the scene.

[0007] Furthermore, the data preprocessing module includes: The image processing unit is used to perform noise reduction, color correction, and adaptive brightness adjustment on the image data captured by the camera unit. The audio processing unit is used to perform noise reduction, speech separation, and volume normalization on the audio data collected by the microphone unit. The depth data processing unit is used to calibrate and remove artifacts from the depth data acquired by the depth sensor unit.

[0008] Furthermore, the multimodal fusion recognition module employs a deep learning model to fuse visual information, auditory information, and three-dimensional spatial information. The fusion processing includes feature-level fusion and decision-level fusion to identify target objects, user gestures, and assess ambient light intensity and noise levels.

[0009] Furthermore, the multimodal fusion recognition module includes: The feature extraction unit is used to extract features of each modality from preprocessed visual information, auditory information, and three-dimensional spatial information; The feature fusion unit is used to concatenate or interactively fuse the extracted multimodal features; The recognition and classification unit is used to identify target objects, user gestures, and environmental features based on the fused features.

[0010] Furthermore, the environment adaptation and user interaction module includes: The environmental state perception unit is used to receive environmental state information output by the multimodal fusion recognition module. The environmental state information includes ambient light intensity and ambient noise level. The display parameter adjustment unit is used to adjust the display brightness, contrast, or transparency of AR content based on the ambient light intensity obtained by the environmental state perception unit. An interaction sensitivity adjustment unit is used to adjust the sensitivity of user interaction based on the ambient noise level obtained by the environmental state perception unit. The instruction parsing unit is used to parse the ordering or browsing instructions issued by the user through gestures or voice.

[0011] Furthermore, the AI ​​recommendation engine utilizes machine learning algorithms, including collaborative filtering, content-based recommendation algorithms, or deep learning recommendation models. The AI ​​recommendation engine makes intelligent recommendations based on user preferences, historical consumption records and preferences in the historical database, real-time inventory information in the beverage / food database, and current environmental factors.

[0012] Furthermore, the backend management and service module includes: The order generation unit is used to integrate user orders into a food order; The inventory update unit is used to send inventory update instructions to the beverage and food databases; The service distribution unit is used to distribute order information to the waiter's terminal; and The payment processing unit is used to receive users' payment requests and connect with third-party payment interfaces to complete the payment.

[0013] Furthermore, the AR display and feedback module includes: The virtual information generation unit is used to generate virtual information such as three-dimensional virtual wine models, personalized recommendation lists, or interactive menus based on the instructions of the environment adaptation and user interaction module and the recommendation results of the AI ​​recommendation engine. The real-time overlay unit is used to overlay generated virtual information onto the real world in real time via the user's front-end device.

[0014] Furthermore, the user front-end device is AR glasses, an AR tablet, or a smartphone equipped with a high-performance camera, microphone, and depth sensor.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention acquires visual, auditory, and 3D spatial information through a multimodal data acquisition module. After optimization by a data preprocessing module, the data is deeply fused and processed by a multimodal fusion recognition module. This avoids the insufficient accuracy of single visual recognition in low-light or occluded environments like bars, achieving high-precision recognition of target objects, user gestures, and environmental features. Leveraging an AI recommendation engine that combines user historical consumption records, real-time preferences, inventory information, and environmental factors, it provides personalized beverage or set meal recommendations, solving the problem of blind recommendations in traditional systems. An environmental adaptation and user interaction module adjusts AR display parameters based on ambient light intensity and interaction sensitivity based on noise levels, enhancing the system's adaptability in complex environments. Simultaneously, an AR display and feedback module provides intuitive virtual information, while a backend management and service module efficiently processes orders, inventory, and payments, comprehensively improving user experience and consumption efficiency. Attached Figure Description

[0016] Figure 1 This is a system architecture diagram of an interactive bar ordering management system based on augmented reality, according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 This invention provides an augmented reality-based interactive bar ordering management system, comprising: A1. User front-end device: The user front-end device integrates a multimodal data acquisition module, which is used to display augmented reality content and collect visual, auditory and three-dimensional spatial information in the environment in real time. A2. Data preprocessing module, connected to the multimodal data acquisition module, is used to perform noise reduction, enhancement and calibration processing on the acquired raw visual information, raw auditory information and raw three-dimensional spatial information; A3. Multimodal fusion recognition module, connected to the data preprocessing module, is used to receive preprocessed visual information, auditory information and three-dimensional spatial information, and to perform deep fusion processing on the visual information, auditory information and three-dimensional spatial information to achieve high-precision recognition of target objects, user gestures and environmental features. A4. Environmental Adaptation and User Interaction Module, connected to the Multimodal Fusion Recognition Module, is used to adaptively adjust the display parameters of AR content and the sensitivity of user interaction based on the environmental status information output by the Multimodal Fusion Recognition Module, and to parse the commands issued by the user through gestures or voice. A5, an AI recommendation engine, is connected to the environment adaptation and user interaction module, as well as to user preferences, historical databases, and beverage and food databases. It is used to intelligently and personally recommend beverages or set meals based on users' historical consumption records, real-time preferences, inventory information, and environmental factors. A6. The back-end management and service module, connected to the environment adaptation and user interaction module and the AI ​​recommendation engine, is used to receive ordering instructions, update real-time inventory, generate orders and distribute them to the waiter's terminal, as well as process payment requests and connect with third-party payment interfaces. A7, the AR display and feedback module, connects with the environment adaptation and user interaction module and the AI ​​recommendation engine. It is used to overlay virtual information onto the real world in real time through the user's front-end device based on the instructions of the environment adaptation and user interaction module and the recommendation results of the AI ​​recommendation engine.

[0019] In this embodiment, AR glasses are selected as the user front-end device. The integrated multimodal data acquisition module can capture bar environment information in real time: the camera captures the placement of drinks on the table, user gestures and ambient light, the microphone collects user voice commands and filters some background music noise, and the depth sensor constructs a three-dimensional model of the table and surrounding space. The collected raw data is transmitted to the data preprocessing module. In response to the problem of blurred images caused by the dim lighting and chaotic lighting in the bar, Gaussian filtering is first used to denoise the visual information, then automatic white balance processing is used to correct the color deviation, and finally brightness adaptive adjustment is performed. The brightness adjustment is calculated using formula (1) to ensure that the image is clear and distinguishable under different lighting conditions.

[0020]

[0021] The parameters in formula (1) are defined as follows: The original brightness value of the image captured by the camera is dimensionless. After being normalized by Min-Max, the value ranges from 0 to 1. It is obtained by calculating the arithmetic mean of the gray values ​​of all pixels in the image. The system's preset target brightness value is dimensionless and is set to 0.5 based on the user's visual comfort in a low-light environment in a bar. It has been normalized. Brightness adjustment coefficient, dimensionless, range of values. Dynamically determined, when To enhance the brightness increase, when hour To preserve the original lighting and shadow details of the image; The adjusted image brightness value is dimensionless and ranges from 0 to 1 after normalization. It is directly used for subsequent image feature extraction.

[0022] For auditory information, environmental noise is removed using spectral subtraction to preserve clear user speech; for three-dimensional spatial information, coordinate system calibration is used to eliminate artifacts caused by occlusion from tables and chairs. The preprocessed data is then fed into a multimodal fusion recognition module, which employs a CNN-RNN hybrid deep learning model to first extract visual texture features. Auditory speech features and three-dimensional spatial structural features Each feature is normalized to the 0-1 range by Min-Max to ensure consistency of dimensions, and then the comprehensive feature is obtained by feature-level weighted fusion. The fusion process is calculated using formula (2): The parameters in formula (2) are defined as follows: The fused multimodal feature vector is dimensionless and has dimensions N+M+P, where N is the visual feature dimension, M is the auditory feature dimension, and P is the three-dimensional feature dimension. Visual feature weights, dimensionless. When there is sufficient sunlight In low light It is dynamically adjusted by the ambient light intensity; The preprocessed visual feature vector is dimensionless and extracted by CNN. It contains information such as wine label texture and gesture contour, and has been normalized. Auditory feature weights, dimensionless. When the noise is low When the noise is high It is dynamically adjusted based on the ambient noise level; The preprocessed auditory feature vector is dimensionless, extracted by RNN, and contains the temporal information of the user's voice commands. It has been normalized. : Three-dimensional feature weights, dimensionless And satisfy When there is little obstruction in the space obscured for a long time ; The preprocessed 3D feature vector is dimensionless, extracted by PointNet, and contains spatial location information of the desktop and hand, and has been normalized.

[0023] Integrating a decision-making voting mechanism, this module can accurately recognize user gestures, such as waving to select drinks, target drinks, and ambient light and noise levels, solving the accuracy problems of traditional single-vision recognition in low light or occlusion conditions. The environmental adaptation and user interaction module receives environmental status information; if the lighting is dim and the light intensity does not exceed 100 lux, it increases the brightness and contrast of the AR display; if the noise level is high and not lower than 60dB, it increases the voice wake-up threshold. It also analyzes user voice commands, such as "order a whiskey," or gesture commands. The AI ​​recommendation engine connects to the user's historical database, combining drink inventory and ambient atmosphere (e.g., recommending a nut combo when listening to soothing music) to generate personalized recommendations. The backend management and service module receives instructions, generates orders, updates inventory, distributes orders to waiter terminals, and connects to third-party payment interfaces to complete payments. The AR display and feedback module overlays a 3D virtual drink model and recommendation list onto the real desktop, allowing users to preview the effect intuitively, thus improving overall recognition stability, interaction smoothness, and recommendation accuracy.

[0024] Furthermore, the multimodal data acquisition module includes: The camera unit is used to capture video streams and image information from the environment in real time; The microphone unit is used to collect ambient sound information and user voice commands; The depth sensor unit is used to acquire three-dimensional depth information of the scene.

[0025] In this embodiment, the camera unit of the multimodal data acquisition module uses a wide-angle high-definition industrial camera. The wide-angle lens can cover the table and the surrounding half-meter range. The 1080P resolution can capture details of the beverage label. The F1.8 large aperture design can still obtain sufficient light in the low-light environment of the bar, ensuring clear video stream and image information. The microphone unit adopts a four-array microphone, which focuses on the user's location through directional sound pickup technology. The sound pickup angle is controlled within ±30°, effectively filtering the conversation and background music of other tables in the bar, and accurately collecting the user's voice commands. The depth sensor unit uses a TOF type sensor. This type of sensor does not rely on ambient light and can quickly acquire three-dimensional depth information of the table, chair and user's hand in dim scenes. The frame rate reaches 15fps, and the error of the generated spatial point cloud data does not exceed 2cm, providing reliable data support for subsequent 3D modeling and gesture recognition, and improving the comprehensiveness and accuracy of data acquisition.

[0026] Furthermore, the data preprocessing module includes: The image processing unit is used to perform noise reduction, color correction, and adaptive brightness adjustment on the image data captured by the camera unit. The audio processing unit is used to perform noise reduction, speech separation, and volume normalization on the audio data collected by the microphone unit. The depth data processing unit is used to calibrate and remove artifacts from the depth data acquired by the depth sensor unit.

[0027] In this embodiment, the image processing unit of the data preprocessing module first uses a bilateral filtering algorithm to remove noise caused by the flickering of bar lights for the image captured by the camera. The filtering radius is set to 3×3. Then, the grayscale world algorithm is used for color correction to solve the problem of color distortion of drinks under different colored lights, such as neon lights. Finally, the brightness is automatically adjusted in combination with formula (1) to ensure that the image is always clear and distinguishable in the changes of brightness. The audio processing unit uses wavelet threshold noise reduction to process the audio collected by the microphone. The wavelet basis is selected as db4 to remove environmental noise. Then, the user's voice and background music are separated by independent component analysis technology, and only the user's voice signal is retained. Then, the volume of the voice signal is uniformly adjusted to the standard range of 60-70dB to avoid the impact of different user speaking volumes on subsequent instruction parsing. The depth data processing unit uses Zhang's calibration method to calibrate the data collected by the depth sensor. After calibration, the error does not exceed 1% to eliminate the sensor's own error. At the same time, the region growing algorithm is used to remove the artifact point cloud caused by object occlusion. The growth threshold is set to 2mm to ensure that the three-dimensional depth data can truly reflect the bar's spatial structure and provide high-quality data for multimodal fusion recognition.

[0028] Furthermore, the multimodal fusion recognition module uses a deep learning model to fuse visual information, auditory information, and three-dimensional spatial information. The fusion processing includes feature-level fusion and decision-level fusion to identify target objects, user gestures, and assess ambient light intensity and noise levels.

[0029] In this embodiment, the multimodal fusion recognition module employs a Transformer-based deep learning model to extract features from the preprocessed visual, auditory, and 3D spatial information: edge and texture features of the visual image are extracted using a CNN. Temporal features of auditory speech are extracted using RNN. Extracting geometric features of 3D point clouds using PointNet Each feature is normalized using Min-Max. In the feature-level fusion stage, the attention mechanism of Transformer is used to focus on features related to target recognition, such as the visual features of user gestures and the auditory features of ordering instructions. The three types of features are then weighted and fused using formula (2). In the decision-level fusion stage, the individual recognition results of each modality, such as visual recognition of a whiskey bottle, auditory recognition of ordering whiskey, and 3D recognition of the user's hand pointing to the bottle, are weighted using a voting mechanism. The voting weights are the same as those in formula (2). Maintaining consistency and comprehensively evaluating recognition results, this fusion processing can accurately identify target objects such as various beverages, user gestures such as clicks and swipes, and accurately assess ambient light intensity such as low light and strong light, as well as noise levels such as high noise and low noise. In the complex environment of a bar, it effectively avoids the limitations of single-modal recognition and improves the stability and accuracy of recognition.

[0030] Furthermore, the multimodal fusion recognition module includes: The feature extraction unit is used to extract features of each modality from preprocessed visual information, auditory information, and three-dimensional spatial information; The feature fusion unit is used to concatenate or interactively fuse the extracted multimodal features; The recognition and classification unit is used to identify target objects, user gestures, and environmental features based on the fused features.

[0031] In this embodiment, the feature extraction unit of the multimodal fusion recognition module extracts features from low to high levels from the preprocessed visual image through convolutional and pooling layers. Low-level features include image edges, while high-level features include edge texture features such as the shape of the beverage and label text, ultimately yielding a visual feature vector. The preprocessed auditory speech signal is then normalized using Min-Max. It is first converted into a Mel spectrogram, and then features such as Mel frequency cepstral coefficients are extracted using a convolutional layer. These features reflect the timbre and semantic information of the speech, resulting in an auditory feature vector. The data is then normalized. The preprocessed 3D spatial data is input into the PointNet network to extract global and local features. Global features include the overall structure of the desktop, while local features include key hand points. These features reflect the position and structural relationships of objects in space, resulting in a 3D feature vector. The feature fusion unit performs normalization processing. It employs a cross-attention mechanism to allow features from different modalities to interact, such as associating the visual features of user gestures with the auditory features of voice commands to highlight key features and suppress irrelevant features. The recognition and classification unit uses a Softmax classifier to classify the fused features. The input classifier outputs target object categories such as Cabernet Sauvignon wine and vodka, user gesture categories such as confirm and cancel, and environmental feature categories such as dim lighting and moderate noise, ensuring accurate recognition results and providing a reliable basis for subsequent environmental adaptive adjustments and user interactions.

[0032] Furthermore, the environment adaptation and user interaction module includes: The environmental state perception unit is used to receive environmental state information output by the multimodal fusion recognition module. The environmental state information includes ambient light intensity and ambient noise level. The display parameter adjustment unit is used to adjust the display brightness, contrast, or transparency of AR content based on the ambient light intensity obtained by the environmental state perception unit. An interaction sensitivity adjustment unit is used to adjust the sensitivity of user interaction based on the ambient noise level obtained by the environmental state perception unit. The instruction parsing unit is used to parse the ordering or browsing instructions issued by the user through gestures or voice.

[0033] In this embodiment, the environmental state perception unit of the environment adaptation and user interaction module receives ambient light intensity and noise level data output by the multimodal fusion recognition module in real time, forming a dynamic environmental state report. The display parameter adjustment unit dynamically adjusts the AR content display parameters based on the light intensity data: when the light intensity is lower than a set threshold, it automatically increases the AR display brightness and contrast to prevent the AR content from becoming blurry due to dim lighting; when the light intensity is higher than the threshold, it appropriately reduces the brightness and increases the transparency to prevent the AR content from reflecting light and causing glare. The interaction sensitivity adjustment unit adjusts the interaction sensitivity based on the noise level data: when the noise level is higher than a set value, it increases the wake-up word recognition threshold for voice interaction to reduce false triggering due to environmental noise, while also enhancing the recognition range of gesture interaction to facilitate user operation via gestures; when the noise level is lower, it lowers the voice wake-up threshold so that even soft user commands can be recognized. The command parsing unit uses natural language processing technology to perform semantic analysis on user voice commands, extracting order requests such as "another order of fries"; it uses a skeletal key point recognition algorithm to parse the commands corresponding to user gestures, such as a clenched fist representing confirmation of ordering, ensuring that user interaction commands are accurately parsed and improving the smoothness of interaction.

[0034] Furthermore, the AI ​​recommendation engine utilizes machine learning algorithms, including collaborative filtering, content-based recommendation algorithms, or deep learning recommendation models. The AI ​​recommendation engine makes intelligent recommendations based on user preferences, historical consumption records and preferences in the historical database, real-time inventory information in the beverage / food database, and current environmental factors.

[0035] In this embodiment, the AI ​​recommendation engine employs a multi-algorithm fusion strategy, combining user preferences with historical databases and databases of beverages and dishes to achieve intelligent recommendations: Using a collaborative filtering algorithm, it analyzes the preferences of groups similar to the current user's consumption habits. If such groups have recently frequently ordered a mojito with squid strips, this set meal is included in the candidate recommendations. A content-based recommendation algorithm is used, based on the user's historical consumption records showing a preference for fruit-flavored cocktails, to filter out cocktails containing fruit ingredients, such as strawberry Daiquiri, from the database. Through a deep learning recommendation model, such as DeepFM, it integrates real-time user preferences (e.g., currently browsing low-alcohol beverages), inventory information (e.g., sufficient stock of a certain low-alcohol beverage), and environmental factors (e.g., the bar is hosting a ladies' night event) to rank the candidate recommendations and generate a personalized recommendation list. During the recommendation process, if a beverage is in short supply, it is automatically excluded from the recommendation, ensuring that the recommended content is feasible. Simultaneously, the recommendation results match user needs and scenarios, improving user satisfaction and increasing the order rate of the bar's food and beverages.

[0036] Furthermore, the back-end management and service module includes: The order generation unit is used to integrate user orders into a food order; The inventory update unit is used to send inventory update instructions to the beverage and food databases; The service distribution unit is used to distribute order information to the waiter's terminal; and The payment processing unit is used to receive users' payment requests and connect with third-party payment interfaces to complete the payment.

[0037] In this embodiment, the order generation unit of the back-end management and service module receives the ordering instructions from the user's front end, automatically integrates the user's selected beverages, dishes, and quantities, adds information such as the user's seat number and ordering time, and generates a standardized electronic order to avoid omissions and errors in manual order recording. After the order is generated, the inventory update unit sends an inventory update instruction to the beverage and dish database in real time, deducts the corresponding inventory quantity, and marks items with inventory below the warning value to remind back-end staff to replenish stock and prevent overselling. The service distribution unit uses local area network wireless communication technology to distribute order information to the nearest waiter's mobile terminal, such as a smart bracelet and mobile APP, according to the waiter's area of ​​responsibility. The terminal reminds the waiter to accept the order in real time and displays the user's seat location, making it convenient for the waiter to quickly locate and deliver the order. The payment processing unit supports integration with mainstream third-party payment interfaces such as WeChat Pay and Alipay. After the user confirms the order on the front end, they can be directly redirected to the payment page. After the payment is completed, the payment result is fed back to the back end in real time, and a payment success and order progress notification is sent to the user, reducing user waiting time, improving payment and service efficiency, and facilitating merchants to manage orders and inventory in real time.

[0038] Furthermore, the AR display and feedback module includes: The virtual information generation unit is used to generate virtual information such as three-dimensional virtual wine models, personalized recommendation lists, or interactive menus based on the instructions of the environment adaptation and user interaction module and the recommendation results of the AI ​​recommendation engine. The real-time overlay unit is used to overlay generated virtual information onto the real world in real time via the user's front-end device.

[0039] In this embodiment, the virtual information generation unit of the AR display and feedback module generates diverse virtual information based on instructions from the environment adaptation and user interaction module, such as displaying recommended beverage details and AI recommendation engine results. This includes generating a scaled-down 3D virtual model of the beverage, showcasing its appearance, bottle label, and shape after being poured into a glass; designing a personalized recommendation list as an interactive interface, with each item containing the beverage name, price, and taste description, allowing users to click and view details; and categorizing the generated interactive menu by type, such as cocktails, wine, and snacks, for easy user browsing. The real-time overlay unit uses SLAM (Simultaneous Localization and Mapping) technology, combined with the user's front-end device's position and orientation data, to accurately overlay the generated virtual information onto the real environment. For example, a 3D beverage model is overlaid next to an empty glass on the user's table, allowing for a direct preview of the beverage's appearance; the recommendation list is overlaid on one side of the table without obstructing the user's view, allowing users to operate the virtual interface via gestures or voice, enhancing the visual experience and ease of use, and helping users make ordering decisions more quickly.

[0040] Furthermore, the user's front-end device is AR glasses, AR tablets, or a smartphone equipped with a high-performance camera, microphone, and depth sensor.

[0041] In this embodiment, the user's front-end device can be flexibly selected according to the bar scene and user needs: When using AR glasses, a lightweight design is adopted, weighing no more than 80g, ensuring comfortable wear. Users do not need to hold the device and can operate the AR interface while talking with friends, adapting to the social scene of a bar. Furthermore, the display lenses of the AR glasses can automatically adjust the transmittance according to the ambient light, with a transmittance adjustment range of 30%-80%, ensuring clear AR content. When using an AR tablet, a 10-inch high-definition touchscreen is used, with a resolution of no less than 2560×1600, supporting multiple people to view and operate simultaneously, suitable for multi-person dining scenarios. Users can order food via touch or voice commands. The tablet's built-in high-performance processor has a main frequency of no less than 2.4GHz, which can quickly process multimodal data and AR rendering. When using a smartphone, it is compatible with mainstream Android 10 and above, and iOS 14 and above. Users can enable the system function by downloading a dedicated APP. The phone's built-in high-definition camera with a resolution of no less than 12 million pixels, microphone and gyroscope assist in obtaining spatial information to achieve data collection and AR display, reducing the user threshold, eliminating the need to purchase additional equipment, expanding the system's applicability, and meeting the usage habits of different users and the diverse needs of bars.

[0042] In summary, this invention acquires visual, auditory, and 3D spatial information through a multimodal data acquisition module. After optimization by a data preprocessing module, the data is deeply fused and processed by a multimodal fusion recognition module. This avoids the insufficient accuracy of single visual recognition in low-light or occluded bar scenarios, achieving high-precision recognition of target objects, user gestures, and environmental features. Leveraging an AI recommendation engine that combines user historical consumption records, real-time preferences, inventory information, and environmental factors, personalized beverage or set meal recommendations are provided, solving the problem of blind recommendations in traditional systems. An environmental adaptation and user interaction module adjusts AR display parameters based on ambient light intensity and interaction sensitivity based on noise levels, enhancing the system's adaptability in complex environments. Simultaneously, an AR display and feedback module provides intuitive virtual information, while a backend management and service module efficiently processes orders, inventory, and payments, comprehensively improving user experience and consumption efficiency.

[0043] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0044] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An augmented reality based bar interactive ordering management system characterized in that, Comprise: A1, user front-end equipment, the user front-end equipment is integrated with multi-modal data acquisition module, is used for displaying augmented reality content and real-time collection environment in visual information, auditory information and three-dimensional space information; A2, data preprocessing module, is connected with multi-modal data acquisition module, is used for carrying out noise reduction processing, enhancement processing and calibration processing to the original visual information, original auditory information and original three-dimensional space information collected; A3, multi-modal fusion recognition module, is connected with data preprocessing module, is used for receiving the visual information, auditory information and three-dimensional space information after preprocessing, and carries out depth fusion processing to visual information, auditory information and three-dimensional space information, to realize the high-precision identification of target object, user gesture and environmental characteristics; A4, environmental adaptation and user interaction module, is connected with multi-modal fusion recognition module, is used for according to environmental state information that multi-modal fusion recognition module exports, adaptively adjusts the display parameter of AR content and the sensitivity of user interaction, and analyzes the instruction that user issues through gesture or voice; A5, AI recommendation engine, is connected with environmental adaptation and user interaction module, and is connected with user preference, historical database and liquor, dish database, is used for according to user historical consumption record, real-time preference, inventory information and environmental factors, intelligently personalized liquor or package is recommended; A6, background management and service module, is connected with environmental adaptation and user interaction module and AI recommendation engine, is used for receiving order instruction, updating real-time inventory, generating order and distributing to waiter terminal, and processing payment request and with third party payment interface is connected; A7, AR display and feedback module, is connected with environmental adaptation and user interaction module and AI recommendation engine, is used for according to the instruction of environmental adaptation and user interaction module and the recommended result of AI recommendation engine, through user front-end equipment, real-time virtual information is superimposed in real world.

2. The augmented reality based bar interactive ordering management system as claimed in claim 1, wherein, The multi-modal data acquisition module comprises: A camera unit for capturing video stream and image information in the environment in real time; A microphone unit for collecting environmental sound information and user voice instructions; A depth sensor unit for obtaining three-dimensional depth information of the scene.

3. The augmented reality based bar interactive ordering management system as claimed in claim 2, wherein, The data preprocessing module comprises: An image processing unit for denoising, color correction and brightness adaptive adjustment of image data captured by the camera unit; An audio processing unit for noise reduction, speech separation and volume normalization of audio data collected by the microphone unit; A depth data processing unit for calibration and artifact removal of depth data obtained by the depth sensor unit.

4. The augmented reality based bar interactive ordering management system as claimed in claim 3, wherein, The multi-modal fusion recognition module uses a deep learning model to fuse the visual, auditory and three-dimensional spatial information, wherein the fusion process includes feature-level fusion and decision-level fusion to identify target objects, user gestures and assess ambient light intensity and noise levels.

5. The augmented reality based bar interactive ordering management system as claimed in claim 4, wherein, The multi-modal fusion recognition module comprises: A feature extraction unit for extracting features of each modality from preprocessed visual, auditory and three-dimensional spatial information; The feature fusion unit is configured to splice or interactively fuse the extracted multi-modal features. The recognition and classification unit is configured to recognize target objects, user gestures, and environmental features based on the fused features.

6. The augmented reality based bar interactive ordering management system as claimed in claim 1, wherein, The environment-adaptive and user-interaction module includes: The environment state perception unit is configured to receive environmental state information output by the multi-modal fusion recognition module, the environmental state information including ambient light intensity and ambient noise level. The display parameter adjustment unit is configured to adjust the display brightness, contrast, or transparency of AR content according to the ambient light intensity obtained by the environment state perception unit. The interaction sensitivity adjustment unit is configured to adjust the sensitivity of user interaction according to the ambient noise level obtained by the environment state perception unit. The instruction analysis unit is configured to analyze the ordering or browsing instructions issued by the user through gestures or voice.

7. The augmented reality based bar interactive ordering management system as claimed in claim 1, wherein, The AI recommendation engine utilizes machine learning algorithms, including collaborative filtering, content-based recommendation algorithms, or deep learning recommendation models. The AI recommendation engine intelligently recommends based on user preferences, historical consumption records and preferences in the user preference and history database, real-time inventory information in the beverage / food database, and current environmental factors.

8. The augmented reality based bar interactive ordering management system as claimed in claim 1, wherein, The background management and service module includes: The order generation unit is configured to integrate the user ordering content into an ordering order. The inventory update unit is configured to send inventory update instructions to the beverage and food databases. The service distribution unit is configured to distribute the ordering order information to the waiter terminal. The payment processing unit is configured to receive the user's payment request and interface with the third-party payment interface to complete the payment.

9. The augmented reality based bar interactive ordering management system as claimed in claim 1, wherein, The AR display and feedback module includes: The virtual information generation unit is configured to generate three-dimensional virtual beverage models, personalized recommendation lists, or interactive menu virtual information based on the instructions of the environment-adaptive and user-interaction module and the recommendation results of the AI recommendation engine. The real-time superposition unit is configured to superimpose the generated virtual information into the real world in real time through the user front-end device.

10. The augmented reality based bar interactive ordering management system as claimed in claim 1, wherein, The user front-end device is an AR glasses, an AR tablet, or a smartphone equipped with high-performance cameras, microphones, and depth sensors.