MR navigation system based on multi-modal AI large model
Through the MR guide system of multimodal AI large model, the problem of inaccurate information acquisition in the existing technology is solved, and personalized guide and environmentally friendly and efficient guide experience is realized.
Patent Information
- Application Number
- CN202510448269.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing MR guide system is difficult to obtain the required information quickly and accurately, resulting in poor experience.
The MR guide system based on multimodal AI big model is adopted, including data acquisition, preprocessing, feature extraction, fusion, analysis and judgment, and AI processing units. Leopard, CogVLM, Emu3 and Ar i a are used for feature extraction and data processing to generate personalized guide words and chat content.
Improve data quality and consistency, generate personalized guides, reduce operating costs, and environmentally friendly eliminates the need for large amounts of energy and raw materials.
Smart Images

Figure CN120429607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tour guide systems, and in particular to an MR tour guide system based on a multimodal AI large model. Background Art
[0002] MR tours, or mixed reality tours, are a new type of guided tour that combines augmented reality (AR) and virtual reality (VR) technologies. They enhance the real-world experience by overlaying virtual images or information within the user's field of view. MR tours utilize mixed reality technology to overlay virtual content onto the real world, allowing users to see virtual objects within their real environment. This technology combines computer vision, sensors, GPS positioning, and other technologies to capture the user's position and movements in real time and adjust the virtual content accordingly.
[0003] The existing navigation methods make it difficult to quickly and accurately obtain the required information, resulting in a poor experience. To this end, we propose an MR navigation system based on a multimodal AI large model. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems mentioned in the above background technology. The present invention provides an MR navigation system based on a multimodal AI large model.
[0005] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:
[0006] An MR tour guide system based on a multimodal AI large model, comprising:
[0007] A data collection unit, wherein the data collection unit is used to collect data;
[0008] A data preprocessing unit, configured to preprocess the data collected by the data acquisition unit;
[0009] A feature extraction unit, wherein the feature extraction unit is used to extract features from the pre-processed data;
[0010] A feature fusion unit, which classifies and fuses the data after feature extraction;
[0011] A data analysis and judgment unit, which is used to analyze and judge the data after feature fusion is completed;
[0012] An AI processing unit, which processes the analyzed data to generate personalized guide words. At the same time, the generative AI generates corresponding chat content;
[0013] The result output unit is used to exit after the navigation is completed.
[0014] Furthermore, the data collected by the data acquisition unit includes: text feature data, image feature data, audio feature data and video feature data.
[0015] Furthermore, the feature extraction unit includes: Leopard, CogVLM, Emu3 and Ar ia, wherein Leopard is used to extract text features, CogVLM is used to extract image features, Emu3 is used to extract audio features, and Ar ia is used to extract video features.
[0016] Furthermore, the data preprocessing unit includes: deleting duplicate data and processing abnormal values.
[0017] Furthermore, the AI processing unit includes: data processing, feature extraction, model training and prediction, wherein the data preprocessing is used to clean the data; the feature extraction is used to extract text data, image data and audio data; the model training is used to minimize the prediction error or loss by optimizing the parameters of the model; the prediction is used to input new data into the trained model to obtain a prediction result.
[0018] Furthermore, the data processing includes the following steps:
[0019] A1) Data cleaning: remove noise, errors or irrelevant information from the data;
[0020] A2) Handling missing values: handle missing data by filling in, deleting, etc.
[0021] A3) Standardized data: standardize data of different dimensions;
[0022] A4) Convert data types: Ensure that the data format meets the requirements of the model;
[0023] A5) Data integration: merging data from different sources into a unified format;
[0024] A6) Data normalization: scaling the data to a specific range, usually between 0 and 1;
[0025] A7) Feature selection: Select the features that are most useful for model prediction and remove irrelevant or redundant features;
[0026] A8) Feature Engineering: Create new features or modify existing features.
[0027] Furthermore, the feature extraction is based on one of principal component analysis, linear discriminant analysis and automatic feature learning.
[0028] Furthermore, the model training includes the following steps:
[0029] B1) Data collection: collecting images, text, audio or video information;
[0030] B2) Data processing: Clean, format and label the collected data;
[0031] B3) Model training: The model starts learning from the data;
[0032] B4) Evaluation and Optimization: After training is completed, the model needs to be evaluated to determine its performance, and the model can be further optimized based on the evaluation results.
[0033] Furthermore, the prediction uses one of statistical methods, machine learning methods and deep learning.
[0034] Furthermore, the result output unit includes: a generation module and a storage module, wherein the generation module is used to generate a record report, and the storage module is used to effectively store the record report.
[0035] The beneficial effects of the present invention are as follows:
[0036] 1. The AI processing unit of the present invention can improve data quality and ensure data integrity and consistency. The AI processing unit is used to process the data that has completed analysis and judgment to generate personalized guide words. At the same time, the corresponding chat content is generated by generative AI, which does not require the consumption of large amounts of energy and raw materials, and does not require complex processes, greatly reducing operating costs and is also beneficial to environmental protection.
[0037] 2. The feature extraction unit of the present invention is used to extract features from the pre-processed data, including Leopard, CogVLM, Emu3 and Ar ia, which can meet the extraction of different features. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a workflow diagram of the present invention;
[0039] Figure 2 It is a workflow diagram of data processing in the present invention;
[0040] Figure 3 It is a workflow diagram of the model training in the present invention. DETAILED DESCRIPTION
[0041] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0042] See also Figure 1 - Figure 3 The present invention provides an MR navigation system based on a multimodal AI large model, comprising:
[0043] Data collection unit, the data collection unit is used to collect data;
[0044] A data preprocessing unit, which is used to preprocess the data collected by the data acquisition unit;
[0045] A feature extraction unit is used to extract features from the pre-processed data;
[0046] Feature fusion unit, which classifies and fuses the data after feature extraction;
[0047] The data analysis and judgment unit is used to analyze and judge the data after feature fusion is completed;
[0048] The AI processing unit is used to process the analyzed data and generate personalized guide words. At the same time, the generative AI generates the corresponding chat content.
[0049] Result output unit, which is used to exit after the tour is completed.
[0050] In this embodiment, preferably, the data collected by the data acquisition unit includes: text feature data, image feature data, audio feature data and video feature data; which is conducive to building a large multimodal model.
[0051] In this embodiment, preferably, the feature extraction units include: Leopard, CogVLM, Emu3 and Ar ia, among which Leopard is used to extract text features, and Leopard focuses on rich text image tasks; CogVLM is used to extract image features, and CogVLM combines the visual language basic model of deep fusion technology, which is particularly suitable for visual question answering and image subtitle generation; Emu3 is used to extract audio features, and Emu3 adopts an autoregressive technology route to uniformly process multiple modal information such as text, images and videos; Ar ia is used to extract video features, and Ar ia can process text, code, images and videos; it can meet the extraction of different features.
[0052] In this embodiment, preferably, the data preprocessing unit includes: deleting duplicate data and processing outliers; deleting duplicate data can remove duplicate records in the data set to avoid interference with the analysis results, and processing outliers to identify and process outliers to avoid their negative impact on the analysis results.
[0053] In this embodiment, preferably, the AI processing unit includes: data processing, feature extraction, model training and prediction, wherein data preprocessing is used to clean the data; feature extraction is used to extract text data, image data and audio data; model training is used to minimize prediction errors or losses by optimizing the parameters of the model; prediction is used to input new data into the trained model to obtain prediction results.
[0054] In this embodiment, preferably, data processing includes the following steps:
[0055] A1) Data cleaning: Remove noise, errors or irrelevant information from the data to improve data quality.
[0056] A2) Handling missing values: Handle missing data by filling in, deleting, and other methods to ensure data integrity and consistency.
[0057] A3) Standardized data: Standardize data of different dimensions to avoid certain features affecting model training due to excessively large numerical ranges.
[0058] A4) Convert data types: Ensure that the data format meets the model requirements, such as converting categorical variables into numerical types.
[0059] A5) Data integration: Merging data from different sources into a unified format for analysis.
[0060] A6) Data normalization: Scale the data to a specific range, usually between 0 and 1, to eliminate the dimensional differences between different features.
[0061] A7) Feature selection: Select the features that are most useful for model prediction and remove irrelevant or redundant features;
[0062] A8) Feature Engineering: Create new features or modify existing features to better suit the needs of the model.
[0063] In this embodiment, feature extraction is preferably based on one of principal component analysis, linear discriminant analysis, and automatic feature learning. Principal component analysis extracts features by identifying the main directions of data variation, projecting high-dimensional data into a low-dimensional space while preserving the main directions of data variation; linear discriminant analysis focuses on enhancing the separability between different categories of data, thereby improving classification performance; automatic feature learning utilizes deep neural networks and other methods to automatically learn feature representations of data without manual intervention.
[0064] In this embodiment, preferably, model training includes the following steps:
[0065] B1) Data collection: Collect images, text, audio or video information; able to meet the collection needs under different modalities.
[0066] B2) Data processing: Clean, format, and label the collected data to ensure data quality and consistency so that the model can learn useful information from it.
[0067] B3) Model training: The model begins to learn from the data; by continuously adjusting the parameters within the model, it can better fit the data and thus complete specific tasks.
[0068] B4) Evaluation and Optimization: After training is completed, the model needs to be evaluated to determine its performance, and the model can be further optimized based on the evaluation results.
[0069] In this embodiment, preferably, the prediction uses one of statistical methods, machine learning methods and deep learning.
[0070] In this embodiment, preferably, the result output unit includes: a generation module and a storage module, wherein the generation module is used to generate a record report, and the storage module is used to effectively store the record report.
[0071] The working principle and use process of the present invention:
[0072] The data collection unit is used to collect data. The data collected by the data collection unit includes text feature data, image feature data, audio feature data and video feature data, which is conducive to building a large multimodal model.
[0073] The data preprocessing unit is used to preprocess the data collected by the data acquisition unit; the data preprocessing unit includes: deleting duplicate data and processing outliers. Deleting duplicate data can remove duplicate records in the data set to avoid interference with the analysis results, and processing outliers to identify and process outliers to avoid their negative impact on the analysis results.
[0074] Feature extraction unit, which is used to extract features from preprocessed data; the feature extraction units include: Leopard, CogVLM, Emu3 and Ar ia, among which Leopard is used to extract text features and focuses on rich text image tasks; CogVLM is used to extract image features and combines the visual language basic model of deep fusion technology, which is particularly suitable for visual question answering and image subtitle generation; Emu3 is used to extract audio features and adopts autoregressive technology to uniformly process multiple modal information such as text, images and videos; Ar ia is used to extract video features and can process text, code, images and videos; it can meet the extraction of different features.
[0075] Feature fusion unit, which classifies and fuses the data after feature extraction;
[0076] The data analysis and judgment unit is used to analyze and judge the data after feature fusion is completed;
[0077] AI processing unit, the AI processing unit is used to process the data that has completed analysis and judgment, and generate personalized guide words. At the same time, the generative AI generates the corresponding chat content.
[0078] The AI processing unit includes: data processing, feature extraction, model training and prediction. Data preprocessing is used to clean up the data. Data processing includes the following steps:
[0079] A1) Data cleaning: Remove noise, errors or irrelevant information from the data to improve data quality.
[0080] A2) Handling missing values: Handle missing data by filling in, deleting, and other methods to ensure data integrity and consistency.
[0081] A3) Standardized data: Standardize data of different dimensions to avoid certain features affecting model training due to excessively large numerical ranges.
[0082] A4) Convert data types: Ensure that the data format meets the model requirements, such as converting categorical variables into numerical types.
[0083] A5) Data integration: Merging data from different sources into a unified format for analysis.
[0084] A6) Data normalization: Scale the data to a specific range, usually between 0 and 1, to eliminate the dimensional differences between different features.
[0085] A7) Feature selection: Select the features that are most useful for model prediction and remove irrelevant or redundant features;
[0086] A8) Feature Engineering: Create new features or modify existing features to better suit the needs of the model.
[0087] Feature extraction is used to extract text, image, and audio data. It relies on principal component analysis (PCA), linear discriminant analysis (LDA), and automatic feature learning. PCA extracts features by identifying the primary directions of data variation, projecting high-dimensional data into a low-dimensional space while preserving the primary direction of variation. LDA focuses on enhancing the separability between different categories of data, thereby improving classification performance. Automatic feature learning utilizes deep neural networks and other methods to automatically learn data feature representations without human intervention.
[0088] Model training is used to minimize prediction error or loss by optimizing the model parameters; model training includes the following steps:
[0089] B1) Data collection: Collect images, text, audio or video information; able to meet the collection needs under different modalities.
[0090] B2) Data processing: Clean, format, and label the collected data to ensure data quality and consistency so that the model can learn useful information from it.
[0091] B3) Model training: The model begins to learn from the data; by continuously adjusting the parameters within the model, it can better fit the data and thus complete specific tasks.
[0092] B4) Evaluation and Optimization: After training is completed, the model needs to be evaluated to determine its performance, and the model can be further optimized based on the evaluation results.
[0093] Prediction is used to input new data into a trained model to obtain prediction results. Prediction uses one of the following methods: statistical methods, machine learning methods, and deep learning.
[0094] The result output unit is used to exit after the navigation ends. The result output unit includes: a generation module and a storage module, wherein the generation module is used to generate a record report, and the storage module is used to effectively store the record report.
[0095] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An MR tour guide system based on a multimodal AI large model, characterized by: include: A data collection unit, wherein the data collection unit is used to collect data; A data preprocessing unit, configured to preprocess the data collected by the data acquisition unit; A feature extraction unit, wherein the feature extraction unit is used to extract features from the pre-processed data; A feature fusion unit, which classifies and fuses the data after feature extraction; A data analysis and judgment unit, which is used to analyze and judge the data after feature fusion is completed; An AI processing unit, which processes the analyzed data to generate personalized guide words. At the same time, the generative AI generates corresponding chat content; The result output unit is used to exit after the navigation is completed.
2. The MR navigation system based on a multimodal AI large model according to claim 1, characterized in that: The data collected by the data acquisition unit includes: text feature data, image feature data, audio feature data and video feature data.
3. The MR navigation system based on a multimodal AI large model according to claim 1, characterized in that: The feature extraction unit includes: Leopard, CogVLM, Emu3 and Aria, wherein Leopard is used to extract text features, CogVLM is used to extract image features, Emu3 is used to extract audio features, and Aria is used to extract video features.
4. The MR tour guide system based on a multimodal AI large model according to claim 1, characterized in that: The data pre-processing unit includes: deleting duplicate data and processing abnormal values.
5. The MR tour guide system based on a multimodal AI large model according to claim 1, characterized in that: The AI processing unit includes: data processing, feature extraction, model training and prediction, wherein the data preprocessing is used to clean the data; the feature extraction is used to extract text data, image data and audio data; the model training is used to minimize the prediction error or loss by optimizing the parameters of the model; The prediction is used to input new data into the trained model to obtain a prediction result.
6. The MR navigation system based on a multimodal AI large model according to claim 5, characterized in that: The data processing includes the following steps: A1) Data cleaning: remove noise, errors or irrelevant information from the data; A2) Handling missing values: handle missing data by filling in, deleting, etc. A3) Standardized data: standardize data of different dimensions; A4) Convert data types: Ensure that the data format meets the requirements of the model; A5) Data integration: merging data from different sources into a unified format; A6) Data normalization: scaling the data to a specific range, usually between 0 and 1; A7) Feature selection: Select the features that are most useful for model prediction and remove irrelevant or redundant features; A8) Feature Engineering: Create new features or modify existing features.
7. The MR tour guide system based on a multimodal AI large model according to claim 5, characterized in that: The feature extraction is based on one of principal component analysis, linear discriminant analysis and automatic feature learning.
8. The MR tour guide system based on a multimodal AI large model according to claim 5, characterized in that: The model training includes the following steps: B1) Data collection: collecting images, text, audio or video information; B2) Data processing: Clean, format and label the collected data; B3) Model training: The model starts learning from the data; B4) Evaluation and Optimization: After training is completed, the model needs to be evaluated to determine its performance, and the model can be further optimized based on the evaluation results.
9. The MR navigation system based on a multimodal AI large model according to claim 5, characterized in that: The prediction uses one of statistical methods, machine learning methods and deep learning.
10. The MR tour guide system based on a multimodal AI large model according to claim 1, characterized in that: The result output unit includes: a generation module and a storage module, wherein the generation module is used to generate a record report, and the storage module is used to effectively store the record report.