Intelligent wearable interaction system based on AI language model and AR technology
The intelligent wearable interaction system, which utilizes AI language models and AR technology, solves the problem of poor interactive experience of intelligent wearable devices in different environments, and achieves more efficient user experience optimization and enhanced adaptability to interactive scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-04-03
AI Technical Summary
Existing smart wearable devices lack effective context awareness during interaction, resulting in poor user experience in different environments and affecting the adaptability of human-computer interaction devices to different interactive situations.
The intelligent wearable interaction system, based on AI language models and AR technology, collects multimodal input data to understand intent, generates interactive intent information, and combines it with environmental context data to generate AR augmented data. The AR display module then performs visualization rendering and real-time adjustments to optimize the user experience.
It improves the adaptability of human-computer interaction devices to interactive scenarios, enhances the dynamism and adaptability of user interaction, and provides personalized and intelligent augmented reality experiences.
Smart Images

Figure CN121785455A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and in particular to an intelligent wearable interactive system based on AI language models and AR technology. Background Technology
[0002] Existing smart wearable devices generally suffer from insufficient context awareness during interaction, typically relying on a single interaction mode, such as voice recognition, touchscreen operation, or gesture control, but failing to effectively adjust the interaction method dynamically according to the user's environment. Understanding of the user's environment, behavioral state, and intent often remains at a superficial data level, failing to achieve deep fusion of multimodal information and semantic understanding, resulting in fragmented and interrupted cross-scenario interactive experiences. Due to the difficulty in dynamically perceiving and understanding the user's real-time context, their feedback mechanisms lack intelligence and predictability, failing to provide appropriate information presentation or operational support, further impacting the user experience of human-computer interaction. In summary, existing technologies suffer from a lack of effective context awareness in wearable interaction systems, leading to poor user interaction experiences in different environments, further affecting the context adaptability of human-computer interaction devices. Summary of the Invention
[0003] The purpose of this application is to provide an intelligent wearable interactive system based on AI language models and AR technology, in order to solve the technical problem in the prior art that the lack of effective context awareness in wearable interactive systems leads to poor user interaction experience in different environments, which further affects the context adaptability of human-computer interaction devices.
[0004] In view of the above problems, this application provides an intelligent wearable interaction system based on AI language models and AR technology. The intelligent wearable interaction system includes: an intent understanding module, used to collect multimodal input data from a user through an intelligent wearable device, and to perform intent understanding on the multimodal input data based on an AI language model to generate interaction intent information; a visualization rendering module, used to generate AR enhanced data based on the interaction intent information and environmental context data, and to perform visualization rendering on the AR enhanced data through the AR display module of the intelligent wearable device to generate user feedback data; and a data adjustment module, used to adjust the AR enhanced data in real time based on the user feedback data to complete the intelligent interaction process.
[0005] Optionally, the feature extraction unit is used to extract features from the multimodal input data to obtain a multimodal feature vector; the joint reasoning unit is used to input the multimodal feature vector into the AI language model for joint reasoning to generate intermediate representation information for intent understanding; and the context-aware correction unit is used to perform context-aware correction based on the intermediate representation information for intent understanding to generate the interaction intent information.
[0006] Optionally, the multi-source parsing subunit is used to perform multi-source parsing based on the multimodal input data to obtain speech stream data, image sequence data, and motion sensing data; the frame-segmentation and windowing processing subunit is used to perform frame-segmentation and windowing processing based on the speech stream data to generate speech feature vectors; the keyframe extraction subunit is used to extract keyframes from the image sequence data, determine multiple keyframes for target detection, and generate visual feature vectors; the filtering and denoising subunit is used to perform filtering and denoising based on the motion sensing data to generate motion feature vectors; and the dimensionality normalization subunit is used to normalize the dimensions of the speech feature vectors, the visual feature vectors, and the motion feature vectors to generate multimodal feature vectors.
[0007] Optionally, the cross-modal attention calculation subunit is used to perform cross-modal attention calculation on the speech feature vector, the visual feature vector, and the motion feature vector to generate a multimodal association mapping network; the hierarchical processing subunit is used to synchronize the multimodal feature vector to the AI language model for hierarchical processing according to the multimodal association mapping network: S1: Traverse the multimodal association mapping network to perform temporal analysis and capture temporal dependencies; S2: Record the interactive evolution of the multimodal feature vector according to the temporal dependencies to construct a dynamic evolution trend of interactive intent; S3: Perform hierarchical reasoning based on the dynamic evolution trend of interactive intent to generate preliminary intent representation information; the cross-modal weight allocation subunit is used to perform cross-modal weight allocation according to the preliminary intent representation information to determine the intermediate representation information of intent understanding.
[0008] Optionally, the semantic parsing unit is used to perform semantic parsing based on the interaction intent information to obtain multiple parsed data, wherein the multiple parsed data includes intent type data, target object data, and operation instructions; the scene geometric constraint unit is used to perform scene geometric constraints based on the environmental context data to determine a feasible placement area; the retrieval and matching unit is used to perform retrieval and matching according to the intent type data and the target object data to determine virtual object AR content data; and the instantiation and configuration unit is used to execute the operation instructions according to the feasible placement area to instantiate and configure the virtual object AR content data to generate the AR augmented data.
[0009] Optionally, the virtual object determination unit is used to receive and analyze the AR augmented data through the AR display module of the smart wearable device to determine the virtual object; the real-time light and shadow rendering unit is used to perform real-time light and shadow rendering on the virtual object to generate shadow data of the virtual object; the visual fusion unit is used to visually fuse the virtual object with the environmental data according to the shadow data of the virtual object to generate AR visual content; and the multi-dimensional acquisition unit is used to overlay the AR visual content onto the user's real field of vision for multi-dimensional acquisition to generate user feedback data.
[0010] Optionally, a channel construction subunit is used to perform geometric transformations on the virtual object according to the shadow data of the virtual object, and establish a geometric processing channel, wherein the geometric processing channel includes a lighting calculation subchannel, a visual processing subchannel, and a visual fusion subchannel; a multi-source lighting calculation subunit is used to perform multi-source lighting calculations on the virtual object through the lighting calculation subchannel to generate first visual content; a color grading subunit is used to perform color grading on the virtual object by combining the visual processing subchannel with environmental data to generate second visual content; and a depth test fusion subunit is used to perform depth test fusion of the first visual content and the second visual content through the visual fusion subchannel to construct the AR visual content.
[0011] Optionally, the user data acquisition subunit is used to project the AR visual content onto the user's retina through an optical see-through display module to collect user physiological data; the motion information capture subunit is used to capture user spatial movement information based on the AR visual content using a depth camera to obtain user behavior interaction data; the preference analysis subunit is used to perform user preference analysis based on the AR visual content to obtain user subjective feedback data; and the structured integration subunit is used to structure and integrate the user physiological data, the user behavior interaction data, and the user subjective feedback data to generate user feedback data.
[0012] Optionally, the eye-tracking unit is used to track the user's eyes based on the user's physiological data to determine the user's gaze point information; the gesture recognition unit is used to recognize gestures based on the user's behavioral interaction data to determine the user's gesture interaction information; the preference analysis unit is used to perform preference analysis based on the user's subjective feedback data to determine the user's personalized association information; the dynamic adjustment unit is used to dynamically adjust the AR augmented data according to the user's personalized association information combined with the user's gaze point information and the user's gesture interaction information to formulate an interaction strategy; the continuous monitoring unit is used to continuously monitor the interaction strategy by executing a real-time feedback loop to obtain user experience indicators; and the strategy optimization unit is used to optimize the interaction strategy by performing reinforcement learning on the intelligent interaction between the AI language model and AR technology according to the user experience indicators to construct an interaction optimization strategy.
[0013] Optionally, the reward signal determination subunit is used to use the user experience index as the reward signal; the state space determination subunit is used to define the state space according to the user gesture interaction information; the action space determination subunit is used to define the action space according to the user gaze point information; and the interaction reinforcement subunit is used to perform reinforcement learning on the intelligent interaction between the AI language model and AR technology according to the reward signal, the state space, and the action space, and construct the interaction optimization strategy.
[0014] The technical solution provided in this application has at least the following technical effects or advantages: An intent understanding module is used to collect multimodal input data from users through a smart wearable device, and to perform intent understanding on the multimodal input data based on an AI language model to generate interactive intent information; a visualization rendering module is used to generate AR augmented data based on the interactive intent information and environmental context data, and to perform visualization rendering on the AR augmented data through the AR display module of the smart wearable device to generate user feedback data; a data adjustment module is used to adjust the AR augmented data in real time based on the user feedback data to complete the intelligent interaction process. In other words, by understanding multimodal input data through an AI language model, generating AR augmented data by combining it with environmental context data, and performing visualization rendering through the AR display module of the smart wearable device, and adjusting the AR augmented data in real time based on user feedback data, the user experience is optimized, the adaptability of the human-computer interaction device to the interactive context is improved, and the dynamism and adaptability of user interaction are enhanced.
[0015] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the structure of an intelligent wearable interactive system based on AI language model and AR technology according to this application.
[0018] Figure 2 This is a schematic diagram of the visualization rendering module in an intelligent wearable interactive system based on AI language models and AR technology, as described in this application.
[0019] Figure labeling: Intent understanding module 11, Visualization rendering module 12, Data adjustment module 13, Semantic parsing unit 21, Scene geometric constraint unit 22, Retrieval and matching unit 23, Instantiation configuration unit 24. Detailed Implementation
[0020] This application provides an intelligent wearable interaction system based on AI language models and AR technology. It addresses the technical problem in existing technologies where the lack of effective context awareness in wearable interaction systems leads to poor user experience in different environments, further impacting the context adaptability of intelligent wearable devices. By using an AI language model to understand multimodal input data and combining it with environmental context data to generate AR-enhanced data, and then visualizing and rendering this data through the AR display module of the intelligent wearable device, the system optimizes the user experience by adjusting the AR-enhanced data in real time based on user feedback. This improves the context adaptability of the human-computer interaction device and enhances the dynamism and adaptability of user interaction.
[0021] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0022] For examples, please refer to the appendix. Figure 1 This application provides an intelligent wearable interactive system based on AI language models and AR technology, wherein the intelligent wearable interactive system based on AI language models and AR technology includes: The intent understanding module 11 is used to collect multimodal input data from users through smart wearable devices, perform intent understanding on the multimodal input data based on an AI language model, and generate interactive intent information.
[0023] Furthermore, the intent understanding module 11 in the intelligent wearable interaction system based on AI language model and AR technology is also used for: a feature extraction unit, used for extracting features from the multimodal input data to obtain a multimodal feature vector; a joint reasoning unit, used for inputting the multimodal feature vector into the AI language model for joint reasoning to generate intermediate representation information for intent understanding; and a context-aware correction unit, used for performing context-aware correction based on the intermediate representation information for intent understanding to generate the interaction intent information.
[0024] Furthermore, the intent understanding module 11 in the intelligent wearable interaction system based on AI language models and AR technology is also configured to: a multi-source parsing subunit, used to perform multi-source parsing based on the multimodal input data to obtain speech stream data, image sequence data, and motion sensing data; a frame-segmentation and windowing processing subunit, used to perform frame-segmentation and windowing processing based on the speech stream data to generate speech feature vectors; a keyframe extraction subunit, used to extract keyframes from the image sequence data, determine multiple keyframes for target detection, and generate visual feature vectors; a filtering and denoising subunit, used to perform filtering and denoising based on the motion sensing data to generate motion feature vectors; and a dimension normalization subunit, used to normalize the dimensions of the speech feature vectors, the visual feature vectors, and the motion feature vectors to generate multimodal feature vectors.
[0025] Furthermore, the intent understanding module 11 in the intelligent wearable interaction system based on AI language models and AR technology is also used for: a cross-modal attention calculation subunit, used for performing cross-modal attention calculation on the speech feature vector, the visual feature vector, and the motion feature vector to generate a multimodal association mapping network; a hierarchical processing subunit, used for synchronizing the multimodal feature vectors to the AI language model for hierarchical processing according to the multimodal association mapping network: S1: traversing the multimodal association mapping network to perform temporal analysis and capture temporal dependencies; S2: recording the interactive evolution of the multimodal feature vectors according to the temporal dependencies to construct a dynamic evolution trend of interactive intent; S3: performing hierarchical reasoning based on the dynamic evolution trend of interactive intent to generate preliminary intent representation information; and a cross-modal weight allocation subunit, used for performing cross-modal weight allocation according to the preliminary intent representation information to determine the intermediate representation information of intent understanding.
[0026] Specifically, smart wearable devices collect multimodal input data from users, meaning input data obtained from multiple sensing channels, including various information sources, which provides a more comprehensive understanding of the user's state and environment. For example, voice data is collected by microphones, visual data is acquired through cameras, and motion data is provided by accelerometers or gyroscopes.
[0027] Multi-source parsing is performed on multimodal input data, which involves decomposing the acquired raw multimodal data into different data streams, such as speech stream data, image sequence data, and motion sensing data. Speech stream data is audio signal acquired through a microphone, typically time-series data containing speech information; image sequence data is a series of continuous image data acquired through a camera, including the scene around the user; motion sensing data is data related to user movement acquired by sensors such as accelerometers and gyroscopes, such as acceleration, rotation angle, and angular velocity.
[0028] Framing and windowing of speech stream data first requires framing, which involves dividing the speech stream data into short frames. The length of each frame depends on the required spectral resolution and application scenario. Framing aims to make the speech signal within each frame approximate a stationary signal, facilitating frequency domain analysis. The length of a speech signal frame can be expressed in several ways. If expressed in time, a frame typically ranges from 15ms to 30ms, with an empirical value of 25ms. A 25ms frame refers to a speech signal with a duration of 25 milliseconds. It can also be expressed in terms of the number of sampling points. If a signal has a sampling rate of 16kHz, a frame consists of 16kHz × 25ms = 400 sampling points. Frame shift is the distance moved during each frame division. Starting from the beginning of the first frame, a frame shift occurs before the next frame begins. Frame shift can also be expressed in two ways: in time (often set to 10ms) or in sampling points (a 16kHz sampling rate signal typically has a frame shift of 160 sampling points).
[0029] Windowing is applied to each frame of the signal to smooth the edges and reduce spectral leakage caused by signal truncation. After framing, there are discontinuities at the beginning and end of each frame; therefore, the more frames there are, the greater the error compared to the original signal. Windowing addresses this issue, making the framed signal continuous and ensuring each frame exhibits periodic function characteristics. Windowing involves applying a window function, such as a Hamming or Hanning window, to each frame to reduce spectral leakage and the impact of signal edge effects. Multiplying each frame of data by a window function causes the frame data to gradually decay to zero at both ends, resulting in better local stability and reduced abrupt changes between frames. The total number of frames that can be divided into for each signal is calculated based on the signal length, frame shift, and frame length. The formula for calculating the number of frames is: Number of frames = (Signal length - Frame length) ÷ Frame shift + 1. For example, suppose there is a speech signal with a length of 1000ms. Choosing 20ms as the frame length and 10ms as the frame shift, according to the formula, there are 98 frames, each 20ms long. A Hamming window is applied to each frame, and a Fourier transform is performed to obtain the spectral data. The spectral data of each frame becomes a speech feature vector, resulting in 98 feature vectors. The windowing operation is relatively simple; each frame of the signal is multiplied sequentially by a window function, which can be directly called from the corresponding tool library. During the framing operation, there may be situations where the remaining signal length is less than one frame. In this case, zero-padding is performed on this segment to make it a full frame, or it can be discarded, as the last frame is at the very end of the sentence and mostly consists of silent segments. A short-time Fourier transform converts each frame of the signal from the time domain to the frequency domain, obtaining the spectral representation of that frame and thus the speech feature vectors. These vectors reflect the spectral characteristics of the speech.
[0030] For example, assuming the audio signal has a sampling rate of 16kHz, then 16,000 data points are sampled per second. If each frame is chosen to be 20ms (320 samples), and the overlap between frames is 50% (160 samples overlap), then one second of the signal will be divided into 80 frames. The first frame of the signal might be from sample 0 to 319, the second frame from sample 160 to 479, and so on. A Hamming window is applied to each frame. Assuming the original signal of the first frame is x1, x2, ..., x... 320 The window functions are w1, w2, ..., w 320 Then the windowed signal is w1x1, w2x2, ..., w 320 x 320 A short-time Fourier transform is performed on each frame of the windowed signal to obtain the spectral features of each frame. Ultimately, the feature vector of each frame represents the speech information of that frame.
[0031] Keyframe extraction from image sequence data involves selecting frames that represent significant changes or distinctive features within a video sequence by calculating the differences between the images. In video or image sequences, keyframes are those image frames that represent important features of the video content or scene, containing key changes or information within the sequence, making them crucial for video processing. For example, if the difference between two adjacent frames is significant—that is, the image content has changed substantially—the current frame can be considered a keyframe. For instance, suppose there are 10 images in a sequence. Images 1 to 5 show almost no change in the scene, while images 6 to 10 show significant changes, such as people walking or objects moving. In this case, images 6, 7, 8, 9, and 10 are selected as keyframes, while images 1 to 5 are ignored.
[0032] Object detection is performed on the extracted keyframes. The purpose of object detection is to identify objects in an image and determine their locations, typically using bounding boxes, such as those used by YOLO. After object detection, each keyframe identifies different objects in the image, such as people, vehicles, and furniture, and marks their locations. The detection results usually include the object category and bounding box coordinates. For example, the bounding box for a person is [100, 50, 200, 150], and the bounding box for a car is [300, 100, 400, 200]. Feature extraction is then performed on the keyframes after object detection to generate visual feature vectors. Typically, feature extraction is done using convolutional neural networks, which learn and extract local and global features from the image. The visual feature vectors will contain key information about objects, scenes, and textures in the image.
[0033] Motion sensing data is filtered and denoised using a low-pass filter to remove high-frequency noise and unnecessary interference signals. Meaningful motion features, such as acceleration, step frequency, and direction, are extracted from the processed motion data to generate a motion feature vector. For example, if the acceleration data recorded by the accelerometer contains noise, a low-pass filter is used to remove the noise, generating a smooth motion feature vector, such as a step frequency of 120 steps / minute, an acceleration of 2.5 m / s², and a direction of southeast.
[0034] Because data from different modalities have different dimensions and numerical ranges, they need to be normalized to ensure they can be compared and fused at the same scale. After normalization, feature vectors from multiple modalities can be merged into a unified multimodal feature vector, representing comprehensive user information. By fusing feature vectors from multiple modalities such as speech, vision, and motion, a multimodal feature vector is obtained to describe the user's multidimensional state and needs. By extracting representative features from multimodal data and fusing them into a unified multimodal feature vector, a comprehensive understanding of user needs and environmental changes can be achieved.
[0035] Cross-modal attention is calculated based on feature vectors from different modalities: speech, vision, and motion. This mechanism automatically learns and adjusts the relative importance of different modalities, integrating key information from speech, vision, and motion data. A multimodal association mapping network is generated to represent the relationships and dependencies between modalities. Features from different modalities are effectively correlated through this network, ensuring accurate transmission of multimodal information. Cross-modal attention mechanisms are typically implemented using deep learning models, dynamically calculating the attention weight for each modality based on the relationships and interactions between features. Specifically, a weighted sum is calculated in the feature vector of each modality to provide the most relevant or important information. For example, suppose a user issues a voice command to open the curtains while interacting with a smart home device. The speech feature vector contains the speech spectrum information of the command, the visual feature vector contains the position image of the curtains, and the motion feature vector shows that the user is moving towards the curtains. Cross-modal attention calculation assigns a higher weight to the speech command, ensuring the importance of speech information in the fusion process, while referencing visual and motion data to execute the command more accurately.
[0036] After cross-modal attention computation is completed, a multimodal association mapping network is constructed by combining the feature information of each modality. This network learns the mapping relationships between modalities by associating feature vectors from different modalities. For example, voice commands may be associated with object locations and user actions. The role of the multimodal association mapping network is to capture the interdependencies between different modalities and take them into account when fusing features. It dynamically adjusts the influence of each modality in the final decision based on contextual information, thereby enhancing the level of intelligence.
[0037] The generated multimodal association mapping network synchronizes all multimodal feature vectors to the AI language model. The AI language model then progressively extracts and understands the user's intent through hierarchical processing of these features. Hierarchical processing means that the AI language model processes and infers the input information sequentially, from low-level feature extraction to high-level semantic understanding, ultimately forming a comprehensive understanding of the user's intent. Temporal analysis is performed by traversing the multimodal association mapping network to capture the dependencies between modalities over time. For example, the user's voice commands, visual input, and motion data may occur sequentially and be correlated.
[0038] By recording the interactive evolution of multimodal feature vectors based on temporal dependencies, a dynamic evolution trend of interaction intent is constructed, reflecting the changes in user intent over time, such as from initial needs to specific operations. The dynamic evolution trend of interaction intent is the changing trend of user intent over time during the interaction process; that is, the change in user needs at different points in time. For example, a user's intent may evolve from issuing a command to adjusting device settings; that is, after issuing a voice command to turn on the light, the user may issue an instruction to dim the light—this is an evolution trend.
[0039] Based on the dynamic evolution of interaction intent, the AI language model will perform hierarchical reasoning on user intent, gradually refining and interpreting it to generate an accurate preliminary intent representation. Cross-modal weighting will then be applied based on this preliminary intent representation, assigning appropriate weights to each modality according to its importance in the interaction, thereby determining the contribution of each modality's information to intent understanding.
[0040] After generating intermediate representation information for intent understanding, the system is corrected based on the current context, taking into account factors such as time, location, and the user's dynamic state to ensure that the smart wearable device's understanding is more consistent with reality. After context-aware correction, interactive intent information is generated, which represents the final understanding of the user's needs and drives subsequent decisions or actions. For example, if a user issues a verbal command to open the curtains, context-aware correction confirms that the user is currently in the bedroom and that their intent is to open the bedroom curtains, ignoring other possible misunderstandings, such as opening the living room curtains. Through the fusion and joint reasoning of multimodal data, the system comprehensively and accurately understands the user's intent, reducing comprehension errors.
[0041] The visualization rendering module 12 is used to generate AR enhanced data based on the interaction intent information and environmental context data, and to perform visualization rendering of the AR enhanced data through the AR display module of the smart wearable device to generate user feedback data.
[0042] Further details are attached. Figure 2 As shown, the visualization rendering module 12 in the intelligent wearable interactive system based on AI language models and AR technology is further configured as follows: a semantic parsing unit 21, configured to perform semantic parsing based on the interaction intent information to obtain multiple parsing data, wherein the multiple parsing data includes intent type data, target object data, and operation instructions; a scene geometric constraint unit 22, configured to perform scene geometric constraints based on the environmental context data to determine a feasible placement area; a retrieval and matching unit 23, configured to perform retrieval and matching according to the intent type data and the target object data to determine virtual object AR content data; and an instantiation and configuration unit 24, configured to execute the operation instructions according to the feasible placement area to instantiate and configure the virtual object AR content data to generate the AR augmented data.
[0043] Specifically, semantic parsing of interaction intent information involves transforming the input intent information into understandable structured data and extracting explicit task components from it, such as intent type data, target object data, and operation instructions. Intent type data categorizes user intents to identify the user's basic needs; target object data provides specific information about the user's target, which may be an object, device, or location; and operation instructions are the user's specific actions on the target object, such as turning it on, off, or adjusting brightness.
[0044] Based on environmental context data, the spatial and geometric constraints in the virtual scene are analyzed to determine which areas can be used to place virtual objects. By analyzing environmental context data, such as spatial dimensions and obstacle positions, feasible placement areas are identified—areas where virtual objects can be successfully placed. Based on intent type data and target object data, a search and matching process is performed to find the corresponding virtual object AR content data, i.e., augmented reality content data related to the virtual object. This typically includes information such as the virtual object's model, texture, and animation, and will be presented on the user's device to create a virtual interactive experience. Based on the previously determined feasible placement areas, the virtual object AR content data is instantiated and configured according to operation instructions. The virtual object is placed in the scene with the correct size, orientation, and position, and corresponding interactions are executed based on user actions. The instantiated and configured virtual object will be displayed on the user's device screen as AR augmented data, blending with the real environment to create an augmented reality interactive experience.
[0045] By extracting explicit target objects and operational instructions from user input and making intelligent decisions based on environmental context and geometric constraints, the system ensures that virtual objects are correctly rendered in the user's AR environment. Through dynamically instantiating virtual objects and generating AR augmented data, a highly intelligent and personalized augmented reality experience is provided.
[0046] Furthermore, the visualization rendering module 12 in the intelligent wearable interactive system based on AI language models and AR technology is also used for: a virtual object determination unit, used to receive and analyze the AR enhanced data through the AR display module of the intelligent wearable device to determine the virtual object; a real-time light and shadow rendering unit, used to perform real-time light and shadow rendering on the virtual object to generate shadow data of the virtual object; a visual fusion unit, used to visually fuse the virtual object with environmental data according to the shadow data of the virtual object to generate AR visual content; and a multi-dimensional acquisition unit, used to overlay the AR visual content onto the user's real field of vision for multi-dimensional acquisition to generate user feedback data.
[0047] Furthermore, the visualization rendering module 12 in the intelligent wearable interactive system based on AI language models and AR technology is also used for: a channel construction subunit, used to perform geometric transformation on the virtual object according to the shadow data of the virtual object, and establish a geometric processing channel, the geometric processing channel including a lighting calculation subchannel, a visual processing subchannel, and a visual fusion subchannel; a multi-source lighting calculation subunit, used to perform multi-source lighting calculation on the virtual object through the lighting calculation subchannel to generate first visual content; a color grading subunit, used to perform color grading on the virtual object in combination with environmental data through the visual processing subchannel to generate second visual content; and a depth test fusion subunit, used to perform depth test fusion on the first visual content and the second visual content through the visual fusion subchannel to construct the AR visual content.
[0048] Furthermore, the visualization rendering module 12 in the intelligent wearable interactive system based on AI language models and AR technology is also used for: a user data acquisition subunit, used to project the AR visual content onto the user's retina through an optical see-through display module to collect user physiological data; a motion information capture subunit, used to capture user spatial motion information based on the AR visual content through a depth camera to obtain user behavior interaction data; a preference analysis subunit, used to perform user preference analysis based on the AR visual content to obtain user subjective feedback data; and a structured integration subunit, used to structure and integrate the user physiological data, the user behavior interaction data, and the user subjective feedback data to generate user feedback data.
[0049] Specifically, the AR display module of a smart wearable device receives and analyzes AR augmented reality data, including information about virtual objects such as 3D models, textures, materials, and lighting effects, to determine the virtual objects to be displayed. The AR display module is a hardware component in a smart wearable device used to display augmented reality content, responsible for overlaying virtual content onto the real world. It is typically presented through a transparent display screen, projection system, or head-mounted display device. Virtual objects are objects generated through computer graphics, such as 3D models, icons, images, or other digital elements. Virtual objects do not depend on physical existence in the real world, but they are embedded into the real world through AR technology, creating an augmented reality experience.
[0050] After defining the virtual object, real-time lighting and shadow rendering is performed on it. This process calculates the impact of light sources on the virtual object, simulating lighting and shadow effects in the real world. Real-time lighting and shadow rendering dynamically calculates and renders the relationship between the virtual object and light sources in the environment as the virtual object interacts with the real-world scene, generating lighting effects and shadows. This not only increases the realism of the virtual object but also reflects the scene's lighting conditions through changes in shadows and lighting. The shadow data of the virtual object includes the shape, size, position, transparency, and blur of the shadow. The rendering process determines the shadow projection based on the relative position of the actual light source and the virtual object. Through real-time lighting and shadow rendering, the shadows of the virtual object are not merely static but dynamically updated according to the movement of the virtual object and changes in the light source.
[0051] Geometric transformations are performed on the shadow data of virtual objects to adjust their position, orientation, size, and other attributes, making them match the objects and light sources in the real-world scene, thus establishing a geometric processing channel. This channel is a processing procedure composed of multiple sub-channels used to handle various visual effects of virtual objects in the AR scene, including a lighting calculation sub-channel, a visual processing sub-channel, and a visual fusion sub-channel. The lighting calculation sub-channel is responsible for calculating the lighting effects of virtual objects in multi-light source environments; the visual processing sub-channel is responsible for color processing and adjustment of virtual objects to match their color with the ambient lighting and atmosphere; and the visual fusion sub-channel is responsible for fusing the results of lighting calculations and color grading to ensure seamless integration of the virtual object with the real world in the final visual presentation.
[0052] In the lighting calculation sub-channel, the lighting effects of a virtual object are calculated by simulating the influence of multiple light sources. Light sources include sunlight and indoor lighting, and lighting effects include highlights, shadows, and reflections. For example, the intensity of natural light is 500 lux, top light is 300 lux, and flashlight light is 100 lux. Through multi-source lighting calculation, the first visual content is obtained, which is the visual effect of the virtual object under different lighting conditions. For example, if the virtual object is a reflective metal surface, the lighting calculation sub-channel will consider the different angles of sunlight and streetlights to calculate the reflection effect of the virtual metal surface, generating the first visual content. The first visual content is the image or visual effect generated by the lighting calculation sub-channel, which mainly describes the appearance of the virtual object under different light sources, including lighting effects such as brightness, shadows, and reflections.
[0053] In the visual processing sub-channel, the virtual object's color is graded based on the environment's color temperature, brightness, and other factors in the scene. This ensures the virtual object's color visually harmonizes with the environment, avoiding visual conflicts between the virtual object and the background. Through color grading, a second visual content is obtained, adjusting the virtual object's color and brightness to make it appear more natural under different lighting conditions. Color grading adjusts the virtual object's hue, saturation, brightness, and other visual attributes according to different lighting conditions to match the virtual object with the real-world scene in different environments, making it look more natural. The second visual content is an image or visual effect generated by the visual processing sub-channel, typically including adjustments to the virtual object's color, brightness, etc., to make the virtual object's visual effect more consistent with the current environment.
[0054] The visual fusion subchannel fuses the outputs from the lighting calculation subchannel and the visual processing subchannel. This involves fusing the first and second visual content through depth testing to ensure the accuracy of the virtual object's depth, preventing occlusion errors or penetration between the virtual object and objects in the real environment. The resulting AR visual content is a complete visual effect that integrates the lighting, color adjustment, and depth information of the virtual object, presented to the user.
[0055] AR visual content is projected onto the user's retina using an optical see-through display module. Lenses, micro-projectors, and other devices directly present virtual images to the user, allowing them to see virtual content integrated with the real world. Simultaneously, physiological sensors monitor the user's physiological responses in real time, capturing physiological reactions such as eye movements, pupil changes, and blink frequency. This allows for understanding the user's attention to the AR content, emotional fluctuations, and reaction intensity. The optical see-through display module is a display technology that projects virtual images directly onto the user's retina using special optical lenses or projection technology, allowing the user to see augmented reality content without a traditional screen or display. It is commonly used in AR glasses, helmets, or other wearable devices. The user's physiological response data is collected by physiological sensors such as eye trackers, heart rate sensors, and brainwave monitoring devices. This data includes eye movements, pupil changes, and heart rate, reflecting the user's emotional state, attention, anxiety level, and other characteristics.
[0056] Depth cameras are used to capture user position changes and movements in space, analyze depth information in the scene, and identify user actions, position changes, postures, etc., accurately recording user behavioral interaction data, including head movements, gestures, gait, body postures, etc. User behavioral interaction data refers to various behavioral data generated by users during interaction with the AR system, including but not limited to eye tracking, gestures, body postures, movements, and gaze duration.
[0057] By analyzing user reactions to AR visual content, and based on their physiological responses and behavioral interaction data, user preferences and emotional states can be inferred. Subjective user feedback data is actively collected through methods such as questionnaires, rating systems, and facial expression analysis to obtain user satisfaction and feedback. This subjective feedback data, provided by users themselves, typically collected through questionnaires, rating systems, facial expression analysis, or voice feedback, reflects users' personal feelings about the AR experience, such as satisfaction, emotional state, and comfort.
[0058] Data from various sources—user physiological response data, user behavioral interaction data, and user subjective feedback data—is structured and integrated to be stored and processed in a unified format, resulting in user feedback data. This processed, structured, and integrated user feedback data, including physiological response data, behavioral interaction data, and subjective feedback data, reflects user experience and preferences. By monitoring user physiological responses, behavioral interactions, and subjective feedback in real time, AR content is adjusted accordingly to enhance user immersion and interactive experience.
[0059] The data adjustment module 13 is used to adjust the AR enhanced data in real time based on the user feedback data to complete the intelligent interaction process.
[0060] Furthermore, the data adjustment module 13 in the intelligent wearable interaction system based on AI language models and AR technology is also used for: an eye-tracking unit, used for eye tracking based on the user's physiological response data to determine the user's gaze point information; a gesture recognition unit, used for gesture recognition based on the user's behavioral interaction data to determine the user's gesture interaction information; a preference analysis unit, used for preference analysis based on the user's subjective feedback data to determine the user's personalized association information; a dynamic adjustment unit, used for dynamically adjusting the AR enhancement data according to the user's personalized association information combined with the user's gaze point information and the user's gesture interaction information to formulate an interaction strategy; a continuous monitoring unit, used for continuously monitoring by executing the interaction strategy in real-time feedback loop to obtain user experience indicators; and a strategy optimization unit, used for strengthening the interaction strategy by performing reinforcement learning on the intelligent interaction of the AI language model and AR technology according to the user experience indicators to construct an interaction optimization strategy.
[0061] Furthermore, the data adjustment module 13 in the intelligent wearable interaction system based on AI language model and AR technology is also used for: a reward signal determination subunit, used to use the user experience index as a reward signal; a state space determination subunit, used to define a state space according to the user gesture interaction information; an action space determination subunit, used to define an action space according to the user gaze point information; and an interaction reinforcement subunit, used to perform reinforcement learning on the intelligent interaction of the AI language model and AR technology according to the reward signal, the state space, and the action space, and construct the interaction optimization strategy.
[0062] Specifically, eye tracking is performed based on user physiological response data. This involves extracting data detected by the eye tracker from the user's physiological response data, including eye movement trajectory, pupil changes, heart rate, etc., recording the user's eye movement trajectory, and determining the specific location of the user's gaze at various moments. For example, if a user is looking at a virtual button on the screen, it is determined that the user is interacting with the virtual button, and this location is captured as the gaze point information. The average time a user gazes at a virtual object is 5 seconds, and during this gaze, the user's pupils show a significant dilation change, with the pupil diameter increasing from 3mm to 4mm, indicating that the user is interested in the virtual object.
[0063] Gesture recognition is performed based on user behavior interaction data. This involves extracting the user's hand movements and gestures from the data to identify their intentions or commands, such as pointing, grasping, or waving. For example, a user might use waving gestures to control the rotation or movement of a virtual object. These gestures are captured and analyzed by a camera to determine the user's interaction intent. If a user performs a drag gesture, moving an object from position A (10cm from hand) to position B (20cm from hand) in the virtual environment, the spatial positional change during this process is captured and recorded by a depth camera.
[0064] Preference analysis is conducted based on user subjective feedback data to determine users' preferences for specific content or interaction methods. In other words, by analyzing user behavioral data and subjective feedback, we can infer users' interests, needs, and emotional states. Personalized user-related information is key information about users' personalized needs derived from preference analysis, including features, content, and interaction methods that users prefer. For example, after completing a feedback questionnaire, a user might give a 9 / 10 satisfaction rating and indicate that the interaction with virtual objects was very smooth and the images were clearly displayed.
[0065] Based on user-personalized association information, user gaze point information, and user gesture interaction information, AR augmented data is dynamically adjusted, changing the displayed content or response method according to the user's current needs, focus, and interaction behavior. The interaction strategy is a response strategy formulated by the smart wearable device based on user data when interacting with the user, determining how to present content or adjust behavior, such as adjusting the position, size, or display hierarchy of virtual objects, or changing the display method of virtual objects to meet the user's current needs. For example, if it detects that the user frequently gazes at a virtual object and adjusts or drags it with gestures, it is inferred that the user is interested in the object, and therefore the visual effects or interactive functions of the object are enhanced.
[0066] The interaction strategy is executed and monitored through a real-time feedback loop to obtain user reactions and interaction effects, including user behavior data, emotional responses, and operational efficiency, forming user experience metrics. User experience metrics are used to measure the quality of the user experience during the interaction process, typically including user satisfaction, operational smoothness, interaction efficiency, and comfort. The effectiveness of the interaction strategy is evaluated through user experience metrics, and adjustments are made based on feedback. For example, if user operations become smoother or satisfaction increases, the strategy is effective; conversely, the interaction method needs to be optimized.
[0067] User experience metrics are used as reward signals to evaluate the quality of a behavior. Interactions are scored using these metrics, and the results are fed back to the learning model to adjust future behaviors. For example, a high reward signal is given if a user completes a task quickly and expresses satisfaction, such as a score of 8 / 10; conversely, a low reward signal is given if the user is dissatisfied. A state space is defined based on user gesture interaction information. Each state describes the user's current interaction, such as their gesture, operation method, and interaction mode. An action space is defined by analyzing user gaze information. Actions in each action space correspond to the user's gaze position or gaze behavior, such as adjusting the display of a virtual object, changing the size of an object, or rotating an object.
[0068] Based on user reward signals, state space, and action space, the interaction strategy is continuously adjusted through reinforcement learning. Different actions are tried, and their effects, such as task completion time and user satisfaction, are evaluated. Future decisions are adjusted based on reward signals. Through repeated learning and optimization, the optimal interaction optimization strategy is gradually built to improve user experience. For example, suppose in a user test, the following data was collected based on user feedback: user satisfaction rating for the interaction is 8.5 / 10, the average time for users to complete the virtual object rotation task is 12 seconds, and the number of errors made by users during the operation is 2. Reward signals are obtained from these user experience indicators. Higher satisfaction ratings and shorter task completion times generate positive rewards, while interactions with higher error frequencies generate negative rewards. Because of the high satisfaction rating (8.5 / 10), a positive reward of +1 is given; the task completion time is 12 seconds, and if the goal is to reduce the task completion time to less than 15 seconds, a positive reward of +0.5 is given; because the user's error frequency is 2 times, a negative reward of -0.4 is given based on the number of errors made during the interaction. Based on these data, the total reward signal is calculated to be +1.1, indicating that the interaction process performed well overall, but there is still room for optimization, such as reducing the error frequency.
[0069] By analyzing users' personalized association information, gaze points, and gesture interaction information, the system can customize personalized interaction strategies for each user, enhancing user engagement and satisfaction. Adjustments to AR-enhanced data based on user behavior and feedback allow the interaction process to flexibly adapt to the needs and emotional states of different users, improving the naturalness and fluency of the interaction.
[0070] In summary, the intelligent wearable interaction system based on AI language models and AR technology provided in this application has the following technical effects: An intent understanding module collects multimodal input data from the user through the intelligent wearable device, performs intent understanding on the multimodal input data based on the AI language model, and generates interaction intent information; a visualization rendering module generates AR enhanced data based on the interaction intent information and environmental context data, and performs visualization rendering on the AR enhanced data through the AR display module of the intelligent wearable device to generate user feedback data; a data adjustment module adjusts the AR enhanced data in real time based on the user feedback data to complete the intelligent interaction process. In other words, by understanding multimodal input data through an AI language model, generating AR enhanced data by combining it with environmental context data, and performing visualization rendering through the AR display module of the intelligent wearable device, the AR enhanced data is adjusted in real time based on user feedback data to optimize the user experience, improve the adaptability of the human-computer interaction device to the interaction context, and enhance the dynamism and adaptability of user interaction.
[0071] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0072] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A smart wearable interactive system based on AI language models and AR technology, characterized in that, include: The intent understanding module is used to collect multimodal input data from users through smart wearable devices, perform intent understanding on the multimodal input data based on an AI language model, and generate interactive intent information. The visualization rendering module is used to generate AR enhanced data based on the interaction intent information and environmental context data, and to perform visualization rendering of the AR enhanced data through the AR display module of the smart wearable device to generate user feedback data. The data adjustment module is used to adjust the AR augmented data in real time based on the user feedback data to complete the intelligent interaction process.
2. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 1, characterized in that, The intent understanding module includes: The feature extraction unit is used to extract features from the multimodal input data to obtain multimodal feature vectors; The joint reasoning unit is used to input the multimodal feature vector into the AI language model for joint reasoning to generate intermediate representation information for intent understanding; The context-aware correction unit is used to perform context-aware correction based on the intermediate representation information of the intent understanding to generate the interaction intent information.
3. The intelligent wearable interactive system based on AI language models and AR technology as described in claim 2, characterized in that, The feature extraction unit includes: The multi-source parsing subunit is used to perform multi-source parsing based on the multimodal input data to obtain speech stream data, image sequence data, and motion sensing data; The frame-segmentation and windowing processing subunit is used to perform frame-segmentation and windowing processing based on the speech stream data to generate speech feature vectors. The keyframe extraction subunit is used to extract keyframes from the image sequence data, determine multiple keyframes for target detection, and generate visual feature vectors. The filtering and denoising subunit is used to filter and denoise the motion sensing data to generate a motion feature vector. The dimension normalization subunit is used to normalize the dimensions of the speech feature vector, the visual feature vector, and the motion feature vector to generate a multimodal feature vector.
4. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 3, characterized in that, The joint reasoning unit includes: A cross-modal attention computation subunit is used to perform cross-modal attention computation on the speech feature vector, the visual feature vector, and the motion feature vector to generate a multimodal association mapping network; The hierarchical processing subunit is used to synchronize the multimodal feature vectors to the AI language model for hierarchical processing according to the multimodal association mapping network. S1: Traverse the multimodal association mapping network to perform temporal analysis and capture temporal dependencies; S2: Record the interactive evolution of the multimodal feature vectors according to the temporal dependency relationship to construct the dynamic evolution trend of interactive intent; S3: Based on the dynamic evolution trend of the interaction intent, perform hierarchical reasoning to generate preliminary intent representation information; A cross-modal weight allocation subunit is used to perform cross-modal weight allocation based on the preliminary intent representation information to determine the intermediate representation information for intent understanding.
5. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 1, characterized in that, The visualization rendering module includes: The semantic parsing unit is used to perform semantic parsing based on the interaction intent information to obtain multiple parsed data, wherein the multiple parsed data includes intent type data, target object data, and operation instructions; The scene geometric constraint unit is used to perform scene geometric constraints based on the environmental context data to determine a feasible placement area. The retrieval and matching unit is used to perform retrieval and matching based on the intent type data and the target object data to determine the virtual object AR content data. An instantiation configuration unit is used to execute the operation instructions according to the feasible placement area to instantiate and configure the AR content data of the virtual object, and generate the AR augmented data.
6. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 1, characterized in that, The visualization rendering module includes: A virtual object determination unit is used to receive and analyze the AR augmented data through the AR display module of the smart wearable device to determine virtual objects; The real-time lighting and shadow rendering unit is used to perform real-time lighting and shadow rendering on the virtual object and generate shadow data of the virtual object. The visual fusion unit is used to visually fuse the virtual object with the environmental data according to the shadow data of the virtual object to generate AR visual content; The multi-dimensional acquisition unit is used to overlay the AR visual content onto the user's real field of vision for multi-dimensional acquisition and generate user feedback data.
7. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 6, characterized in that, The visual fusion unit includes: The channel construction subunit is used to perform geometric transformations on the virtual object according to the shadow data of the virtual object and establish a geometric processing channel. The geometric processing channel includes a lighting calculation subchannel, a visual processing subchannel, and a visual fusion subchannel. A multi-source lighting calculation subunit is used to perform multi-source lighting calculations on virtual objects through the lighting calculation sub-channel to generate first visual content; The color grading subunit is used to perform color grading on virtual objects by combining the visual processing subchannel with environmental data, and generate second visual content. The depth test fusion subunit is used to perform depth test fusion of the first visual content and the second visual content through the visual fusion subchannel to construct the AR visual content.
8. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 6, characterized in that, The multidimensional acquisition unit includes: The user data acquisition subunit is used to project the AR visual content onto the user's retina through an optical see-through display module to collect the corresponding physiological data of the user. The motion information capture subunit is used to capture user spatial motion information based on the AR visual content using a depth camera, and to obtain user behavior interaction data. The preference analysis subunit is used to perform user preference analysis based on the AR visual content and obtain user subjective feedback data. The structured integration subunit is used to structure and integrate the user's physiological response data, the user's behavioral interaction data, and the user's subjective feedback data to generate user feedback data.
9. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 8, characterized in that, The data adjustment module includes: An eye-tracking unit is used to track the user's eyes based on the user's physiological response data and determine the user's gaze point information; A gesture recognition unit is used to perform gesture recognition based on the user behavior interaction data to determine user gesture interaction information. The preference analysis unit is used to perform preference analysis based on the user's subjective feedback data to determine the user's personalized association information. The dynamic adjustment unit is used to dynamically adjust the AR augmented data according to the user's personalized association information, the user's gaze point information, and the user's gesture interaction information, and to formulate an interaction strategy. The continuous monitoring unit is used to execute the interaction strategy in real-time feedback loop to continuously monitor and obtain user experience indicators. The strategy optimization unit is used to perform reinforcement learning on the intelligent interaction between the AI language model and AR technology according to the user experience indicators to optimize the interaction strategy and construct the interaction optimization strategy.
10. The intelligent wearable interactive system based on AI language model and AR technology as described in claim 9, characterized in that, The strategy optimization unit includes: A reward signal determination subunit is used to use the user experience metric as a reward signal; The state space determination subunit is used to define the state space based on the user gesture interaction information. The action space determination subunit is used to define the action space based on the user gaze point information; The interaction reinforcement subunit is used to perform reinforcement learning on the intelligent interaction between the AI language model and AR technology according to the reward signal, the state space, and the action space, and to construct the interaction optimization strategy.