AI glasses and smart phone data interaction system under intelligent scene perception
By enabling bidirectional data transmission and dynamic task allocation between AI glasses and smartphones, the problem of low collaboration efficiency is solved, achieving balanced resource utilization and improved real-time perception capabilities, ensuring efficient processing and energy efficiency optimization in complex scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- RUISHI DIGITAL (SHENZHEN) TECHNOLOGY CO LTD
- Filing Date
- 2025-10-21
- Publication Date
- 2026-05-29
AI Technical Summary
AI glasses suffer from low collaboration efficiency with smartphones, uneven resource utilization, and insufficient real-time perception and processing capabilities in complex scenarios. This results in high communication latency, high energy consumption, and an inability to fully utilize the advantages of dual-end computing, affecting the application effect of smart devices in complex scenarios.
The intelligent scene perception module reads position, orientation, and attitude angle data, performs scene data initialization on the smartphone, activates the scene perception sensor of the AI glasses, performs perception processing based on scene framework data, dynamically configures perception tasks through a two-end communication link, and uses resource status perception results for task allocation and adaptive processing.
It achieves efficient collaboration between AI glasses and smartphones, dynamic load balancing, improves the real-time performance and energy efficiency of scene perception, and ensures efficient processing and resource optimization in dynamic environments.
Smart Images

Figure CN121397131B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data interaction technology, specifically to a data interaction system between AI glasses and smartphones under intelligent scene perception. Background Technology
[0002] With the development of Augmented Reality (AR) and Artificial Intelligence (AI), smart devices such as AI glasses and smartphones have become important tools. AI glasses, with their portability and real-time interaction capabilities, have shown great potential in fields such as navigation, education, industrial maintenance, and healthcare. However, AI glasses are limited by computing resources, battery life, and sensor capabilities, making it impossible for them to independently complete complex scene perception and data processing tasks. In contrast, smartphones possess more powerful processing capabilities, richer sensor arrays, and mature communication modules, but their mobility and real-time performance are inferior to AI glasses. Therefore, how to achieve efficient collaboration between AI glasses and smartphones to improve the overall performance of intelligent scene perception has become a key issue. Existing data interaction between smart devices largely relies on one-way transmission or simple command synchronization, lacking dynamic resource coordination and task allocation mechanisms. For example, AI glasses may only send raw sensor data to smartphones for processing, resulting in high communication latency, high energy consumption, and failure to fully utilize the computing advantages of both devices. Furthermore, scene perception tasks typically require real-time responses to environmental changes, making it difficult to guarantee processing efficiency and accuracy under resource-constrained conditions, thus affecting the application effectiveness of smart devices in complex scenarios such as dynamic navigation, multi-target recognition, and real-time decision-making.
[0003] Therefore, current technologies suffer from technical problems such as low efficiency in collaboration between AI glasses and smartphones, uneven resource utilization, and insufficient real-time perception and processing capabilities in complex scenarios. Summary of the Invention
[0004] This application provides a data interaction system and method for AI glasses and smartphones under intelligent scene perception, which solves the technical problems of low collaboration efficiency, uneven resource utilization, and insufficient real-time perception and processing capabilities in complex scenarios in the prior art. It achieves the technical effects of realizing efficient collaboration between AI glasses and smartphones, dynamic load balancing, and improving the real-time performance and energy efficiency of scene perception.
[0005] This application provides a data interaction system between AI glasses and a smartphone under intelligent scene perception. The system includes: a scene perception module, used to read the position data, orientation data, and attitude angle data of the AI glasses, perform scene data initialization on the smartphone, and extract scene frame data; a perception processing module, used to activate the scene perception sensor of the AI glasses after the scene frame data is sent to the AI glasses through a dual-end communication link, perform scene perception based on the scene frame data, and establish a low-dimensional scene state vector; a discrimination module, used to perform resource state perception of the AI glasses and the smartphone, and dynamically configure perception tasks based on the resource state perception results and the low-dimensional scene state vector; and an interaction processing module, used to segment the low-dimensional scene state vector according to the perception task, package the low-dimensional scene state vector identified by the smartphone, and upload it to the smartphone through the dual-end communication link, and perform scene recognition processing using the smartphone and the AI glasses respectively.
[0006] In a possible implementation, the AI glasses and smartphone data interaction system under intelligent scene perception further performs the following processing: collecting processor utilization, battery remaining, and sensor occupancy rates of the AI glasses to establish a resource state description vector for the AI glasses; collecting computing idle time and cache usage data of the smartphone to establish a resource state description vector for the smartphone; acquiring heat generation data of the AI glasses and smartphone, and establishing a resource state penalty factor based on the heat generation data; normalizing the resource state description vectors of the AI glasses and smartphone, and then performing principal component extraction to obtain resource load feature vectors for the AI glasses and smartphone respectively; calculating the rate of change of the resource load feature vector using a sliding window to establish a time-series state matrix of resource state trends; performing multi-scale fusion on the time-series state matrix, and then penalizing and compensating through the resource state penalty factor to establish a resource state perception result.
[0007] In a possible implementation, the AI glasses and smartphone data interaction system under intelligent scene perception further performs the following processing: clustering analysis on the low-dimensional scene state vector, the clustering analysis including clustering analysis based on spatial location, target semantics, and dynamic features, and establishing multi-level clustering results; obtaining the communication bandwidth, latency, and power consumption of the two-end communication link, and establishing a transmission penalty coefficient; using the transmission penalty coefficient as a constraint term, performing adaptive task adaptation analysis on the smartphone and AI glasses using the resource state perception results and the multi-level clustering results, and establishing adaptation analysis results; and completing the dynamic configuration of the perception task based on the adaptation analysis results.
[0008] In a possible implementation, the AI glasses and smartphone data interaction system under intelligent scene perception further perform the following processing: establishing a spatial pose vector based on the location data, orientation data, and attitude angle data; calling the local map and environmental semantic database on the smartphone based on the spatial pose vector to retrieve environmental semantic blocks within the current field of view; performing environmental complexity analysis on the environmental semantic blocks to establish a first extraction constraint; obtaining the accuracy setting of prior perception and establishing a second extraction constraint based on the accuracy setting; after fusing the first and second extraction constraints, performing feature compression and structural simplification of the environmental semantic blocks to establish scene framework data composed of geometric contours and semantic label boundaries.
[0009] In a possible implementation, the AI glasses and smartphone data interaction system under intelligent scene perception also performs the following processing: verifying the authenticity and timeliness of the scene framework data and establishing a verification result; if the verification result is a pass result, directly activating the scene perception sensor, using the scene framework data as a priori template, performing modal alignment and feature decoupling of visual and acoustic sensing signals, using a cross-modal attention mechanism to perform semantic consistency fusion of multi-source perception information, and establishing a scene semantic feature set; performing feature dimensionality reduction and state embedding calculation on the scene semantic feature set, and projecting the scene semantic feature set into a low-dimensional scene state vector.
[0010] In a possible implementation, the AI glasses and smartphone data interaction system under intelligent scene perception further performs the following processing: if the verification result is a verification failure, a verification acquisition command is activated; the scene perception sensor is controlled to acquire data according to the verification acquisition command to establish an acquisition dataset; the same-position feature of the scene frame data is extracted from the acquisition dataset, and feature verification comparison is performed; the feature verification comparison result is used as a frame template to perform linkage feature analysis of the acquisition dataset, establish a scene semantic feature set, and establish a low-dimensional scene state vector according to the scene semantic feature set.
[0011] In a possible implementation, the AI glasses and smartphone data interaction system under intelligent scene perception further includes: a perception enhancement module, used to perform active gaze perception of the user based on the AI glasses, establish attention meta-events, send the attention meta-events to the smartphone, establish scene attention, perform scene trigger analysis based on the scene attention results, establish enhanced content, and send the enhanced content back to the AI glasses for display.
[0012] In a possible implementation, the AI glasses and smartphone data interaction system under intelligent scene perception further includes: a short-term scene prediction module, used to obtain the user's preset resource library, perform short-term scene prediction based on the scene recognition result and the preset resource library, establish predicted interaction content, and push the predicted interaction content to the AI glasses for display.
[0013] This application proposes an AI glasses and smartphone data interaction system with intelligent scene perception. The system comprises a scene perception module for reading the AI glasses' position, orientation, and attitude angle data and initializing the smartphone's scene data; a perception processing module for activating the AI glasses' scene perception sensors and performing scene perception; a discrimination module for performing resource status perception on both the AI glasses and smartphone and dynamically configuring perception tasks; and an interaction processing module for packaging and uploading the low-dimensional scene state vector identified by the smartphone, and performing scene recognition processing on each module. This system addresses the technical problems of low collaboration efficiency between AI glasses and smartphones, uneven resource utilization, and insufficient real-time perception and processing capabilities in complex scenarios in existing technologies. It achieves efficient collaboration between AI glasses and smartphones, dynamic load balancing, and improved real-time scene perception and energy efficiency. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments of this disclosure will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the platform according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0015] Figure 1 This is a schematic diagram of the structure of an AI glasses and smartphone data interaction system under intelligent scene perception provided in an embodiment of this application.
[0016] Figure 2 This is a schematic diagram illustrating the execution flow of the discrimination module in the AI glasses and smartphone data interaction system under intelligent scene perception provided in an embodiment of this application.
[0017] Figure labeling: Scene perception module 10, perception processing module 20, discrimination module 30, interaction processing module 40. Detailed Implementation
[0018] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0019] This application provides an AI glasses and smartphone data interaction system under intelligent scene perception. The system performs bidirectional data transmission between the AI glasses and the smartphone through a dual-end communication link.
[0020] Preferably, the dual-end communication link refers to the data transmission channel established between the AI glasses and the smartphone, used to perform bidirectional data transmission between the AI glasses and the smartphone. The dual-end communication link is implemented based on short-range wireless communication protocols, such as Bluetooth or Wi-Fi, ensuring reliable and orderly data transmission. Bidirectional transmission means that both the AI glasses and the smartphone send and receive data streams. The data stream from the smartphone to the AI glasses may include smartphone initialization data, perception task instructions, and predicted augmented / interactive content. The data stream from the AI glasses to the smartphone may include sensor data, scene feature data, resource status information from the AI glasses, and task execution result data. The AI glasses and the smartphone process tasks in parallel according to task allocation and transmit the necessary results bidirectionally, ultimately presenting them to the user on the glasses.
[0021] like Figure 1 As shown, the AI glasses and smartphone data interaction system under intelligent scene perception includes:
[0022] The scene perception module 10 is used to read the position data, orientation data, and attitude angle data of the AI glasses, and then perform scene data initialization on the smartphone and extract scene frame data.
[0023] Preferably, the AI glasses use built-in GPS, Wi-Fi positioning, or Bluetooth beacon components to read three-dimensional coordinates as position data to define their location in space. The magnetometer inside the AI glasses provides the orientation data, i.e., the azimuth angle pointed to by the embedded lens. The inertial measurement unit inside the AI glasses acquires attitude angle data such as pitch, roll, and yaw angles to accurately describe the tilt and rotation of the AI glasses in space. Then, the smartphone receives the position, orientation, and attitude angle data from the AI glasses and performs scene data initialization. This involves loading and preparing prior environmental data for storage in memory, typically including pre-downloaded local map data such as indoor floor plans, building structure diagrams, or high-precision outdoor maps, as well as environmental semantic data, containing object information with semantic tags bound to map coordinates, such as the door of meeting room A, a coffee machine of brand B, and a fire hydrant in area C. Based on the smartphone's initialized environmental data and the AI glasses' real-time pose, the smartphone performs data retrieval and simplified calculations to generate scene framework data.
[0024] Furthermore, the specific configuration of the scene perception module 10 also includes: establishing a spatial pose vector based on the position data, direction data, and attitude angle data; calling the local map and environmental semantic database on the smartphone based on the spatial pose vector to retrieve environmental semantic blocks within the current field of view; performing environmental complexity analysis on the environmental semantic blocks to establish a first extraction constraint; obtaining the accuracy setting of prior perception and establishing a second extraction constraint based on the accuracy setting; and after fusing the first and second extraction constraints, performing feature compression and structural simplification of the environmental semantic blocks to establish scene framework data composed of geometric contours and semantic label boundaries.
[0025] Preferably, the position, orientation, and pose angle data of the AI glasses are unified and fused into a spatial pose vector to completely and accurately describe the AI glasses' 3D position and orientation in the global coordinate system. Then, using the spatial pose vector as camera parameters, frustum culling calculations are performed on the local map data and environmental semantic data from the smartphone. Only multiple environmental semantic blocks currently within the possible field of view of the AI glasses lens are retrieved; that is, all tagged objects or structures and their 3D data currently within the theoretical field of view of the AI glasses. Next, environmental complexity analysis of the environmental semantic blocks is performed, which involves quantitative analysis of the retrieved environmental semantic blocks, including calculating the number, density, and type of objects within the current field of view. A complexity index is generated based on the current objective environment as the first extraction constraint; higher complexity results in greater simplification, while lower complexity allows for the retention of more details.
[0026] Preferably, the system acquires the user-preset accuracy, such as high-precision mode, balanced mode, or low-power mode. A second extraction constraint is configured based on the user-preset accuracy and resource requirements, representing the task's detail requirements. For example, the constraint for high-precision mode is minimizing simplification, while the constraint for low-power mode is maximizing simplification. The first and second extraction constraints are fused using a weighted average to generate simplified control instructions. These instructions are then used to perform feature compression and structural simplification of the environmental semantic blocks. Specifically, for geometric contours, complex 3D meshes are represented with fewer vertices and faces, such as using a cube instead of a complex computer host, or directly extracting the bounding box of the object as its geometric contour. Semantic labels are merged or filtered; for example, in simplified mode, individual labels such as keyboard and mouse are ignored, while the overall label of the desk is retained. Finally, scene framework data composed of geometric contours and semantic label boundaries is obtained. The geometric contour is a simplified geometric body representing the spatial extent of an object, and the semantic label boundary is a semantic label closely bound to that geometric body. This ensures that the smartphone transmits a large amount of data to the AI glasses with minimal data volume and efficiently guides its perception tasks.
[0027] The perception processing module 20 is used to activate the scene perception sensor of the AI glasses after the scene frame data is sent to the AI glasses through the dual-end communication link, and to perform scene perception based on the scene frame data to establish a low-dimensional scene state vector.
[0028] Preferably, the AI glasses receive scene frame data from the smartphone via a dual-end communication link and use it as a control signal to trigger the AI glasses' power management or task scheduling unit. This triggers the scene perception sensor, which is in low-power standby mode, to switch to active working mode as needed. The AI glasses' scene perception sensor then uses the scene frame data for scene perception, including rapidly comparing and matching real-time image data with the scene frame data, spatially aligning multiple modal data, focusing on the corresponding region of the image based on the scene frame data and decoupling relevant features, fusing multiple semantic visual features, depth information, and sound cues through a cross-modal attention mechanism to ensure fusion accuracy, and finally integrating the obtained multi-dimensional features into high-dimensional features, performing feature dimensionality reduction, and projecting them onto a preset low-dimensional vector space. Each vector corresponds to a specific state of the scene, ultimately outputting a low-dimensional scene state vector. This ensures scene perception efficiency and accuracy while reducing communication bandwidth requirements and power consumption.
[0029] Furthermore, the specific configuration of the perception processing module 20 also includes: verifying the authenticity and timeliness of the scene framework data and establishing a verification result; if the verification result is a pass result, directly activating the scene perception sensor, using the scene framework data as a priori template, performing modal alignment and feature decoupling of visual and acoustic sensing signals, using a cross-modal attention mechanism to perform semantic consistency fusion of multi-source perception information, and establishing a scene semantic feature set; performing feature dimensionality reduction and state embedding calculation on the scene semantic feature set, and projecting the scene semantic feature set into a low-dimensional scene state vector.
[0030] Preferably, the scene frame data undergoes authenticity verification. This involves the AI glasses rapidly activating scene perception sensors to capture data and extract key features such as edges, corners, and significant semantic features. These features are then quickly matched with the geometric contours and semantic label boundaries in the scene frame data. If the matching degree exceeds a preset threshold—for example, if the position and shape of the object predicted by the frame data largely match the position and shape of the detected object in the image—then the authenticity verification is passed, ensuring that the scene frame data is consistent with the current physical environment. Next, the scene frame data undergoes timeliness verification. This involves checking whether there is any potentially outdated dynamic object information in the scene frame data. For example, the scene frame data indicates that a chair is in a certain position, but the scene perception sensors of the AI glasses indicate that the position is currently empty. If the number and significance of such cases exceed a preset threshold, the timeliness verification fails, indicating that the environment has changed significantly since the scene frame data was generated. Finally, the verification result is output, including a pass result or a fail result.
[0031] Preferably, if the verification result is satisfactory, the scene perception sensor is directly activated. Using the scene frame data as a priori template, modal alignment of visual and acoustic sensor signals is performed. Specifically, the raw data from sensors such as cameras and microphones are spatiotemporally aligned and registered to ensure that the image captured by the camera, the sound segment recorded by the microphone, and the pose of the glasses are all in the same coordinate system at every moment. Then, feature decoupling is performed, that is, for each region of interest marked by the scene frame data, independent features are extracted from different modal data. For example, for the screen area marked by the scene frame data indicating that someone is speaking, features such as whether the screen content has changed and whether there is lip movement are decoupled from the visual modality, and features such as the direction of the sound source and the speech content are decoupled from the acoustic modality. Then... Semantic consistency fusion is performed through a cross-modal attention mechanism. When visual and acoustic features are highly correlated, i.e., lip reading and speech are synchronized and homologous, high attention weights are obtained. Then, all modal features weighted by attention weights are integrated to generate a scene semantic feature set, which contains high-dimensional feature representations of mutual verification and supplementary information between modalities. Then, principal component analysis or autoencoder is used to reduce the dimensionality of the features, remove redundant information from the scene semantic feature set, retain the most discriminative core features, and project the dimensionality-reduced features into a preset low-dimensional embedding space. Semantically similar scene states are mapped to vector positions that are close to each other, and different scenes are mapped to positions that are far apart. Finally, a low-dimensional scene state vector is output, thereby ensuring efficient intelligent scene perception.
[0032] Furthermore, the specific configuration of the perception processing module 20 also includes: if the verification result is a verification failure result, then activating a verification acquisition command; controlling the scene perception sensor to acquire data according to the verification acquisition command, and establishing an acquisition dataset; extracting the same position features of the scene frame data from the acquisition dataset, and performing feature verification comparison; using the feature verification comparison result as a frame template, performing linkage feature analysis of the acquisition dataset, establishing a scene semantic feature set, and establishing a low-dimensional scene state vector based on the scene semantic feature set.
[0033] Preferably, if the verification result is a failure, indicating that the scene frame data is unreliable, a verification acquisition command is activated. This command is more comprehensive and energy-intensive than the command in the normal perception mode. Then, according to the verification acquisition command, the scene perception sensor is controlled to acquire data at a higher resolution, more modalities, or for a longer duration to obtain more complete raw multimodal sensor data, such as image sequences, depth maps, point clouds, audio clips, etc., to determine the acquisition dataset. Next, the same-location features of the scene frame data are extracted in the acquisition dataset, that is, the visual and depth features of the same location as those described by the failed scene frame data are determined, and these features are compared with the features in the scene frame data to obtain the feature verification comparison results, including pass / fail judgments and specific difference reports. Then, the feature verification comparison results are used as a frame template to perform linked feature analysis of the acquisition dataset, including associative perception in the acquisition dataset, such as joint analysis of the visual appearance, three-dimensional shape, and material of the area, focusing on multimodal recognition of unknown objects and determination of semantic relationships, thereby outputting a scene semantic feature set. Finally, feature dimensionality reduction and state embedding calculations are performed on the scene semantic feature set to generate a low-dimensional scene state vector that accurately reflects the current state of the environment.
[0034] The discrimination module 30 is used to perform resource status perception of AI glasses and smartphones, and dynamically configures perception tasks based on resource status perception results and low-dimensional scene state vectors.
[0035] Furthermore, such as Figure 2 As shown, the specific configuration of the discrimination module 30 also includes: collecting and acquiring the processor utilization rate, battery remaining amount, and sensor occupancy rate of the AI glasses to establish a resource status description vector for the AI glasses; collecting the computing idle time and cache occupancy data of the smartphone to establish a resource status description vector for the smartphone; acquiring the heat generation data of the AI glasses and the smartphone, and establishing a resource status penalty factor based on the heat generation data; normalizing the resource status description vectors of the AI glasses and the smartphone, and then performing principal component extraction to obtain the resource load feature vectors of the AI glasses and the smartphone respectively; calculating the rate of change of the resource load feature vector using a sliding window to establish a time-series state matrix of resource status trends; performing multi-scale fusion on the time-series state matrix, and then penalizing and compensating it through the resource status penalty factor to establish a resource status perception result.
[0036] Preferably, resource monitoring is performed on both the AI glasses and the smartphone to obtain the processor utilization, battery level, and sensor occupancy rate of the AI glasses. Processor utilization refers to the percentage of CPU / GPU usage, reflecting computational load; battery level refers to the percentage of remaining power, reflecting energy constraints; sensor occupancy rate refers to whether sensors such as cameras are occupied by the current task or other tasks, reflecting hardware channel availability. These data are then combined to generate a resource status description vector for the AI glasses, used to describe the instantaneous static resource status of the AI glasses. Computational idle time and cache usage data are collected from the smartphone. Computational idle time characterizes the phone's available computing power and complements processor utilization; cache usage data refers to the occupancy of memory or dedicated cache, reflecting potential bottlenecks in data processing. These data are also combined to construct a resource status description vector for the smartphone.
[0037] Preferably, temperature sensors are used to acquire heat data from AI glasses and smartphones. Excessive temperature may cause system frequency reduction, resulting in decreased actual performance even if other indicators appear good. Then, a resource status penalty factor is established based on the heat data. The resource status penalty factor is positively correlated with temperature; for example, the higher the temperature, the larger the resource status penalty factor, which is used to discount the effective performance of AI glasses and smartphones. The resource status description vectors of AI glasses and smartphones are normalized, including mapping indicators with different dimensions and ranges to the same scale through min-max normalization or Z-score standardization, so that different indicators can be compared and weighted. Then, principal component analysis is used to extract principal components from the normalized resource status vectors. That is, principal components that are independent and can represent most of the resource status information are extracted from multiple original indicators that may be correlated, thereby obtaining the resource load feature vectors of AI glasses and smartphones respectively.
[0038] Preferably, a sliding window is preset, and the rate of change of the resource load feature vector arranged in time series according to the sliding window is calculated. That is, the first derivative of the principal component of the data within each sliding window is calculated as the rate of change to quantify the direction and speed of change of resource load, such as whether the CPU load is rising rapidly or falling slowly. Then, the resource load feature vector at the current moment and its rate of change vector within the sliding window are integrated into a time series state matrix, which includes both the current state and the recent trend. Then, different weights are assigned to the trends of different time scales of the time series state matrix for multi-scale fusion. Finally, a resource state penalty factor is applied to the fused time series state matrix for penalty compensation. For example, the load value in the matrix is multiplied by the penalty factor, and finally, the resource state perception result is generated.
[0039] Furthermore, the specific configuration of the discrimination module 30 also includes: performing cluster analysis on the low-dimensional scene state vector, the cluster analysis including cluster analysis based on spatial location, target semantics, and dynamic features, and establishing multi-level clustering results; obtaining the communication bandwidth, latency, and power consumption of the dual-end communication link, and establishing a transmission penalty coefficient; using the transmission penalty coefficient as a constraint term, and using the resource state perception results and the multi-level clustering results to perform adaptive task adaptation analysis for smartphones and AI glasses, and establishing adaptation analysis results; and completing the dynamic configuration of the perception task based on the adaptation analysis results.
[0040] Preferably, based on spatial location, target semantics, and dynamic features, cluster analysis is performed on the low-dimensional scene state vector representing the entire scene. Specifically, adjacent objects or regions in space are clustered into one category for localized processing; semantically similar objects are clustered into another category to invoke a dedicated recognition model; and objects with similar motion states are clustered into another category to distinguish between static environments and dynamic targets. This further identifies sub-regions or sub-targets with similar characteristics, obtaining multi-level clustering results. That is, the low-dimensional scene is divided into multiple different subsets from different perspectives, for example, simultaneously divided into static background clusters, dynamic pedestrian clusters, near-field table clusters, and distant wall clusters. The communication bandwidth, latency, and power consumption of the two-way communication link are obtained to determine the data transmission rate, real-time data transmission, and energy cost of data transmission, respectively. A transmission cost function is constructed based on a weighted comprehensive analysis of communication bandwidth, latency, and power consumption to evaluate the transmission penalty coefficient. When bandwidth is low, latency is high, and power consumption is high, the transmission penalty coefficient is high, meaning the data transmission cost is high.
[0041] Preferably, adaptive task adaptation analysis for smartphones and AI glasses is performed based on resource status perception results and multi-level clustering results. Specifically, the transmission penalty coefficient quantifies the communication cost, the resource status perception results quantify the computational cost and availability of the glasses and the phone, and the multi-level clustering results quantify the characteristics and structure of the task itself. Then, using the transmission penalty coefficient as a constraint, adaptive task adaptation analysis is performed based on linear programming or heuristic optimization algorithms. Specifically, for each multi-level clustering result, the total cost of processing on the glasses and the phone is evaluated, where the processing cost on the glasses is determined based on the resource load of the glasses and the computational complexity of the task. The processing cost on the mobile device is determined based on the product of the mobile device's resource load, the computational complexity of the task, the transmission penalty coefficient, and the amount of data in the task cluster. Then, the processing costs on the glasses and the mobile device are compared, and the processing device with the lower total cost is selected for each level of clustering results. Finally, the adaptation analysis results are determined, which include a clear task allocation mapping table. Based on the adaptation analysis results, specific configuration instructions are generated and sent to the AI glasses or smartphone. Static backgrounds and specific pedestrian tracking are processed locally by the AI glasses, while complex text recognition and large-scale semantic segmentation are processed by the smartphone, ensuring that the perception tasks are dynamically and optimally configured on the AI glasses and smartphone.
[0042] The interaction processing module 40 is used to segment the low-dimensional scene state vector according to the perception task, package the low-dimensional scene state vector identified by the smartphone, and upload it to the smartphone through a dual-end communication link, and perform scene recognition processing using the smartphone and AI glasses respectively.
[0043] Preferably, the low-dimensional scene state vector is logically segmented based on the dynamic configuration results of the perception task. This includes identifying task clusters within the low-dimensional scene state vector that require processing by the smartphone. The low-dimensional scene state vectors identified by the smartphone are then packaged into data packets containing data sequence numbers, task IDs, timestamps, etc., and uploaded to the smartphone via a dual-end communication link. The AI glasses and the smartphone then work in parallel according to the assigned tasks to perform scene recognition processing. Specifically, after receiving the data, the smartphone parses and reads the corresponding low-dimensional scene state vectors and calls an artificial intelligence model to perform deep analysis and high-precision recognition on these vectors. This may include complex object recognition, scene semantic segmentation, or accessing cloud databases for information retrieval, and generating high-precision scene recognition results on the smartphone. Simultaneously, the AI glasses process the low-dimensional scene state vectors that were not uploaded locally, including performing tasks with high real-time requirements and relatively low computational load, such as basic object tracking, spatial plane detection, or simple gesture recognition, and generating fast and lightweight scene recognition results on the AI glasses. This ensures that computationally intensive tasks are performed on the smartphone and real-time tasks are performed on the AI glasses, improving the real-time performance and energy efficiency of scene recognition perception.
[0044] Furthermore, the AI glasses and smartphone data interaction system under intelligent scene perception also includes a perception enhancement module, which is used to perform active gaze perception of the user based on the AI glasses, establish attention meta-events, send the attention meta-events to the smartphone, establish scene attention, perform scene trigger analysis based on the scene recognition results according to the scene attention, establish enhanced content, and send the enhanced content back to the AI glasses for display.
[0045] Preferably, the perception enhancement module achieves proactive visual enhancement collaboration by introducing scene triggers. The AI glasses, through built-in eye-tracking sensors, monitor the user's eye movements in real time and accurately calculate the user's gaze point and gaze duration. For example, when it detects that the user is continuously gazing at a specific object or area for more than 3 seconds, a structured attention meta-event is generated, including the spatial coordinates of the gaze target, the start time and duration of the gaze, and the preliminary identification of the gazed object. The structured meta-event is then sent to the smartphone through a two-way communication link. The smartphone fuses it with the scene recognition results currently being processed to clearly identify the user's current focus of interest. Then, based on the attention meta-event and scene semantics, intelligent push or enhanced content feedback is performed, such as displaying detailed descriptions, AR tags, and instant translation results. Finally, the generated enhanced content is sent back to the AI glasses. After receiving it, the display driver component of the AI glasses accurately maps the virtual enhanced content onto the real object that the user is gazing at and displays it through optical lenses, achieving the fusion of virtual and reality.
[0046] Furthermore, the AI glasses and smartphone data interaction system under intelligent scene perception also includes a short-term scene prediction module, which is used to obtain the user's preset resource library, perform short-term scene prediction based on the scene recognition results and the preset resource library, establish predicted interaction content, and push the predicted interaction content to the AI glasses for display.
[0047] Preferably, the short-term scene prediction module reads a preset resource library from the smartphone's storage space, namely a digital user model and content database, which may include frequently performed operations or viewed information in specific scenarios, user personal preference settings, pre-associated digital content, and a scene-action rule library. Then, based on the scene recognition results and the preset resource library, the module uses the computing power of the mobile phone to perform short-term scene prediction on the environment in which the AI glasses are located. That is, the rule-based inference engine matches and analyzes the current scene recognition results with the preset resource library, outputs short-term scene prediction results, and determines the predicted interaction content. For example, if the mobile phone predicts that the user is about to enter a meeting scene, it pushes the meeting schedule to the AI glasses in advance for display. When the AI glasses detect scene switching signals such as acoustic echo or sudden changes in lighting, it quickly activates the pre-fetched content for seamless switching and pushes the predicted interaction content to the AI glasses' graphics cache through a dual-end communication link, thereby realizing a predictive preloading and rendering engine for zero-wait AR experience to improve the smoothness of the interaction system.
[0048] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A data interaction system between AI glasses and a smartphone under intelligent scene perception, characterized in that, The system performs bidirectional data transmission between the AI glasses and the smartphone via a dual-end communication link, and the system includes: The scene perception module is used to read the position data, orientation data, and attitude angle data of the AI glasses, and then perform scene data initialization on the smartphone and extract scene frame data. The perception processing module is used to activate the scene perception sensor of the AI glasses after the scene frame data is sent to the AI glasses through the dual-end communication link, perform scene perception based on the scene frame data, and establish a low-dimensional scene state vector. The discrimination module is used to perform resource status perception of AI glasses and smartphones, and dynamically configures perception tasks based on resource status perception results and low-dimensional scene state vectors. The interactive processing module is used to segment the low-dimensional scene state vector according to the perception task, package the low-dimensional scene state vector identified by the smartphone, and upload it to the smartphone through a dual-end communication link, and perform scene recognition processing using the smartphone and AI glasses respectively. The discrimination module performs resource status perception of the AI glasses and smartphone, including: Collect data on the processor utilization, battery level, and sensor occupancy rate of the AI glasses, and establish a resource status description vector for the AI glasses. Collect data on smartphone computing idle time and cache usage to establish a resource status description vector for smartphones; Acquire heat generation data from AI glasses and smartphones, and establish a resource status penalty factor based on the heat generation data; After normalizing the resource status description vectors of AI glasses and smartphones, principal component extraction is performed to obtain the resource load feature vectors of AI glasses and smartphones respectively. The rate of change of the resource load feature vector is calculated using a sliding window to establish a time-series state matrix of resource status trends. The time-series state matrix is then fused at multiple scales and penalized using the resource status penalty factor to establish a resource status perception result.
2. The AI glasses and smartphone data interaction system under intelligent scene perception as described in claim 1, characterized in that, The discrimination module dynamically configures perception tasks based on resource status perception results and low-dimensional scene state vectors, including: Cluster analysis is performed on the low-dimensional scene state vector, including cluster analysis based on spatial location, target semantics, and dynamic features, to establish multi-level clustering results; Obtain the communication bandwidth, latency, and power consumption of the two-way communication link, and establish the transmission penalty coefficient; Using the transmission penalty coefficient as a constraint, the resource state perception results and the multi-level clustering results are used to perform adaptive task adaptation analysis on smartphones and AI glasses, and adaptation analysis results are established. Dynamic configuration of perception tasks is completed based on the adaptation analysis results.
3. The AI glasses and smartphone data interaction system under intelligent scene perception as described in claim 1, characterized in that, The scene perception module extracts scene framework data, including: A spatial pose vector is established based on the position data, orientation data, and attitude angle data; Based on the spatial pose vector, the local map and environmental semantic database on the smartphone are invoked to retrieve environmental semantic blocks within the current field of view; Analyze the environmental complexity of the execution environment semantic block and establish the first extraction constraints; Obtain the accuracy setting of prior perception, and establish a second extraction constraint based on the accuracy setting; After fusing the first and second extraction constraints, feature compression and structural simplification of the environmental semantic blocks are performed to establish scene framework data composed of geometric contours and semantic label boundaries.
4. The AI glasses and smartphone data interaction system under intelligent scene perception as described in claim 1, characterized in that, In the perception processing module, the scene perception sensor of the AI glasses is activated, and scene perception is performed based on scene framework data to establish a low-dimensional scene state vector, including: The authenticity and timeliness of the scenario framework data are verified, and verification results are established. If the verification result is a pass, the scene perception sensor is directly activated. Using the scene framework data as a priori template, modal alignment and feature decoupling of visual and acoustic sensing signals are performed. A cross-modal attention mechanism is used to perform semantic consistency fusion of multi-source perception information to establish a scene semantic feature set. The scene semantic feature set is subjected to feature dimensionality reduction and state embedding calculations, and the scene semantic feature set is projected into a low-dimensional scene state vector.
5. The AI glasses and smartphone data interaction system under intelligent scene perception as described in claim 4, characterized in that, The perception processing module, in establishing the verification result, also includes: If the verification result is a failure, then the verification collection command is activated; The scene perception sensor is controlled to collect data according to the verification and acquisition command, and a collection dataset is established. The collected dataset is used to extract features from the same location within the scene framework data, and feature verification and comparison are performed. Using the feature verification and comparison results as a framework template, perform linked feature analysis on the collected dataset to establish a scene semantic feature set, and establish a low-dimensional scene state vector based on the scene semantic feature set.
6. The AI glasses and smartphone data interaction system under intelligent scene perception as described in claim 1, characterized in that, The system also includes: The perception enhancement module is used to perceive the user's active gaze based on the AI glasses, establish attention meta-events, send the attention meta-events to the smartphone, establish scene attention, perform scene trigger analysis based on the scene attention results, establish enhanced content, and send the enhanced content back to the AI glasses for display.
7. The AI glasses and smartphone data interaction system under intelligent scene perception as described in claim 1, characterized in that, The system also includes: The short-term scene prediction module is used to obtain the user's preset resource library, perform short-term scene prediction based on the scene recognition results and the preset resource library, establish predicted interactive content, and push the predicted interactive content to the AI glasses for display.