Method for improving man-machine interaction efficiency by sensing environmental information through AI
By collecting images and sensor data in real time in a dynamic industrial environment, combining AI technology to analyze scene and device status, adopting an end-cloud collaboration architecture to generate dynamic recommendation instructions, solving the problem of low human-computer interaction efficiency and achieving efficient and accurate user guidance.
Patent Information
- Application Number
- CN202510283045.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-08-29
AI Technical Summary
In a dynamic industrial environment, the existing technology cannot effectively utilize environmental information, resulting in low human-computer interaction efficiency and poor real-time performance. Users need to frequently operate manually and lack automation assistance, so they cannot quickly adjust recommendations according to environmental changes.
Through mobile devices, real-time acquisition of environmental images and sensor data, combined with AI technology to identify scenes and device status, adopt end-cloud collaborative architecture for lightweight processing and in-depth analysis, generate differentiated recommendation instructions, and use AR visualization and voice prompts to guide users' operations.
It significantly improves human-computer interaction efficiency, reduces user decision-making time, avoids misoperation, improves recommendation accuracy and real-timeness, improves development efficiency by 60%, and reduces manual input errors by 65%.
Smart Images

Figure FT_1
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of artificial intelligence (AI) applications, human-computer interaction technology, industrial automation, and the Internet of Things (IoT) technology, and is particularly suitable for intelligent interaction scenarios in dynamic industrial environments. The system uses mobile devices (such as mobile phones or AR glasses) to collect environmental images and sensor data in real time, combines AI technology to identify scenes and device status, and dynamically recommends differentiated processes based on user roles to improve interaction efficiency. The system adopts an end-to-cloud collaborative architecture, with local devices performing lightweight processing and the cloud performing in-depth analysis to ensure a balance between real-time performance and computing resources. For the use of large models, both local models and third-party APIs are used through APIs. Background Art
[0002] Currently, human-computer interaction technologies primarily rely on manual user interaction or proactive information querying. This is particularly true in dynamic industrial environments, where users frequently switch tasks or rely on experience to determine the next action. Existing technologies suffer from significant deficiencies in terms of insufficient utilization of environmental information, low interaction efficiency, and poor real-time performance. Users must manually perceive scenarios and make decisions, lacking automated assistance. Business function recommendations are not dynamically integrated with user identity and device status, leading to redundant operations. Traditional systems are unable to quickly adjust recommendations based on environmental changes (such as machine status or material location). An intelligent, dynamic solution is urgently needed. Summary of the Invention
[0003] The present invention relates to a method for improving the efficiency of human-computer interaction through AI perception of environmental information. Environmental images and sensor data are collected in real time through mobile devices, and AI technology is combined to identify scenes and equipment status, and differentiated processes are dynamically recommended according to user roles. The system adopts an end-cloud collaborative architecture, with local devices performing lightweight processing and the cloud performing in-depth analysis to ensure a balance between real-time performance and computing resources. Through multimodal data fusion, the accuracy of recommendations is improved, the user's decision-making time is reduced, and human errors are avoided. Expanded application scenarios: The present invention is not only applicable to industrial workshops and warehouse management scenarios, but can also be extended to scenarios in multiple fields such as manufacturing, logistics, and medical care to improve the efficiency of human-computer interaction in various industries. Core technology: The core technology of the present invention includes scene recognition model, equipment status monitoring, and recommendation algorithm. The scene recognition model is based on deep learning image classification technology, and equipment status monitoring obtains machine operation data through the IoT interface. The recommendation algorithm is based on the dynamic decision-making of rule engine and reinforcement learning to ensure the accuracy and real-time performance of recommendations.
[0004] Technical solution: A method for improving human-computer interaction efficiency by perceiving environmental information through AI, characterized by comprising the following steps: a. The user enters the scene with the device and collects environmental data.
[0005] b. Perform edge preprocessing and data normalization on the collected image data and sensor data.
[0006] c. Upload the processed data to the cloud for AI analysis to generate environmental semantics and device status analysis.
[0007] d. Match user roles and permissions, generate recommended instructions, and deliver them to users through AR visualization, voice prompts, or mobile notifications.
[0008] e. After the user performs the operation, the results are fed back and the model and rule base are updated.
[0009] A method for improving human-computer interaction efficiency by using AI to perceive environmental information, characterized by comprising: Scenario-driven intelligent recommendations: AI generates action suggestions through semantic understanding of the environment.
[0010] b. Dual-dimensional matching of roles and scenarios: The recommendation logic considers both user identity permissions and current environmental requirements.
[0011] c. End-to-cloud collaborative architecture: lightweight processing on local devices + deep AI analysis on the cloud. DETAILED DESCRIPTION
[0012] As shown in the accompanying drawings Figure 1 The figure shows the overall flow chart of the system, which illustrates the entire process of a method for improving the efficiency of human-computer interaction through AI perception of environmental information. Starting from the moment the user enters the scene with the device, the device collects environmental data, including image data and sensor data, which is then pre-processed at the edge and standardized before being uploaded to the cloud for AI analysis. Cloud-based AI analysis includes scene classification and device status analysis, and generates environmental semantics through multimodal data fusion. It combines user roles and permissions to generate recommended instructions (such as AR visualization, voice prompts, mobile notifications, etc.), ultimately guiding the user to perform operations. The results of user operations are fed back to the system and used to update the model and rule base, thereby continuously optimizing the recommendation effect: The user enters the scene with the device and collects environmental data.
[0013] Perform edge preprocessing and data normalization on acquired image data and sensor data.
[0014] The processed data is uploaded to the cloud for AI analysis to generate environmental semantics and device status analysis.
[0015] Match user roles and permissions, generate recommended instructions, and deliver them to users through AR visualization, voice prompts, or mobile notifications.
[0016] After the user performs the operation, the results are fed back and the model and rule base are updated.
[0017] Beneficial effects: The present invention significantly improves development efficiency and accuracy through multi-layer context integration, node-specific prompt words, dynamic adaptation mechanism and low-invasive design, and solves problems such as context fragmentation, the contradiction between universality and specificity, and development efficiency bottlenecks.
[0018] Through the automated process of the present invention, the efficiency of code-free platform page development is improved by about 60%, and manual input errors are reduced by about 65%. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 The overall system flow chart shown here illustrates the entire process of a method for improving human-computer interaction efficiency through AI-powered environmental perception. Starting with the user entering a scene with a device, the device collects environmental data, including image and sensor data. This data is then pre-processed and standardized at the edge before being uploaded to the cloud for AI analysis. Cloud-based AI analysis includes scene classification and device status analysis. It generates environmental semantics through multimodal data fusion, and combines user roles with permissions to generate recommended instructions (such as AR visualization, voice prompts, and mobile notifications), ultimately guiding the user to perform actions. User action results are fed back to the system and used to update the model and rule base, thereby continuously optimizing recommendation effectiveness.
Claims
1. A method for improving the efficiency of human-computer interaction by perceiving environmental information through AI, characterized in that include: Environmental data collection: Real-time collection of environmental images and sensor data via mobile devices. Edge preprocessing: Performs noise reduction and object detection on collected image data, and standardizes sensor data. Cloud AI analysis: Uploads preprocessed data to the cloud, where AI models are used for scene classification and device status analysis. Environmental semantic generation: Generates environmental semantic information based on the results of scene classification and device status analysis. User role matching: Matches the generated environmental semantic information with user roles and permissions. Recommended instruction generation: Generates recommended instructions based on the matching results and presents them to users through AR visualization, voice prompts, and mobile notifications. User action execution: Users perform actions based on recommended instructions and provide feedback to the system. Model and rule base update: Updates the AI model and rule base based on user action results to optimize subsequent recommendations. Multimodal data fusion: Combines images, device logs, and historical user behavior data to improve recommendation accuracy. End-to-cloud collaborative architecture: Balances real-time performance with computing resources through lightweight local device processing and in-depth cloud AI analysis.
Citation Information
Cited By
Internet of Things medical equipment intelligent guiding system and method based on multi-mode interaction
CN121364915A
An internet of things medical device intelligent guiding system and method based on multi-modal interaction
CN121364915B