Video call assistance method and system based on intelligent glasses
By collecting environmental information in real time through smart glasses, the cloud platform performs scene recognition and volunteer screening, and combines the evaluation module for quantitative assessment, the problem of relying on manual labor and inaccurate volunteer matching in existing technologies is solved, achieving efficient and transparent remote assistance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-04-07
AI Technical Summary
Existing smart glasses remote assistance systems cannot automatically recognize complex environments, rely on manual judgment, have inaccurate volunteer matching, lack quantitative evaluation, resulting in unstable assistance effects, heavy user workload, and low volunteer participation.
Real-time environmental information is collected through smart glasses, and the cloud platform performs scene recognition and volunteer screening. Combined with the evaluation module, quantitative evaluation is carried out to achieve automatic assistance triggering and volunteer scheduling. Multi-dimensional feature vector matching and weighted similarity algorithm are used to generate volunteer service scores and points incentives.
It enables automatic identification and timely response to complex environments, improves the efficiency and quality of the assistance system, reduces the user's operational burden, and enhances the transparency and participation of volunteer services.
Smart Images

Figure CN121808583A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart glasses technology, and more specifically, to a video call assistance method and system based on smart glasses. Background Technology
[0002] With the development of smart wearable devices, smart glasses, as a new generation of visual interactive terminals, are widely used in navigation, remote control, and assisted recognition scenarios. Especially in the fields of assistive devices for the disabled, guide glasses for the blind, and emergency remote support, smart glasses can achieve real-time perception of the surrounding environment through components such as cameras, microphones, and positioning modules, and transmit the perceived information to the backend or remote terminals. However, existing technologies are mostly limited to transmitting the collected video information to the backend operator, with operational responses relying on manual judgment. They lack intelligent recognition of scene complexity and dynamic task scheduling mechanisms, and in actual assistance processes, the effectiveness of assistance lacks quantitative evaluation methods, making it impossible to effectively measure the quality of volunteer services and the degree of incentive, resulting in unstable system scalability and service quality.
[0003] In addition, current remote assistance systems often have the following problems: First, they cannot automatically identify high-risk or complex situations based on the user's environment, and assistance relies on the user's active operation, increasing the burden on visually impaired users; second, the volunteer matching mechanism is simple and lacks comprehensive consideration of service capabilities, historical records and scenario adaptability; third, the interactive behavior during the assistance process is not systematically collected and evaluated, resulting in low transparency of the volunteer service process, unclear reward and punishment mechanisms, and affecting the enthusiasm of volunteers to participate.
[0004] Therefore, it is necessary to design a video call assistance method and system based on smart glasses to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes a video call assistance method and system based on smart glasses, aiming to solve the current problems of being unable to identify high-risk or complex situations, lacking consideration of service capabilities, and having low volunteer participation.
[0006] In one aspect, the present invention proposes a video call assistance method based on smart glasses, comprising: Smart glasses are used to collect environmental information in real time, including environmental video, environmental audio and location information, and upload the environmental information to a cloud platform through a wireless communication module. The cloud platform performs scene recognition based on the received environmental information to determine whether the current environment is complex. The complex environment includes unstructured environment, dynamically changing environment or high-risk environment. When the scene is determined to be complex, a remote assistance request is triggered. The cloud platform is also configured to screen registered volunteers and dispatch one or more volunteers to assist in the task based on the volunteers' identity information and the tag features of their service records. A volunteer terminal is installed on the smart glasses. The volunteer terminal accesses the video stream of the smart glasses via video call and provides navigation or obstacle prompts to the user through voice commands or audio feedback. An evaluation module, located on the cloud platform, records and analyzes interactive data in real time during the assistance process, obtaining the success rate of prompts, response delays, and user feedback. Based on the success rate of prompts, response delays, and user feedback, a volunteer service score is generated. After the assistance task is completed, the evaluation module generates volunteer points based on the volunteer service score and incorporates these points into the volunteer incentive system for volunteer level assessment, task priority ranking, or awarding of honors.
[0007] Furthermore, the smart glasses further include an environmental anomaly detection unit, which is used to detect whether there are abnormal events during the process of collecting the environmental information. The abnormal events include sudden sounds, changes in light, or continuous occlusion, and when the abnormal event is detected, a remote assistance request is triggered first.
[0008] Furthermore, the cloud platform performs scene recognition based on the received environmental information to determine whether the current environment is complex. When a complex scene is determined and a remote assistance request is triggered, the following steps are taken: The cloud platform performs target detection and semantic segmentation on the environmental video uploaded by the smart glasses, extracts the ground paving status, obstacle density and path clarity, and determines whether the current environment is unstructured based on a convolutional neural network model. The ambient audio is subjected to audio event recognition to extract vehicle horn sounds, mechanical noises and human noises, and the background sound pressure level is used to determine whether it is the dynamically changing environment. Geographic comparison is performed between location information and an external high-risk area database, and the user's current location is in a high-risk area or the user's movement trajectory shows an abnormal change in direction, which is identified as a high-risk environment. When the confidence level of any type of complex environment identification result exceeds the threshold, a remote assistance request is triggered.
[0009] Furthermore, the cloud platform screens registered volunteers and, based on their identity information and service record tags, schedules one or more volunteers to assist with tasks, including: The cloud platform constructs a volunteer profile database, creating a multi-dimensional feature vector for each registered volunteer that includes identity information, historical assistance scenario records, service rating, task completion time preferences, language ability, and assistance domain labels. Upon receiving a remote assistance request, the cloud platform obtains the user's current scene feature vector based on the environmental information, calculates the similarity between the current scene feature vector and the volunteer profile feature vector, and uses a weighted cosine similarity algorithm to determine the priority list of candidate volunteers. The task access request is pushed to one or more volunteers with the highest scheduling scores, and a video call connection is established after the volunteer responds.
[0010] Furthermore, when the cloud platform screens registered volunteers and schedules one or more volunteers to assist with tasks based on their identity information and service record tags, it also includes: Volunteers who are currently in a task execution state or in an unservice state are removed from the priority list, and a scheduling score is calculated for the candidate volunteers. The scheduling score is obtained by weighting and summing the similarity score, idle state weight, historical response time and service score according to a preset ratio. Push the task access request to one or more volunteers with the highest scheduling score, and establish a video call connection after the volunteer responds. If the first batch of pushes does not receive a response, the system will automatically switch to the next batch of candidate volunteers for scheduling within the preset timeout period, until the assistance task is connected or fails due to timeout.
[0011] Furthermore, when the volunteer terminal accesses the video stream of the smart glasses via video call and provides navigation or obstacle prompts to the user through voice commands or audio feedback, it includes: The volunteer terminal establishes an end-to-end encrypted video call channel via the WebRTC protocol, and the video call channel supports multi-bitrate adaptive encoding. The voice interaction module in the volunteer terminal supports speech-to-text and real-time audio synthesis, allowing volunteers to generate voice prompts through text input or directly broadcast navigation instructions to users via voice calls. The audio feedback information is broadcast on the smart glasses via a bone conduction speaker or a directional voice module.
[0012] Furthermore, the evaluation module records and analyzes interaction data in real time during the assistance process, obtaining the success rate of prompts, response latency, and user feedback. When generating a volunteer service score based on the success rate of prompts, response latency, and user feedback, the module includes: During the assistance process, the evaluation module collects interactive data in real time, including voice command timestamps, user location change trajectories, task completion markers, and user's immediate feedback information, including user active voice confirmation, head movement detection feedback, or touch confirmation input. The success rate of the prompt is calculated based on a preset matching rule, which includes: if the user completes the corresponding action within a preset time window after the volunteer issues a navigation or obstacle instruction, it is determined to be a successful prompt, and the success rate of the prompt is the ratio of the number of successful prompts to the total number of prompts; The response delay was calculated based on the time difference between the time the assistance request was sent and the time of the volunteer's first voice response, and the average value of multiple responses was regressed. User feedback is generated based on subjective evaluation scores after interaction, and the user feedback score range is standardized for score normalization. The evaluation module calculates the volunteer service score by weighting and summing the success rate of the prompts, response latency, and user feedback.
[0013] Furthermore, when the evaluation module generates volunteer points based on the volunteer service score, it includes: The evaluation module combines the volunteer service score with task characteristics to quantify the score. The task characteristics include the task urgency level, assistance duration, environmental complexity label, and task completion rate. The evaluation module calls the integral generation function model, which has multiple adjustment factors, including: task complexity adjustment factor α, service score adjustment factor β, and historical service stability factor γ. Volunteer points = α × β × γ × basic points unit × task duration × completion coefficient; Once the volunteer points are generated, they are recorded in the volunteer database and updated synchronously with their historical accumulated points.
[0014] Compared with existing technologies, the advantages of this invention are as follows: By constructing an intelligent video call assistance system that collaboratively integrates smart glasses, a cloud platform, volunteer terminals, and an evaluation module, it overcomes the limitations of existing remote assistance solutions, which are characterized by "passive transmission, manual response, and ambiguous evaluation." The smart glasses collect environmental video, audio, and location information in real time, and the cloud platform performs scene recognition, enabling automatic judgment and proactive assistance triggering in unstructured, dynamically changing, or high-risk environments. This effectively reduces the user's operational burden and improves response timeliness. The cloud platform performs multi-dimensional screening and intelligent matching based on volunteer profile characteristics, ensuring a high degree of compatibility in terms of language, experience, and scenario suitability. Volunteers directly access the user's first-person perspective through video calls and receive navigation guidance via voice or audio prompts, achieving efficient and intuitive remote assistance. The evaluation module quantitatively analyzes the success rate of prompts, response latency, and user feedback, generating a volunteer service score and awarding points accordingly. This constructs an objective incentive system, promoting the continuous improvement of volunteer service quality and the healthy operation of the platform.
[0015] On the other hand, this application also provides a video call assistance method based on smart glasses, applied to the aforementioned video call assistance system based on smart glasses, comprising: The system collects environmental information in real time, including environmental video, environmental audio, and location information, and uploads it to the cloud platform via a wireless communication module. Based on the received environmental information, scene recognition is performed to determine whether the current environment is a complex environment. The complex environment includes unstructured environment, dynamically changing environment, or high-risk environment. When a complex environment is identified, a remote assistance request is triggered. Registered volunteers are screened, and a multi-dimensional feature vector is constructed based on the volunteer's identity information, service record, idle status and tag features that match the user. The similarity is calculated with the user's current scene features to generate a scheduling score, and one or more volunteers are scheduled to join the assistance task. The video call stream of the smart glasses is accessed through the volunteer terminal, and navigation or obstacle prompts are provided to the user through voice commands or audio feedback. The audio prompts are played through bone conduction speakers or directional voice modules. During the assistance process, interactive data is collected in real time, including voice commands, user action feedback and task completion status, and a volunteer service score is generated based on the success rate of prompts, response delay and user feedback. After the assistance task is completed, volunteer points are generated based on the service score and task characteristics. These volunteer points are used for volunteer level assessment, task priority ranking, or incentive reward system.
[0016] Furthermore, when performing scene recognition based on the received environmental information to determine whether the current environment is complex, this includes: Image target detection and semantic segmentation are performed on the uploaded environmental video to extract ground paving, obstacle density and path clarity, and convolutional neural network is used to determine whether it is an unstructured environment; The system performs audio event recognition on ambient audio, extracts typical sound source features, and determines whether the environment is dynamically changing based on the background noise level. The location information is compared with a preset high-risk area database to determine whether the user is in a high-risk area; If the confidence level of the identification result of any type of complex environment exceeds the preset threshold, it is determined to be a complex environment and a remote assistance request is triggered.
[0017] It is understandable that the above-mentioned video call assistance methods and systems based on smart glasses have the same beneficial effects, and will not be elaborated further here. Attached Figure Description
[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a structural block diagram of a video call assistance system based on smart glasses provided in an embodiment of the present invention; Figure 2 A flowchart of a video call assistance method based on smart glasses provided in an embodiment of the present invention. Detailed Implementation
[0019] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] In traditional smart glasses remote assistance systems, environmental perception information transmission relies on manual judgment and triggering mechanisms, failing to automatically identify complex scene features. This leads to increased delays in assistance requests and higher false trigger rates. The lack of dynamic task scheduling capabilities and the reliance on simple tag-based volunteer matching without integrating historical service records and scene adaptability analysis result in insufficient accuracy and real-time performance of navigation instructions. Furthermore, the lack of interactive data quantification and evaluation methods during the assistance process prevents a closed-loop feedback loop for service quality, impacting the effectiveness of volunteer incentive mechanisms.
[0021] For example, when visually impaired users enter unstructured environments, the ground paving and obstacle distribution exhibit highly dynamic changes, and existing systems cannot automatically determine environmental complexity through video stream semantic segmentation and audio event recognition. Users must manually trigger assistance requests, increasing their operational burden and causing response delays. When the system randomly assigns volunteers, it does not consider their historical service ratings and language ability tags, which may lead to a mismatch between navigation instructions and user actions. During assistance, voice command timestamps and user location trajectories are not recorded in real time, making it impossible to generate an objective service rating after the task is completed. Volunteer level assessment relies on subjective evaluations, reducing the fairness of task priority ranking.
[0022] If these issues are not addressed, delays in identifying complex environments will increase user security risks; the lack of a dynamic scheduling mechanism will lead to low utilization of volunteer resources; and the failure rate of task execution in critical scenarios will increase. The lack of a quantitative evaluation system will weaken volunteer participation, exacerbate service quality fluctuations, and create a bottleneck in system scalability. In the long run, user trust and system reliability will continue to decline.
[0023] For this, please refer to Figure 1As shown, this application proposes a video call assistance system based on smart glasses, comprising: smart glasses for real-time collection of environmental information, including environmental video, environmental audio, and location information, and uploading the environmental information to a cloud platform via a wireless communication module; the cloud platform for scene recognition based on the received environmental information, determining whether the current environment is complex, including unstructured environments, dynamically changing environments, or high-risk environments, and triggering a remote assistance request when a complex scene is determined; the cloud platform is also configured to screen registered volunteers, and dispatch one or more volunteers to access the system based on the volunteer's identity information and service record tag characteristics. The system includes: a volunteer terminal (installed on smart glasses) that accesses the smart glasses' video stream via video call and provides navigation or obstacle prompts to the user through voice commands or audio feedback; and an evaluation module (located on a cloud platform) that records and analyzes interaction data in real time during the assistance process, obtaining the success rate of prompts, response latency, and user feedback, and generating a volunteer service score based on these metrics. Upon completion of the assistance task, the evaluation module generates volunteer points based on the volunteer service score, which are then incorporated into the volunteer incentive system for volunteer level assessment, task priority ranking, or awarding of honors.
[0024] Smart glasses, in particular, are wearable devices with environmental awareness capabilities. They are used to collect environmental video, audio, and location information in real time through cameras, microphones, and positioning modules. Specifically, they can be implemented using hardware modules that integrate multimodal sensors, such as those equipped with wide-angle cameras, directional microphones, and GPS chips. Their function is to acquire dynamic data about the user's environment in real time, providing basic input for subsequent scene recognition.
[0025] The cloud platform refers to a data processing and task scheduling center deployed on a remote server. It is used for scene recognition based on environmental video, audio, and location information to determine the existence of unstructured, dynamically changing, or high-risk environments. Specifically, it can be implemented using object detection algorithms, audio event classification models, and geofencing technology. Its function is to automatically identify complex scenes and trigger assistance requests, reducing the burden of manual operation on the user.
[0026] Volunteer selection involves calculating the matching degree based on volunteer identity information and tag features in historical service records. This can be achieved using multi-dimensional feature vector construction and similarity matching algorithms, such as weighted cosine similarity calculation based on identity authentication data, service ratings, and scene tags. Its purpose is to improve the adaptability of volunteers to the current scenario and optimize assistance efficiency.
[0027] The volunteer terminal refers to the interactive interface integrated into smart glasses, used to provide real-time assistance through video streaming and voice command feedback. Specifically, it can establish an encrypted video channel using the WebRTC protocol and integrate speech-to-text and audio synthesis modules. Its function is to achieve low-latency two-way interaction and ensure user privacy through bone conduction or directional voice modules.
[0028] The evaluation module, embedded in the cloud platform, is a data analysis unit used to record the success rate of prompts, response latency, and user feedback in real time during the interaction process. Specifically, it can be implemented using timestamp comparison, action matching rules, and standardized subjective evaluation algorithms. Its purpose is to quantify the quality of volunteer services and provide an objective basis for points-based incentives.
[0029] Among them, volunteer points are quantitative indicators generated based on service ratings and task characteristics. They are used for level assessment and task priority ranking in the incentive system. Specifically, they can be calculated using a weighted function model combined with task complexity, service duration, and completion coefficients. Their purpose is to enhance volunteer participation through a transparent evaluation mechanism.
[0030] This application utilizes a collaborative architecture between smart glasses and a cloud platform to achieve automatic identification and dynamic volunteer scheduling in complex scenarios. Simultaneously, it combines real-time interactive data collection and quantitative evaluation mechanisms to construct a closed-loop incentive system for volunteer service quality. This solution addresses the problems of existing technologies, such as reliance on manual environmental identification, low volunteer matching accuracy, and a lack of objective evaluation of assistance effectiveness, thereby improving system response efficiency and service scalability.
[0031] The working process and principle of this application are as follows: smart glasses collect environmental information in real time, including environmental video, environmental audio, and location information, and upload this information to a cloud platform via a wireless communication module. After receiving the environmental information, the cloud platform performs scene recognition to determine whether the current environment is complex. Complex environments include unstructured environments, dynamically changing environments, or high-risk environments. When a complex scene is determined, the cloud platform triggers a remote assistance request.
[0032] The cloud platform screens registered volunteers and, based on their identity information and service record tags, assigns one or more volunteers to assist with tasks. Volunteer terminals are installed on smart glasses, accessing the smart glasses' video stream via video calls and providing navigation or obstacle prompts to users through voice commands or audio feedback.
[0033] The evaluation module, hosted on a cloud platform, records and analyzes interaction data in real time during the assistance process, obtaining information such as success rate of prompts, response latency, and user feedback. Based on this data, a volunteer service score is generated. Upon completion of the assistance task, the evaluation module generates volunteer points based on the volunteer service score, which are then incorporated into the volunteer incentive system for volunteer level assessment, task priority ranking, or awarding of honors.
[0034] This working principle enables automatic identification of complex environments, intelligent scheduling of volunteers, and quantitative assessment of service quality, thereby improving the efficiency and quality of remote assistance.
[0035] As a preferred embodiment, the solution of this application is specifically implemented as follows: The smart glasses are equipped with a high-definition camera, microphone array, and GPS module to collect environmental video, audio, and location information in real time. The collected data is uploaded to a cloud server in real time via a 5G network.
[0036] The cloud platform employs a deep learning model for scene recognition. It performs object detection and semantic segmentation on the video stream, extracting features such as ground condition and obstacle density. It also performs sound source identification and noise analysis on the audio. The platform compares the results with a pre-set database of high-risk areas, incorporating GPS positioning information. When the confidence level of any complex environment recognition result exceeds a preset threshold, a remote assistance request is triggered.
[0037] The cloud platform constructs a volunteer profile database, including multi-dimensional features such as identity information, historical service records, ratings, and time-of-day preferences. Machine learning algorithms are used to calculate the match between volunteers and the current scenario, generating a priority list. The system then sends task requests to the available volunteers with the highest match scores.
[0038] The volunteer terminal establishes an encrypted video call with the smart glasses via the WebRTC protocol. Volunteers can generate navigation instructions via voice or text input, which are then played to the user through the bone conduction speakers of the smart glasses.
[0039] The evaluation module records interaction data in real time, including command timestamps, user location changes, and task completion markers. The system calculates the success rate of prompts and response latency, and collects user feedback. A service score is calculated based on preset weights. After the task is completed, the evaluation module converts the service score into points rewards, taking into account factors such as task difficulty and duration.
[0040] Through the above-described solution, this application achieves automatic identification and timely response in complex environments, avoiding the inconvenience of users manually triggering assistance requests. The intelligent scheduling mechanism improves the matching degree between volunteers and tasks, enhancing the accuracy and real-time performance of navigation instructions. The quantitative evaluation system provides an objective basis for the quality of volunteer services, promoting the effective operation of the incentive mechanism. It also improves the reliability and user experience of the remote assistance system, providing visually impaired users with safer and more efficient navigation support.
[0041] In some of the solutions described above in this application, smart glasses trigger complex scene recognition and assistance requests from the cloud platform through environmental information collection. However, during the environmental information collection process, sudden abnormal events may not be detected in time, resulting in a delay in triggering assistance requests and increasing user security risks.
[0042] This application further proposes that the smart glasses include an environmental anomaly detection unit, which is used to detect whether there are abnormal events during the process of collecting environmental information. Abnormal events include sudden sound, changes in light or continuous occlusion, and triggers a remote assistance request first when an abnormal event is detected.
[0043] The environmental anomaly detection unit includes a multimodal sensor fusion module, which integrates a voiceprint recognition sensor, a light intensity detection sensor, and an image occlusion detection algorithm. Sudden sound detection is achieved through real-time sound pressure level change monitoring; an anomaly marker is triggered when the ambient audio decibel level rises by more than 30 dB within 0.5 seconds. Light change detection uses a photosensitive element to capture illuminance fluctuations at a 100 Hz sampling frequency; an anomaly signal is generated when the illuminance change rate exceeds 2000 lux per second. Continuous occlusion detection calculates the occlusion area percentage in the image using a camera inter-frame difference algorithm; an abnormal state is determined when the occlusion area exceeds 60% in five consecutive frames.
[0044] Specifically, the environmental anomaly detection unit deploys an edge computing unit locally on the smart glasses. This unit synchronously processes data from three sensors at a 10ms cycle. When any sensor triggers an anomaly, the edge computing unit immediately sends an interrupt signal to the wireless communication module, forcibly activating a high-priority data transmission channel. At this time, the environmental video stream resolution automatically switches to 720p / 60fps mode, the audio sampling rate is increased to 48kHz, and a metadata data packet with a red alert marker is sent to the cloud platform. When the cloud platform receives a request with a red alert marker, it skips the regular scene recognition process and directly activates the emergency response protocol of the volunteer scheduling module. This protocol shortens the volunteer matching timeout threshold from the usual 15 seconds to 3 seconds and prioritizes calling volunteers with emergency event handling tags. At the data transmission level, after an anomaly is triggered, the smart glasses activate a dual-link redundant transmission mechanism, simultaneously uploading data through cellular networks and Wi-Fi channels to ensure real-time video stream transmission is maintained as long as at least one link is available.
[0045] As a preferred embodiment, the solution of this application is implemented as follows: The smart glasses include an environmental anomaly detection unit. The environmental anomaly detection unit detects the presence of abnormal events during the collection of environmental information. Abnormal events include sudden noises, changes in light, or continuous obstruction. Upon detecting an abnormal event, a remote assistance request is triggered first.
[0046] Specifically, the environmental anomaly detection unit can be implemented in the following ways: Sudden Sound Detection: Utilizing the built-in microphone array of the smart glasses, ambient audio is collected in real time. Short-time Fourier transform is used to perform time-frequency analysis on the audio signal, extracting the frequency, amplitude, and duration characteristics of the sound. Sound intensity and duration thresholds are set; when a sound event exceeding the threshold is detected, it is determined to be a sudden sound.
[0047] Light change detection: Using the light sensor of the smart glasses, the ambient light intensity is continuously monitored. The rate of change of light intensity over a short period of time is calculated, and when the rate of change exceeds a preset threshold, it is determined to be a sudden change in light intensity.
[0048] Continuous occlusion detection: Video frame sequences are captured using the front-facing camera of the smart glasses. Image differencing and moving target detection are performed on consecutive frames. When a large area of continuous occlusion is detected, it is determined to be a continuous occlusion event.
[0049] Multimodal fusion: This method integrates sound, light, and visual information to construct an anomaly detection model. Machine learning algorithms, such as support vector machines or random forests, are used to classify the fused features, improving the accuracy of anomaly detection.
[0050] Priority Trigger: When any type of abnormal event is detected, the environmental anomaly detection unit immediately sends a high-priority remote assistance request to the cloud platform, which includes the type of abnormal event, timestamp, and relevant sensor data.
[0051] Through the above technical solutions, this application enables smart glasses to automatically detect and rapidly respond to abnormal environmental events. This improves the system's sensitivity to potential hazards and reduces the user's burden of manually triggering assistance. Furthermore, multimodal information fusion and machine learning algorithms enhance the accuracy and robustness of abnormal event identification. The priority triggering mechanism ensures timely remote assistance upon detecting anomalies, improving user safety in complex or hazardous environments.
[0052] In some of the solutions mentioned above in this application, the cloud platform performs scene recognition based on the received environmental information to determine the complex environment. However, in the actual execution process, there are problems such as the single dimension of environmental feature extraction and insufficient judgment basis, which leads to insufficient accuracy and timeliness of complex environment recognition, and may delay or erroneously trigger remote assistance requests.
[0053] This application further proposes a cloud platform that performs scene recognition based on received environmental information to determine whether the current environment is complex. When a complex scene is determined and a remote assistance request is triggered, the cloud platform performs object detection and semantic segmentation on the environmental video uploaded by the smart glasses, extracts the ground paving status, obstacle density, and path clarity, and determines whether the current environment is unstructured based on a convolutional neural network model; performs audio event recognition on the environmental audio, extracts vehicle horn sounds, mechanical noises, and human noises, and determines whether the environment is dynamically changing based on the background sound pressure level; performs geographical comparison based on location information and an external high-risk area database, and identifies a high-risk environment when the user's current location is in a high-risk area or the user's movement trajectory shows an abnormal change; and triggers a remote assistance request when the confidence level of any type of complex environment recognition result exceeds a threshold.
[0054] The system employs object detection and semantic segmentation to extract ground paving status, obstacle density, and path clarity from environmental videos. Ground paving status is identified using image segmentation algorithms to determine road material and continuity; obstacle density is determined by object detection algorithms to count the number of obstacles per unit area; and path clarity is assessed by edge detection algorithms to analyze path boundary integrity. Audio event recognition utilizes spectral analysis and voiceprint matching to separate characteristic waveforms of vehicle horns, mechanical noise, and human noise from environmental audio. Background sound pressure level is calculated in real-time using a decibel meter to determine dynamic noise levels. Geographic comparison matches the GPS coordinates of the smart glasses with polygonal geofences in a high-risk area database, and anomaly direction changes are identified by calculating the standard deviation of the movement direction using trajectory point sequences.
[0055] Specifically, the convolutional neural network model fuses features of ground paving status, obstacle density, and path clarity using pre-trained weight parameters, outputting a confidence score for unstructured environments. The audio event recognition module extracts acoustic features using Mel-frequency cepstral coefficients and uses a support vector machine classifier to determine the confidence level of dynamically changing environments. The high-risk environment recognition module combines geofence matching results and the rate of change of turning angle in movement trajectories to generate a confidence score for high-risk environments. When any confidence score exceeds a preset threshold, the cloud platform immediately triggers a remote assistance request. For example, when the confidence level of an unstructured environment reaches 0.85, the system will still prioritize triggering assistance requests even if the confidence levels of dynamically changing and high-risk environments do not reach the threshold. Through multi-dimensional environmental feature extraction and independent threshold judgment mechanisms, the system ensures high coverage and fast response speed in complex environment recognition.
[0056] As a preferred embodiment, the solution of this application is specifically implemented as follows: The cloud platform performs object detection and semantic segmentation on environmental videos uploaded by smart glasses. Specifically, the YOLOv5 object detection algorithm is used to process video frames, extracting features such as ground paving status, obstacle density, and path clarity. Further, the DeepLabV3+ semantic segmentation model is used for pixel-level scene classification, obtaining semantic labels for ground, obstacles, and buildings. Based on this, a convolutional neural network model is used to determine whether the current environment is unstructured.
[0057] The system performs audio event recognition on ambient audio. Specifically, the VGGish sound classification model is used to extract features such as vehicle horns, mechanical noise, and human noise. Furthermore, a sound pressure level measurement module calculates the background noise level to determine whether the environment is dynamically changing.
[0058] Geographic comparison is performed based on location information and an external high-risk area database. Specifically, GPS coordinates collected by the smart glasses are matched with preset geofences for high-risk areas. Furthermore, the user's movement trajectory is analyzed, and abnormal changes in trajectory are identified as high-risk environments.
[0059] Finally, a remote assistance request is triggered when the confidence level of any type of complex environment identification result exceeds a preset threshold. For example, the confidence threshold for unstructured environment identification is set to 0.8, the confidence threshold for dynamically changing environment identification is set to 0.7, and the confidence threshold for high-risk environment identification is set to 0.9.
[0060] Through the above technical solution, this application achieves intelligent identification and timely response to complex environments. The cloud platform can perform scene analysis based on multimodal data to accurately determine whether the user is in an unstructured environment, a dynamically changing environment, or a high-risk environment. Therefore, the system can proactively trigger remote assistance requests without requiring manual user intervention, improving the timeliness and accuracy of assistance. This intelligent scene recognition mechanism effectively reduces the burden on visually impaired users and enhances the system's ability to cope with complex situations.
[0061] In some of the solutions mentioned above in this application, the cloud platform relies solely on identity information and service record tags for simple matching when screening volunteers, lacking a comprehensive consideration of the multidimensional capabilities of volunteers. This may result in the scheduling results being insufficiently adapted to the user's current scenario needs, affecting assistance efficiency and user experience.
[0062] This application further proposes a cloud platform for constructing a volunteer profile database. For each registered volunteer, a multi-dimensional feature vector is created, including identity information, historical assistance scene records, service rating, task completion time preferences, language ability, and assistance domain labels. Upon receiving a remote assistance request, the cloud platform obtains the user's current scene feature vector based on environmental information, calculates the similarity between the current scene feature vector and the volunteer profile feature vector, and uses a weighted cosine similarity algorithm to determine a priority list of candidate volunteers. The platform then pushes the task access request to one or more volunteers with the highest scheduling scores and establishes a video call connection after the volunteer responds.
[0063] The multidimensional feature vector includes identity information, historical assistance scenario records, service rating, task completion time preferences, language ability, and assistance domain labels. Each dimension is assigned a different weight, which is dynamically adjusted based on the influence of different features in historical task data on task success rate. The weighted cosine similarity algorithm calculates the cosine of the angle between the current scene feature vector and the volunteer profile feature vector, and combines this with the weights of each dimension to obtain a comprehensive similarity score. A higher score indicates a higher fit. During the priority list generation process, candidate volunteers are arranged in descending order of score, and a threshold is set to filter low-scoring items to ensure that only candidates who meet the fit criteria are retained.
[0064] Specifically, when a user triggers a remote assistance request, the cloud platform extracts environmental video, audio, and location information to generate a scene feature vector. This vector includes elements such as ground paving status, obstacle density, and high-risk area markers. The multi-dimensional feature vector in the volunteer profile database reflects volunteers' experience and capabilities in similar scenarios through historical assistance scenario records and service ratings. Task completion time preferences and language proficiency ensure that volunteers are available and communication is seamless during the user's current time period. In the weighted cosine similarity calculation, service ratings and assistance domain labels have higher weight coefficients than other dimensions, prioritizing matching volunteers with matching professional fields and high service quality. After the candidate volunteer priority list is generated, the system automatically sends access requests to the highest-ranked volunteers. If a candidate does not respond within a preset time, the request is forwarded to the next candidate until the task is completed or timeout occurs.
[0065] As a preferred embodiment, the solution of this application is specifically implemented as follows: A cloud platform constructs a volunteer profile database, creating a multi-dimensional feature vector for each registered volunteer. This multi-dimensional feature vector includes identity information, historical assistance scenario records, service ratings, task completion time preferences, language proficiency, and assistance domain tags. Identity information includes basic attributes such as age, gender, and occupation. Historical assistance scenario records include the type, number, and frequency of completed tasks. Service ratings are derived from user reviews of previous tasks. Task completion time preferences record the volunteer's commonly used service time periods. Language proficiency includes the languages spoken and their level of fluency. Assistance domain tags cover different categories such as visual impairment assistance, emergency rescue, and professional consultation.
[0066] Upon receiving a remote assistance request, the cloud platform obtains the user's current scene feature vector based on environmental information. This feature vector includes dimensions such as geographical location, environmental complexity, and task urgency. The cloud platform then calculates the similarity between the current scene feature vector and the volunteer profile feature vector, using a weighted cosine similarity algorithm to determine a priority list of candidate volunteers. The weighted cosine similarity algorithm assigns different weights to different feature dimensions, such as giving higher weights to language ability and assistance domain labels.
[0067] The system will send task access requests to volunteers ranked high in the scheduling score. Specifically, the system selects the top 5 volunteers with the highest similarity and sends them task access invitations. If a volunteer responds, a video call connection is immediately established. If no response is received within 15 seconds, the system automatically sends invitations to the next batch of 5 candidate volunteers until a connection is successfully established or the 3-minute timeout limit is reached.
[0068] Through the aforementioned technical solutions, this application achieves accurate volunteer matching and efficient task scheduling, thereby improving the response speed and service quality of remote assistance. Specifically, by constructing multi-dimensional feature vectors, the system can comprehensively consider the volunteers' abilities, experience, and preferences, thus selecting the most suitable volunteers for the current scenario. Furthermore, a weighted cosine similarity algorithm is used for matching to ensure the accuracy of the selection results. The mechanism of pushing task requests in batches ensures both rapid response and avoids excessive disruption to volunteers. Therefore, this application not only improves the efficiency of remote assistance but also enhances the enthusiasm of volunteers.
[0069] In some of the solutions described above in this application, after generating a priority list of candidate volunteers based on the similarity between the scene feature vector and the volunteer profile feature vector, directly pushing task requests may cause volunteers to be unable to respond because they are in the process of executing a task or are not in a serviceable period, resulting in scheduling delays or failures and affecting the efficiency of assisting in task execution.
[0070] This application further proposes removing volunteers currently in task execution or in an unservice state from the priority list, and calculating a scheduling score for candidate volunteers. The scheduling score is obtained by weighting and summing similarity score, idle state weight, historical response time, and service score according to a preset ratio. One or more volunteers with the highest scheduling scores are pushed with task access requests, and a video call connection is established after the volunteer responds. If the first batch of pushes does not receive a response, the system automatically switches to the next batch of candidate volunteers for scheduling within a preset timeout period until the task is connected or fails due to timeout.
[0071] Volunteers in the task execution or inactive state are removed through a real-time status polling mechanism to ensure that volunteers in the candidate list have immediate response capabilities. The scheduling score calculation uses a multi-dimensional weighted model: similarity score reflects scenario matching degree, idle state weight is calculated based on volunteer terminal online status and task load, historical response time is based on the average response time of past tasks, and service score is taken from historical evaluation data. The preset timeout period is dynamically adjusted according to the task's urgency level; for example, it is set to 30 seconds for high-risk environment tasks and 60 seconds for regular tasks. Automatic batch switching uses a queue polling mechanism; a countdown starts after each push, and if there is no response within the timeout period, the next batch of candidate volunteers is pushed.
[0072] Specifically, during the scheduling process, after the priority list is generated, unavailable volunteers are first filtered out, retaining candidates with service capabilities. Then, candidates are scored based on multiple factors, with idle status accounting for 30%, similarity score for 40%, historical response time for 20%, and service score for 10%. After scoring, they are arranged in descending order to generate a scheduling queue. The system sends an access request to the volunteer at the front of the queue; if no response is received within 30 seconds, the request is automatically pushed to subsequent volunteers. This process is repeated until the task is accepted or the queue is exhausted. Through status filtering and dynamic scoring mechanisms, volunteers with high response willingness and high service quality are prioritized, while a batch switching mechanism avoids long waiting times caused by single pushes, improving task connection efficiency. For example, in urgent tasks, the system prioritizes scheduling volunteers with scores above 85 and historical response times below 10 seconds, switching batches within a 15-second timeout period to ensure rapid task allocation.
[0073] As a preferred embodiment, the solution of this application is specifically implemented as follows: Volunteers currently in a task execution state or not in a service state are removed from the priority list. A scheduling score is calculated for the candidate volunteers. The scheduling score is obtained by weighting the similarity score, idle state weight, historical response time, and service score according to a preset ratio.
[0074] The task access request is sent to one or more volunteers with the highest scheduling scores. A video call connection is established after the volunteer responds.
[0075] If the first batch of pushes does not receive a response, the system will automatically switch to the next batch of candidate volunteers for scheduling within the preset timeout period, until the assistance task is connected or fails due to timeout.
[0076] Specifically, the cloud platform first removes unavailable volunteers from the priority list, including volunteers currently performing other tasks and those set to temporarily not accept tasks. Then, it calculates a scheduling score for the remaining candidate volunteers. The scheduling score consists of four indicators: similarity score to the current scenario (weight 40%), volunteer's current idle status (weight 20%), historical average response time (weight 20%), and historical service score (weight 20%). These indicators are added together according to preset weights to obtain the final scheduling score.
[0077] The system will send task access requests to the top 3 volunteers with the highest scores in the first batch. If a volunteer responds, a video call connection will be established immediately. If no one responds within 15 seconds, the system will automatically switch to volunteers ranked 4th to 6th for the second batch. This process will continue until a volunteer answers or the request is deemed a timeout failure after 5 rounds of requests (75 seconds).
[0078] Through the above technical solutions, this application achieves intelligent screening and dynamic scheduling of volunteers. A multi-dimensional scoring mechanism selects the most suitable volunteers for the current scenario for delivery. Batch delivery and automatic switching mechanisms improve task connection efficiency. Simultaneously considering volunteers' historical performance and current status ensures service quality. This enhances the response speed and service quality of the remote assistance system.
[0079] In some of the solutions mentioned above in this application, the volunteer terminal accesses the video stream of the smart glasses via video call and provides navigation or obstacle prompts to the user through voice commands or audio feedback. However, during video transmission, there is a problem of screen stuttering or delay when the network fluctuates due to the fixed bit rate. In addition, the voice interaction method is singular and cannot adapt to the operating habits of volunteers in different scenarios. At the same time, the audio feedback may affect the user's reception due to environmental interference.
[0080] This application further proposes that the volunteer terminal establishes an end-to-end encrypted video call channel via the WebRTC protocol, and the video call channel supports multi-bitrate adaptive encoding; the voice interaction module in the volunteer terminal supports speech-to-text and real-time audio synthesis, allowing volunteers to generate voice prompts through text input, or directly broadcast navigation instructions to users through voice calls; audio feedback information is broadcast on the smart glasses through bone conduction speakers or directional voice modules.
[0081] The WebRTC protocol uses DTLS-SRTP encryption to establish an end-to-end communication link, ensuring data security during video stream transmission. Multi-bitrate adaptive encoding dynamically adjusts video resolution and frame rate based on real-time network bandwidth. For example, it switches to 480p resolution and 15fps frame rate when network bandwidth is below 1Mbps, and increases to 720p resolution and 30fps frame rate when bandwidth is above 2Mbps. The voice interaction module integrates a speech recognition engine and a TTS synthesis engine. Volunteers can choose to input voice commands via microphone or manually input text commands. Text commands are converted into audio streams by the TTS engine and transmitted synchronously with the video stream. Bone conduction speakers transmit sound signals by vibrating the cheekbones, avoiding external sound interference with the surrounding environment. The directional voice module uses beamforming technology to concentrate and direct sound waves to the user's ear canal area.
[0082] Specifically, after the video call channel is established via the WebRTC protocol, the cloud platform monitors the network transmission quality in real time and feeds it back to the encoder. The encoder dynamically adjusts the video encoding parameters based on a bandwidth prediction algorithm. For example, it automatically reduces the bitrate to maintain call fluency when a network latency exceeds 200ms. Volunteers can choose between voice input mode or text input mode on the terminal interface. In text input mode, the navigation instructions typed by the volunteer are converted into standardized voice prompts by the TTS engine, such as "There are steps two meters ahead, please turn right." The converted voice data is encapsulated in RTP packets with low latency and transmitted to the smart glasses. After receiving the audio stream, the smart glasses select the audio output mode according to the ambient noise level. In noisy environments, the directional voice module is activated, focusing the voice within the user's hearing range through a preset 40-degree beam angle. In quiet environments, it switches to bone conduction mode to reduce sound leakage. As a result, the stability of video transmission, the flexibility of voice interaction, and the clarity of audio feedback are optimized, solving the problem of transmission delays in assistive commands and difficulties in receiving audio due to environmental interference in complex network environments.
[0083] As a preferred embodiment, the solution of this application is specifically implemented as follows: The volunteer terminal establishes an end-to-end encrypted video call channel via the WebRTC protocol. This channel supports multi-bitrate adaptive encoding, which can dynamically adjust the video quality according to network conditions. For example, when network bandwidth is sufficient, a high-definition video stream of 1080p resolution and 30fps frame rate is used; when network conditions are poor, it automatically reduces to 720p resolution and 15fps frame rate to ensure smooth call quality.
[0084] The voice interaction module in the volunteer terminal supports speech-to-text and real-time audio synthesis. Volunteers can generate voice prompts through text input, and the system will convert the text into natural speech in real time; or they can directly broadcast navigation instructions to users via voice calls. The speech synthesis uses a deep learning model, which can simulate various timbres and tones to adapt to different scenario needs.
[0085] Audio feedback is delivered via bone conduction speakers or a directional voice module on the smart glasses. Bone conduction speakers use skull vibrations to transmit sound, enabling clear information delivery even in noisy environments without affecting the user's perception of external sounds. The directional voice module uses phased array technology to create directional sound waves, forming an audible sound field only at the user's ear, thus avoiding disturbing others.
[0086] Through the above technical solutions, this application achieves high-quality, low-latency video call connections between volunteers and users, ensuring the security and stability of the communication process. The voice interaction module provides flexible command input methods, improving the operational efficiency of volunteers. Bone conduction and directional voice technology ensure the clarity and privacy of audio feedback, enhancing the user's information reception capabilities in complex environments. Therefore, this solution improves the communication quality and user experience of remote assistance, providing visually impaired users with more effective and convenient navigation and obstacle warning services.
[0087] In some of the solutions mentioned above in this application, an evaluation module is proposed to record and analyze interactive data in real time to generate volunteer service scores. However, in the specific implementation process, it is still difficult to accurately quantify the success rate of prompts, response delays and user feedback, and to establish a unified scoring standard, which leads to a lack of objectivity and operability in the evaluation of volunteer service quality.
[0088] This application further proposes an evaluation module that collects interactive data in real time during the assistance process, including voice command timestamps, user location change trajectories, task completion markers, and immediate user feedback information. Feedback information includes user-initiated voice confirmation, head movement detection feedback, or touch confirmation input. The success rate of prompts is calculated based on preset matching rules, which include: if the user completes the corresponding action within a preset time window after the volunteer issues a navigation or obstacle instruction, it is considered a successful prompt; the success rate is the ratio of successful prompts to the total number of prompts. The response latency is calculated based on the time difference between the assistance request sending time and the volunteer's first voice response time, and average regression processing is performed on multiple responses. User feedback is generated based on subjective evaluation scores after interaction, and the user feedback score range is standardized for score normalization. The evaluation module performs a weighted summation of the prompt success rate, response latency, and user feedback to obtain a volunteer service score.
[0089] The voice command timestamps are collected with millisecond-level precision via a time synchronization server, ensuring the temporal correlation between user actions and commands. User location change trajectories are generated using a fusion of GPS positioning data from the smart glasses' built-in inertial navigation module. Task completion markers are automatically set by the cloud platform based on preset task nodes. The preset time windows in the matching rules are dynamically adjusted according to the task type; for example, a 10-second time window corresponds to navigation commands, while a 5-second time window corresponds to obstacle warning commands. The average regression processing uses a sliding window algorithm to calculate a moving average of three consecutive response delays to eliminate random fluctuations. User feedback standardization maps the original scores to the 0-1 range through linear transformation, eliminating differences in rating scales among different users.
[0090] Specifically, during the task execution phase, the evaluation module continuously receives voice command timestamp data uploaded by the smart glasses and matches it with the user's action trigger time. When a volunteer issues the command "Turn left ahead," the system records the command timestamp as T0. If the user responds with a left-turn action via head movement detection within T0+10 seconds, the prompt is considered successful. The response delay calculation module extracts the assistance request sending time T1 and the volunteer's first voice response time T2, calculating the difference between T2 and T1 as the single-time delay data. When the volunteer responds to three commands during the task, the average of the three delay data is taken as the final delay index. After completing the assistance, the user submits a 1-5 star rating through the touch input interface. This rating is standardized and converted to 0.8 points for weighted calculation. The evaluation module sets the prompt success rate weight to 0.5, the response delay weight to 0.3, and the user feedback weight to 0.2, outputting a service score in the range of 0-100 through a weighted summation formula. This score directly reflects the timeliness, accuracy, and user satisfaction of the volunteer's service.
[0091] As a preferred embodiment, the solution of this application is specifically implemented as follows: The evaluation module records and analyzes interaction data in real time during the assistance process. Specifically, the evaluation module collects voice command timestamps, user location change trajectories, task completion markers, and real-time user feedback. User feedback includes user-initiated voice confirmation, head movement detection feedback, or touch confirmation input.
[0092] Furthermore, the evaluation module calculates the success rate of prompts. The success rate is calculated based on preset matching rules. These rules include: if a user completes the corresponding action within a preset time window after a volunteer issues a navigation or obstacle instruction, it is considered a successful prompt. The success rate is derived from the ratio of successful prompts to the total number of prompts.
[0093] The response latency was calculated based on the time difference between the time the assistance request was sent and the time of the volunteer's first voice response. Average regression was performed on multiple responses.
[0094] Therefore, user feedback is generated based on subjective evaluations and scores given after interaction. The score ranges for user feedback are standardized for score normalization.
[0095] Specifically, the evaluation module calculates a volunteer service score by weighting and summing the success rate of prompts, response latency, and user feedback.
[0096] For example, the evaluation module can set a weight of 0.4 for prompt success rate, 0.3 for response latency, and 0.3 for user feedback. After normalizing the prompt success rate, response latency, and user feedback respectively, they are multiplied by their corresponding weights and summed to obtain the final volunteer service score.
[0097] Through the aforementioned technical solution, this application achieves a quantitative assessment of volunteer service quality. The assessment module collects interactive data in real time, calculates objective indicators such as prompt success rate and response latency, and combines this with user subjective feedback to generate a comprehensive volunteer service score. This assessment method improves the objectivity and comprehensiveness of the score, providing a reliable basis for subsequent volunteer incentives and task allocation. Simultaneously, the real-time assessment mechanism helps to promptly identify and improve service issues, enhancing overall service quality. Furthermore, the standardized scoring system allows for horizontal comparison of performance among different volunteers, promoting healthy competition and motivating volunteers to continuously improve their service levels.
[0098] In some of the solutions mentioned above in this application, the evaluation module generates volunteer points based on volunteer service scores, but the point calculation method is too simplistic and cannot be dynamically adjusted based on the urgency of the task, the complexity of the environment, and the stability of the volunteer's historical service. As a result, the quantitative results of the points cannot accurately reflect the actual contribution of the volunteers, which affects the fairness and effectiveness of the incentive mechanism.
[0099] This application further proposes that when the evaluation module generates volunteer points based on volunteer service scores, it combines volunteer service scores with task characteristics for point quantification. Task characteristics include task urgency level, assistance duration, environmental complexity label, and task completion degree. The point generation function model is invoked, which has multiple adjustment factors, including a task complexity adjustment factor α, a service score adjustment factor β, and a historical service stability factor γ. Volunteer points = α × β × γ × basic point unit × task time length × completion degree coefficient. After the volunteer points are generated, they are recorded in the volunteer database and updated synchronously with their historical accumulated points.
[0100] The task complexity adjustment factor α ranges from 0.5 to 2.0, and is dynamically adjusted according to the task urgency level and environmental complexity label. Tasks in high-risk areas or unstructured environments correspond to higher α values. The service score adjustment factor β is positively correlated with the service score. For every 10% increase in the service score, the β value increases linearly by 0.1. The historical service stability factor γ is calculated by taking the variance of the volunteer's service scores for the most recent ten tasks. When the variance is below the threshold, γ is set to 1.2, and when the variance exceeds the threshold, γ is set to 0.8. The basic unit of integration is set to a fixed value. The task time is converted to minutes proportionally. The completion coefficient is set according to the proportion of task objectives achieved. When the objectives are fully achieved, the coefficient is 1.0, and when they are partially achieved, the coefficient is 0.6 to 0.9.
[0101] Specifically, after the assistance task is completed, the evaluation module extracts the urgency level label, assistance duration data, environmental complexity classification results, and task completion index from the task features, and obtains the volunteer service score corresponding to the task. The integral generation function model calls the preset adjustment factor calculation rules, binding the task complexity adjustment factor α with the environmental complexity label. High-risk environment tasks automatically trigger α=1.8, and unstructured environment tasks trigger α=1.5. The service score adjustment factor β is calculated by interpolation based on the percentile level of the current service score. For example, when the service score is 85 points, β=1.3. The historical service stability factor γ is calculated by querying historical service records in the volunteer database and calculating the standard deviation of the scores of the most recent ten tasks. When the standard deviation is less than 5, γ=1.2. The basic unit of integral is set to 10 points / minute. When the task lasts for 15 minutes and the completion coefficient is 0.9, the integral calculation result is 1.8×1.3×1.2×10×15×0.9=454.8 points. The score is updated to the volunteer database in real time and added to the historical score to generate a new total score, which is used to update the volunteer level and task scheduling priority.
[0102] As a preferred embodiment, the specific implementation of this application's solution is as follows: After the volunteer completes the nighttime navigation task in a high-risk area, the evaluation module calls the integral generation function model for calculation. The task urgency level is marked as Level 2, the assistance duration is 18 minutes, the environmental complexity label is set to a dynamic obstacle-dense scene, and the task completion coefficient is determined by the system to be 0.95. In the integral generation function model, the task complexity adjustment factor α dynamically takes a value of 0.8 based on the environmental complexity label, the service score adjustment factor β takes a value of 1.2 based on the 92 service scores obtained by the volunteer, and the historical service stability factor γ is calculated to be 0.9 based on the standard deviation of the volunteer's past 30 service scores. The basic point unit is set to 10 points / minute, and the final volunteer point = 0.8 × 1.2 × 0.9 × 10 × 18 × 0.95, with a calculation result of 147.7 points. This point is recorded in the volunteer database and added to its historical accumulated points of 2350 points for updating. The updated total point of 2497.7 points is synchronized to the task scheduling system to improve the priority ranking of the volunteer's subsequent tasks.
[0103] Through the above technical solution, this application realizes an integral quantification mechanism based on multi-dimensional task characteristics and dynamic adjustment factors, effectively solving the problems of single scoring dimensions and lack of quantification of task differences in traditional volunteer incentive systems. By introducing a dynamic coupling calculation model of task complexity and service quality, the integral generation process can objectively reflect the true value of service contributions in different scenarios, improving the transparency and fairness of the volunteer incentive system. Furthermore, by combining historical service stability factors to dynamically correct the integrals, the excessive influence of occasional high-scoring tasks on long-term service evaluation is avoided, establishing a long-term incentive mechanism that balances service quality stability and task difficulty.
[0104] In some of the schemes mentioned above in this application, when the evaluation module generates volunteer points based on volunteer service scores, it only considers the single factor of service score and does not combine task characteristics and the stability of volunteers' historical service. As a result, the point quantification cannot accurately reflect the differences in task complexity and the long-term service quality of volunteers, affecting the fairness and effectiveness of the incentive mechanism.
[0105] This application further proposes that when the evaluation module generates volunteer points based on volunteer service scores, it combines volunteer service scores with task characteristics for point quantification. Task characteristics include task urgency level, assistance duration, environmental complexity label, and task completion degree. The point generation function model is invoked, which has multiple adjustment factors, including a task complexity adjustment factor α, a service score adjustment factor β, and a historical service stability factor γ. Volunteer points are calculated by α×β×γ×basic point unit×task time length×completion degree coefficient, and the generated volunteer points are recorded in the volunteer database and updated synchronously with historical accumulated points.
[0106] Among them, the task complexity adjustment factor α is dynamically adjusted according to the environmental complexity label and the task urgency level. The environmental complexity label is classified into levels based on the identification results of unstructured environment, dynamically changing environment or high-risk environment. The task urgency level is determined by the matching degree between the user's location and the high-risk area database. The service score adjustment factor β is linearly or nonlinearly mapped according to the normalized results of volunteer service scores. The historical service stability factor γ is obtained by calculating the inverse of the variance of the volunteer's historical service scores. The smaller the variance, the larger the value of γ. The basic unit of integration is preset by the system. The task time length is recorded in minutes to record the assistance duration. The completion coefficient is calculated comprehensively based on the task completion mark and the user's final feedback.
[0107] Specifically, when the evaluation module generates volunteer scores, it first obtains the environmental complexity labels and high-risk area matching results from the scene recognition module. Combined with the task duration and user feedback on task completion, it determines the task feature parameters. The task complexity adjustment factor α is assigned a tiered value based on the level of the environmental complexity label; for example, a high-risk environment corresponds to α=1.5, a dynamically changing environment to α=1.2, and an unstructured environment to α=1.0. Simultaneously, the α value increases by 0.1 for each level of task urgency. The service score adjustment factor β maps volunteer service scores to the 0.8-1.5 range; for example, a score of 90 corresponds to β=1.5, and a score of 60 corresponds to β=0.8. The historical service stability factor γ is calculated by taking the reciprocal of the standard deviation of the past 10 service scores and normalizing it to the 0.9-1.1 range. The basic scoring unit is set to 10 points / minute, the task duration is the actual assistance time, and the completion coefficient is calculated based on whether the user reached the target location and the feedback score, with a value of 0.6-1.0. The integral generation function substitutes the above parameters into the formula for calculation. For example, in a task with α=1.2, β=1.3, γ=1.05, task duration of 30 minutes, and completion coefficient of 0.9, the volunteer's points = 1.2 × 1.3 × 1.05 × 10 × 30 × 0.9 = 442.26 points. This score, accumulated with historical scores, is used to update the volunteer's level and task scheduling priority, thereby achieving fair incentives based on multiple dimensions.
[0108] As a preferred embodiment, the solution of this application is implemented as follows: After the user wears smart glasses, the built-in camera of the glasses captures environmental video at a rate of 30 frames per second, the microphone acquires environmental audio at a sampling rate of 16kHz, and the GPS module continuously updates the location coordinates. The above data is transmitted to the cloud server via a 5G network. The recognition engine deployed in the cloud analyzes the video stream frame by frame and uses the ResNet-50 model to perform semantic segmentation of the ground paving status. When an unpaved road is detected and the density of obstacles exceeds 3 per square meter, it is determined to be an unstructured environment. The audio processing unit extracts acoustic features through Mel-frequency cepstral coefficients. If mechanical noise exceeding 75 decibels for more than 10 seconds is detected, it is determined to be a dynamically changing environment. The location coordinates are compared in real time with a high-risk area database. If the user enters the fenced area of a construction site, high-risk environment recognition is triggered. When the confidence level of any environment recognition exceeds 85%, the cloud generates an assistance request containing scene type, location coordinates, and user ID.
[0109] The volunteer dispatch module filters candidates with experience in building navigation and a service rating higher than 4.5 from the profile database. It calculates the similarity between the current scene features and the feature vectors of the user's needs, prioritizing volunteers with records of assisting at construction sites. The dispatch system pushes task notifications to the three volunteers with the highest matching scores. If the first volunteer does not respond within 15 seconds, it automatically switches to the next candidate. After establishing a connection, the volunteer terminal receives a 720P video stream via H.265 encoding, marks obstacle locations on the interactive interface, and generates voice commands. The command text is converted into audio by a TTS engine and transmitted to the bone conduction component of the smart glasses at a bitrate of 48kbps.
[0110] During the assistance process, the system records the time difference between issuing each navigation command and the user's action. A valid prompt is recorded when the user completes the step avoidance maneuver within 5 seconds. At the end of the task, the user triggers a satisfaction rating by triple-clicking a mirror image. The system generates a volunteer service rating based on an 85% prompt success rate, an average response delay of 2.3 seconds, and a user rating of 4.8. The points calculation module increases the base points by 1.5 times based on the high-risk environment label. Combined with the 120-second assistance duration, a final score of 180 points is generated and updated in the volunteer's service profile.
[0111] Through the above technical solutions, this application realizes an intelligent assistance triggering mechanism in complex environments, effectively reducing the user's operational burden through multimodal data fusion and identification; the volunteer scheduling strategy based on multidimensional feature matching improves the scene adaptability and increases the assistance response accuracy by 30%; the real-time quantitative service evaluation system provides objective data support for the volunteer incentive mechanism and promotes continuous optimization of service quality; the points generation model is dynamically adjusted in combination with task characteristics to enhance the enthusiasm of volunteers and the sustainability of the system.
[0112] The above embodiments construct an intelligent video call assistance system that collaboratively integrates smart glasses, a cloud platform, volunteer terminals, and an evaluation module, overcoming the limitations of existing remote assistance solutions characterized by "passive transmission, manual response, and ambiguous evaluation." The smart glasses collect environmental video, audio, and location information in real time, and the cloud platform performs scene recognition, enabling automatic judgment and proactive assistance triggering in unstructured, dynamically changing, or high-risk environments. This effectively reduces user workload and improves response timeliness. The cloud platform performs multi-dimensional screening and intelligent matching based on volunteer profile characteristics, ensuring a high degree of compatibility in terms of language, experience, and scenario suitability. Volunteers directly access the user's first-person perspective through video calls and receive navigation guidance via voice or audio prompts, achieving efficient and intuitive remote assistance. The evaluation module quantitatively analyzes the success rate of prompts, response latency, and user feedback, generating a volunteer service score and awarding points accordingly. This constructs an objective incentive system, promoting continuous improvement in volunteer service quality and the healthy operation of the platform.
[0113] In another preferred embodiment based on the above embodiments, see [reference] Figure 2 As shown, this embodiment provides a video call assistance method based on smart glasses, applied to the aforementioned video call assistance system based on smart glasses, including: S100: Collects environmental information in real time, including environmental video, environmental audio and location information, and uploads it to the cloud platform through the wireless communication module; S200: Based on the received environmental information, perform scene recognition to determine whether the current environment is a complex environment. Complex environments include unstructured environments, dynamically changing environments, or high-risk environments. When a complex environment is identified, a remote assistance request is triggered. S300: Screen registered volunteers, construct multi-dimensional feature vectors based on volunteers' identity information, service records, idle status and tag features that match users, calculate similarity with the user's current scene features, generate scheduling scores, and schedule one or more volunteers to access assistance tasks. S400: Accesses the video call stream of smart glasses through a volunteer terminal and provides navigation or obstacle prompts to users through voice commands or audio feedback. Audio prompts are played through bone conduction speakers or directional voice modules. S500: During the assistance process, it collects interactive data in real time, including voice commands, user action feedback and task completion status, and generates a volunteer service score based on the success rate of prompts, response delay and user feedback. S600: After the assistance task is completed, volunteer points are generated based on the service score and task characteristics. Volunteer points are used for volunteer level assessment, task priority ranking, or incentive reward system.
[0114] Furthermore, when performing scene recognition based on the received environmental information and determining whether the current environment is a complex environment, the process includes: performing image target detection and semantic segmentation on the uploaded environmental video, extracting ground paving, obstacle density and path clarity, and determining whether it is an unstructured environment through a convolutional neural network; The system performs audio event recognition on ambient audio, extracts typical sound source features, and determines whether the environment is dynamically changing based on the background noise level. The location information is compared with a pre-defined database of high-risk areas to determine whether the user is in a high-risk area. If the confidence level of the identification result of any type of complex environment exceeds the preset threshold, it is determined to be a complex environment and a remote assistance request is triggered.
[0115] In some of the solutions mentioned above in this application, a scheme based on environmental information to identify the scene and trigger a remote assistance request is proposed. However, when judging complex environments, there may be cases of misjudgment or omission of single-dimensional data, which may lead to inaccurate or delayed triggering of assistance requests, affecting user safety and user experience.
[0116] This application further proposes to perform image target detection and semantic segmentation on uploaded environmental videos, extract ground paving, obstacle density, and path clarity, and use convolutional neural networks to determine whether it is an unstructured environment; to perform audio event recognition on environmental audio, extract typical sound source features, and determine whether it is a dynamically changing environment based on background noise level; to compare location information with a preset high-risk area database to determine whether the user is in a high-risk area; and if the confidence level of the recognition result of any type of complex environment exceeds a preset threshold, it is determined to be a complex environment and a remote assistance request is triggered.
[0117] In this system, image object detection uses the YOLOv5 model to detect obstacles in video frames in real time, semantic segmentation uses the DeepLabv3+ model to perform pixel-level classification of ground paving textures, and obstacle density calculation is based on the coverage ratio of detection boxes per unit area. Audio event recognition extracts Mel-spectrum features using a pre-trained VGGish model and combines it with a dynamic time warping algorithm to match vehicle horn sound feature templates. Geographic location comparison uses GPS coordinates and high-risk area polygons in GeoJSON format to perform spatial relationship calculations, and movement trajectory anomaly detection is based on the deviation between the Kalman filter predicted path and the actual trajectory. The confidence threshold is dynamically adjusted according to the environment type: 85% for unstructured environments, 75% for dynamically changing environments, and 90% for high-risk environments.
[0118] Specifically, when the video stream captured by the smart glasses is input into the cloud platform, the video processing module simultaneously performs object detection and semantic segmentation. The ground paving status is quantified by the proportion of flat tiles and gravel roads in the segmentation results. Areas with more than three detection boxes per square meter of obstacles are marked as high-density areas. The audio analysis module extracts MFCC features from a continuous 5-second audio clip. When more than three vehicle horns are detected and the background noise exceeds 65 decibels, it is determined to be a dynamically changing environment. Location information is compared with the high-risk area database every 2 seconds. A high-risk judgment is triggered when the user is within the construction area defined by the electronic fence for 10 consecutive seconds, or when the movement trajectory shows three directional changes of more than 15 degrees within 30 seconds. The three types of recognition results are input into the corresponding convolutional neural network classifiers. When the output probability of any classifier exceeds a preset threshold, a remote assistance request command is immediately generated and pushed to the scheduling module through a message queue. This multi-dimensional cross-validation mechanism effectively avoids misjudgment by a single sensor, while balancing the recognition sensitivity and specificity of different environment types through dynamic threshold settings.
[0119] As a preferred embodiment, the solution of this application is implemented as follows: During the user's movement while wearing smart glasses, environmental video is uploaded to the cloud platform at a rate of 30 frames per second via the glasses' built-in camera. The video stream is input into a pre-trained ResNet-50 convolutional neural network for semantic segmentation, segmenting unpaved areas, temporary obstacle piles, and discontinuous path boundaries. When the obstacle density exceeds 3 per square meter and the path clarity is below 0.7, it is determined to be an unstructured environment. Simultaneously, environmental audio is extracted using MFCC features and input into an LSTM network to identify continuous mechanical noise exceeding 85dB and sudden vehicle horn sounds. When the background sound pressure level fluctuates by more than 20dB within 10 seconds, it is determined to be a dynamically changing environment. Location information is obtained through a GPS module and matched with coordinates in a pre-loaded high-risk area GIS database. If the user's current location overlaps with a 500-meter buffer zone of construction sites or transportation hubs in the database, it is marked as a high-risk area. When the confidence score of any of the above three types of detections exceeds the 0.85 threshold, the cloud platform immediately sends an assistance request instruction to the volunteer dispatch module.
[0120] Through the above technical solutions, this application effectively solves the problems of reliance on manual operation for environmental recognition and low accuracy in complex scene judgment in existing technologies. By employing a collaborative judgment mechanism based on multi-dimensional environmental data analysis and preset thresholds, it achieves automated identification of unstructured road surfaces, sudden noise interference, and high-risk areas, reducing the operational burden of users manually triggering requests. The combination of semantic segmentation and voiceprint recognition for dual verification improves the reliability of complex environment judgment, ensuring that assistance requests are triggered only in necessary scenarios and avoiding the waste of volunteer resources on ineffective tasks. The integration of geofencing technology and dynamic noise monitoring enables rapid emergency response when users approach dangerous areas or encounter sudden environmental changes, enhancing the safety level of visually impaired users in outdoor scenarios.
[0121] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0122] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0123] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0124] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A video call assistance system based on smart glasses, characterized in that, include: Smart glasses are used to collect environmental information in real time, including environmental video, environmental audio and location information, and upload the environmental information to a cloud platform through a wireless communication module. The cloud platform performs scene recognition based on the received environmental information to determine whether the current environment is complex. The complex environment includes unstructured environment, dynamically changing environment or high-risk environment. When the scene is determined to be complex, a remote assistance request is triggered. The cloud platform is also configured to screen registered volunteers and dispatch one or more volunteers to assist in the task based on the volunteers' identity information and the tag features of their service records. A volunteer terminal is installed on the smart glasses. The volunteer terminal accesses the video stream of the smart glasses via video call and provides navigation or obstacle prompts to the user through voice commands or audio feedback. An evaluation module is set up on the cloud platform. During the assistance process, the evaluation module records and analyzes the interaction data in real time, obtains the prompt success rate, response delay and user feedback during the assistance process, and generates a volunteer service score based on the prompt success rate, response delay and user feedback. After the assistance task is completed, the evaluation module generates volunteer points based on the volunteer service score, and incorporates the volunteer points into the volunteer incentive system for volunteer level assessment, task priority ranking, or honor distribution.
2. The video call assistance system based on smart glasses according to claim 1, characterized in that, The smart glasses further include an environmental anomaly detection unit, which is used to detect whether there are abnormal events during the process of collecting environmental information. The abnormal events include sudden sounds, changes in light, or continuous obstruction, and when the abnormal event is detected, a remote assistance request is triggered first.
3. The video call assistance system based on smart glasses according to claim 1, characterized in that, The cloud platform performs scene recognition based on the received environmental information to determine whether the current environment is complex. When a complex scene is determined and a remote assistance request is triggered, the following is included: The cloud platform performs target detection and semantic segmentation on the environmental video uploaded by the smart glasses, extracts the ground paving status, obstacle density and path clarity, and determines whether the current environment is unstructured based on a convolutional neural network model. The ambient audio is subjected to audio event recognition to extract vehicle horn sounds, mechanical noises and human noises, and the background sound pressure level is used to determine whether it is the dynamically changing environment. Geographic comparison is performed between location information and an external high-risk area database, and the user's current location is in a high-risk area or the user's movement trajectory shows an abnormal change in direction, which is identified as a high-risk environment. When the confidence level of any type of complex environment identification result exceeds the threshold, a remote assistance request is triggered.
4. The video call assistance system based on smart glasses according to claim 1, characterized in that, The cloud platform screens registered volunteers and, based on their identity information and service record tags, schedules one or more volunteers to assist with tasks, including: The cloud platform constructs a volunteer profile database, creating a multi-dimensional feature vector for each registered volunteer that includes identity information, historical assistance scenario records, service rating, task completion time preferences, language ability, and assistance domain labels. Upon receiving a remote assistance request, the cloud platform obtains the user's current scene feature vector based on the environmental information, calculates the similarity between the current scene feature vector and the volunteer profile feature vector, and uses a weighted cosine similarity algorithm to determine the priority list of candidate volunteers. The task access request is pushed to one or more volunteers with the highest scheduling scores, and a video call connection is established after the volunteer responds.
5. The video call assistance system based on smart glasses according to claim 4, characterized in that, The cloud platform screens registered volunteers and, based on their identity information and service record tags, schedules one or more volunteers to assist with tasks. This also includes: Volunteers who are currently in a task execution state or in an unservice state are removed from the priority list, and a scheduling score is calculated for the candidate volunteers. The scheduling score is obtained by weighting and summing the similarity score, idle state weight, historical response time and service score according to a preset ratio. Push the task access request to one or more volunteers with the highest scheduling score, and establish a video call connection after the volunteer responds. If the first batch of pushes does not receive a response, the system will automatically switch to the next batch of candidate volunteers for scheduling within the preset timeout period, until the assistance task is connected or fails due to timeout.
6. The video call assistance system based on smart glasses according to claim 1, characterized in that, When the volunteer terminal accesses the video stream of the smart glasses via video call and provides navigation or obstacle prompts to the user through voice commands or audio feedback, it includes: The volunteer terminal establishes an end-to-end encrypted video call channel via the WebRTC protocol, and the video call channel supports multi-bitrate adaptive encoding. The voice interaction module in the volunteer terminal supports speech-to-text and real-time audio synthesis, allowing volunteers to generate voice prompts through text input or directly broadcast navigation instructions to users via voice calls. The audio feedback information is broadcast on the smart glasses via a bone conduction speaker or a directional voice module.
7. The video call assistance system based on smart glasses according to claim 1, characterized in that, The evaluation module records and analyzes interaction data in real time during the assistance process, obtaining the success rate of prompts, response latency, and user feedback. When generating a volunteer service score based on the success rate of prompts, response latency, and user feedback, the module includes: During the assistance process, the evaluation module collects interactive data in real time, including voice command timestamps, user location change trajectories, task completion markers, and user's immediate feedback information, including user active voice confirmation, head movement detection feedback, or touch confirmation input. The success rate of the prompt is calculated based on a preset matching rule, which includes: if the user completes the corresponding action within a preset time window after the volunteer issues a navigation or obstacle instruction, it is determined to be a successful prompt, and the success rate of the prompt is the ratio of the number of successful prompts to the total number of prompts; The response delay was calculated based on the time difference between the time the assistance request was sent and the time of the volunteer's first voice response, and the average value of multiple responses was regressed. User feedback is generated based on subjective evaluation scores after interaction, and the user feedback score range is standardized for score normalization. The evaluation module calculates the volunteer service score by weighting and summing the success rate of the prompts, response latency, and user feedback.
8. The video call assistance system based on smart glasses according to claim 7, characterized in that, When the evaluation module generates volunteer points based on the volunteer service score, it includes: The evaluation module combines the volunteer service score with task characteristics to quantify the score. The task characteristics include the task urgency level, assistance duration, environmental complexity label, and task completion rate. The evaluation module calls the integral generation function model, which has multiple adjustment factors, including: task complexity adjustment factor α, service score adjustment factor β, and historical service stability factor γ. Volunteer points = α × β × γ × basic points unit × task duration × completion coefficient; Once the volunteer points are generated, they are recorded in the volunteer database and updated synchronously with their historical accumulated points.
9. A video call assistance method based on smart glasses, applied to the video call assistance system based on smart glasses as described in any one of claims 1-8, characterized in that, include: The system collects environmental information in real time, including environmental video, environmental audio, and location information, and uploads it to the cloud platform via a wireless communication module. Scene recognition is performed based on the received environmental information to determine whether the current environment is a complex environment, including unstructured environment, dynamically changing environment or high-risk environment; When the environment is identified as complex, a remote assistance request is triggered. Registered volunteers are screened, and a multi-dimensional feature vector is constructed based on the volunteer's identity information, service record, idle status and tag features that match the user. The similarity is calculated with the user's current scene features to generate a scheduling score, and one or more volunteers are scheduled to join the assistance task. The video call stream of the smart glasses is accessed through the volunteer terminal, and navigation or obstacle prompts are provided to the user through voice commands or audio feedback. The audio prompts are played through bone conduction speakers or directional voice modules. During the assistance process, interactive data is collected in real time, including voice commands, user action feedback and task completion status, and a volunteer service score is generated based on the success rate of prompts, response delay and user feedback. After the assistance task is completed, volunteer points are generated based on the service score and task characteristics. These volunteer points are used for volunteer level assessment, task priority ranking, or incentive reward system.
10. The video call assistance method based on smart glasses according to claim 9, characterized in that, When performing scene recognition based on received environmental information to determine whether the current environment is complex, the following includes: Image target detection and semantic segmentation are performed on the uploaded environmental video to extract ground paving, obstacle density and path clarity, and convolutional neural network is used to determine whether it is an unstructured environment; The system performs audio event recognition on ambient audio, extracts typical sound source features, and determines whether the environment is dynamically changing based on the background noise level. The location information is compared with a preset high-risk area database to determine whether the user is in a high-risk area; If the confidence level of the identification result of any type of complex environment exceeds the preset threshold, it is determined to be a complex environment and a remote assistance request is triggered.