Remote driving control method and system and medium

By using large language models and multimodal fusion technology for scene understanding and risk assessment, the problem of insufficient risk identification in remote driving control systems under complex traffic environments is solved, achieving higher risk identification accuracy and lower cognitive burden.

CN121963028APending Publication Date: 2026-05-01ZHIJI AUTOMOTIVE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHIJI AUTOMOTIVE TECH CO LTD
Filing Date
2025-12-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing remote driving control systems struggle to understand scenarios and identify risks in complex traffic environments, resulting in a heavy cognitive burden, insufficient generalization ability, and decreased reliability in extreme environments.

Method used

By employing a large language model for scene understanding and combining it with multimodal fusion technology, risk levels are quantified through a multi-dimensional risk assessment algorithm, and targeted early warning information is generated, thereby reducing cognitive burden.

Benefits of technology

It improves the accuracy and generalization ability of remote driving control systems in complex traffic environments, reduces the cognitive burden of remote safety operators, and avoids alarm fatigue caused by excessive warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963028A_ABST
    Figure CN121963028A_ABST
Patent Text Reader

Abstract

The invention relates to a remote driving control method and system and a medium, and relates to the field of intelligent driving, and the method comprises the steps: obtaining multi-channel video stream data of a remote vehicle; scene understanding is carried out on the multi-channel video stream data based on a visual language model, and traffic participants in a traffic scene and behavior characteristics of the traffic participants are identified; according to the scene understanding result, performing risk level quantification on the identified potential risk scene through a multi-dimensional risk assessment algorithm, and generating a structured risk assessment result including risk description, severity and time urgency; and based on the structured risk assessment result, generating targeted early warning information and presenting the targeted early warning information to a remote security officer through a human-computer interaction interface. According to the invention, by introducing the scene understanding ability of the large language model and the multi-modal fusion technology, the risk identification accuracy and generalization ability of the remote driving control system in a complex traffic environment can be improved, and the cognitive burden of a remote safety officer can be reduced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

A remote driving control method, system and medium Technical Field

[0001] This invention relates to the field of intelligent driving, and in particular to a remote driving control method, system and medium. Background Technology

[0002] Remote driving technology refers to a system where an operator remotely controls a vehicle via a communication network. Its core components include onboard sensors, a communication transmission system, a remote control center, and a user interface. As a supplementary technology to autonomous driving, remote driving enables safe vehicle operation through remote human intervention in complex scenarios where autonomous driving systems cannot handle them. Current visual recognition technologies in remote driving systems mainly include rule-based traditional computer vision algorithms, deep learning-based object detection and segmentation technologies, and fusion perception systems combining multi-sensor data. These technologies perform well in recognizing structured scenes and predefined objects, but they have limitations in scene understanding and risk prediction in complex traffic environments.

[0003] Existing remote driving control systems face numerous technical challenges. Remote safety operators must focus on monitoring multiple video streams for extended periods, resulting in a heavy cognitive load, increased attention deficit, and fatigue, impacting risk identification capabilities. Traditional computer vision methods focus on object detection rather than scene understanding, making it difficult to comprehend the behavioral intentions of traffic participants or their interactions with the environment. Existing algorithms achieve high recognition rates for predefined risk scenarios, but their generalization ability is insufficient when facing novel or rare risk situations in open worlds. Furthermore, most of these systems employ deterministic judgment criteria, lacking the ability to reasonably handle uncertainty and quantify risk levels, leading to a significant decrease in reliability under complex environments such as extreme weather and varying lighting conditions. Summary of the Invention

[0004] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a remote driving control method, system and medium. By introducing the scene understanding capability of large language models and multimodal fusion technology, the accuracy and generalization capability of risk identification of remote driving control systems in complex traffic environments can be improved, while reducing the cognitive burden of remote safety drivers.

[0005] To achieve the above objectives, the present invention adopts the following technical solution.

[0006] In a first aspect, the present invention provides a remote driving control method, which adopts the following technical solution: acquiring multi-channel video stream data of a remote vehicle; performing scene understanding on the multi-channel video stream data based on a visual language model to identify traffic participants and their behavioral characteristics in the traffic scene; quantifying the risk level of the identified potential risk scenes through a multi-dimensional risk assessment algorithm based on the scene understanding results, generating a structured risk assessment result including risk description, severity, and time urgency; and generating targeted early warning information based on the structured risk assessment result and presenting it to a remote safety officer through a human-computer interaction interface.

[0007] Furthermore, in the aforementioned remote driving method, the scene understanding based on the visual language model for the multi-channel video stream data includes: preprocessing the multi-channel video stream data, including denoising, enhancement, and key area detection; performing basic risk screening using a lightweight visual language model; transmitting the preprocessed video data and basic screening results to the cloud; and performing deep scene understanding analysis using a complete visual language model deployed in the cloud.

[0008] Furthermore, in the aforementioned remote driving control method, the deep scene understanding analysis includes: basic-level analysis, identifying road elements and their attributes; relation-level analysis, analyzing the spatial relationships and interaction patterns between elements; and semantic-level analysis, understanding the overall semantics of the scene and the meaning of potential risks.

[0009] Furthermore, in the aforementioned remote driving method, the multi-dimensional risk assessment algorithm includes: calculating collision probability parameters; quantifying time urgency parameters; assessing the severity of consequences parameters; determining the difficulty of hazard avoidance parameters; and calculating and assessing certainty parameters.

[0010] Furthermore, in the aforementioned remote driving method, the risk level quantification includes: calculating a comprehensive risk score based on the five parameters; classifying the risk level into four levels—reminder, warning, emergency, and danger—based on the comprehensive risk score; and configuring a corresponding warning trigger threshold for each risk level.

[0011] Furthermore, in the aforementioned remote driving control method, the generation of targeted early warning information includes: selecting the early warning presentation method according to the risk level, wherein the reminder level uses visual markers on the interface, the warning level uses visual markers and voice prompts, the emergency level uses visual markers, voice alarms and interface flashing, and the danger level uses multimodal all-round early warning; displaying the risk location and type in the remote operation interface through augmented reality technology; and adjusting the early warning intensity according to the current cognitive load of the remote safety officer.

[0012] Furthermore, the aforementioned remote driving control method also includes: monitoring the response behavior of remote safety personnel to warning information; analyzing the accuracy and effectiveness of warnings; adjusting risk assessment parameters based on the analysis results; and storing new risk scenarios in a scenario library for system optimization.

[0013] Furthermore, in the aforementioned remote driving control method, monitoring the remote safety officer's response to warning information includes: recording the safety officer's acceptance, ignoring, or intervention in response to warnings; statistically analyzing the response time for warnings at different risk levels; assessing the safety officer's cognitive load changes over different time periods; and establishing a database of safety officer operational behavior patterns.

[0014] Furthermore, in the aforementioned remote driving control method, the identification of traffic participants and their behavioral characteristics in a traffic scene includes: identifying the position and movement status of vehicles, pedestrians, and non-motorized vehicles; analyzing the attention direction and behavioral intention of each traffic participant; predicting the possible movement trajectory of each participant; constructing a relationship graph among traffic participants; and calculating the similarity between the current scene and historical high-risk scenes based on historical driving data.

[0015] Secondly, the present invention provides a remote driving control system, which adopts the following technical solution: a data acquisition module for acquiring multi-channel video stream data of a remote vehicle; a scene understanding module for performing scene understanding on the multi-channel video stream data based on a visual language model, and identifying traffic participants and their behavioral characteristics in the traffic scene; a risk assessment module for quantifying the risk level of the identified potential risk scenes based on the scene understanding results using a multi-dimensional risk assessment algorithm, and generating a structured risk assessment result including risk description, severity, and time urgency; and an early warning presentation module for generating targeted early warning information based on the structured risk assessment result and presenting it to the remote safety officer through a human-computer interaction interface.

[0016] Thirdly, the present invention provides a readable storage medium, which adopts the following technical solution: a readable storage medium storing computer instructions, wherein the computer instructions, when executed by a processor, implement the remote driving method as described in any one of the first aspects above.

[0017] In summary, compared with existing technologies, this invention offers at least one of the following beneficial technical effects: The remote driving control method of this invention achieves a deep understanding of traffic scenes through a visual language model, enabling the identification of traffic participants' behavioral characteristics and intentions, and exhibiting stronger scene understanding capabilities compared to traditional object detection methods. The multi-dimensional risk assessment algorithm quantifies risks from multiple dimensions, including collision probability, time urgency, severity of consequences, difficulty of avoidance, and assessment certainty, generating structured risk assessment results that can more accurately determine risk levels. The targeted early warning information presentation method based on risk levels can provide appropriate intensity of warnings under different risk conditions, avoiding alarm fatigue caused by excessive warnings, while ensuring that high-risk scenarios receive timely attention, thereby reducing the cognitive burden on remote safety operators and improving the accuracy and efficiency of risk identification and response. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 is a flowchart of a specific embodiment of a remote driving control method of the present invention.

[0020] Figure 2 is a flowchart of another specific embodiment of a remote driving control method of the present invention.

[0021] Figure 3 is a flowchart of another specific embodiment of a remote driving control method of the present invention.

[0022] Figure 4 is a flowchart of another specific embodiment of a remote driving control method of the present invention.

[0023] Figure 5 is a flowchart of a specific application scenario of the remote driving control method of the present invention.

[0024] Figure 6 is a structural schematic diagram of a specific embodiment of the remote driving control system of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Furthermore, it should be understood that the specific embodiments described herein are only for illustration and explanation of this application and are not intended to limit this application.

[0026] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments of this application. Furthermore, the descriptions of each embodiment in the following embodiments have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0027] The method steps described in this embodiment of the invention can be executed in the order described in the specific implementation, or the execution order of each step can be adjusted according to actual needs, provided that the technical problem can be solved. These are not listed one by one here.

[0028] The present invention will be further described in detail below with reference to the accompanying drawings.

[0029] Referring to Figure 1, this embodiment of the invention provides a remote driving control method 100, which includes multiple steps to achieve intelligent monitoring and risk warning of a remote vehicle. In some embodiments, the remote driving control method 100 begins at step 102, acquiring multi-channel video stream data of the remote vehicle. The multi-channel video stream data is acquired at 1080p resolution and 30fps frame rate, providing high-quality visual input for subsequent scene understanding and risk assessment.

[0030] The remote driving method 100 continues with step 104, performing scene understanding on multi-channel video stream data based on a visual language model. The visual language model possesses powerful semantic understanding capabilities, enabling it to handle complex traffic scenes and understand the relationships between elements within the scene. In some implementations, the visual language model employs a dedicated model with a parameter scale of 13B, optimized for driving scenarios.

[0031] As shown in Figure 1, the remote driving method 100 continues to execute step 106, identifying traffic participants and their behavioral characteristics in the traffic scene. Traffic participants include road elements such as vehicles, pedestrians, and non-motorized vehicles, and their behavioral characteristics cover multiple dimensions such as location information, motion state, attention direction, and behavioral intention. The visual language model can understand the complex behavioral patterns of traffic participants, such as the hesitant posture of a pedestrian crossing the road or the abnormal driving trajectory of the vehicle in front.

[0032] The remote driving method 100 then proceeds to step 108, quantifying the risk level of the identified potential risk scenarios using a multi-dimensional risk assessment algorithm. The multi-dimensional risk assessment algorithm comprehensively considers parameters across five dimensions: collision probability, time urgency, severity of consequences, difficulty of avoidance, and assessment certainty. In some implementations, the algorithm employs a Bayesian network to quantify uncertainty and adaptively adjusts the assessment threshold based on the driving environment, vehicle speed, and safety operator status.

[0033] Referring again to Figure 1, remote driving method 100 executes step 110 to generate a structured risk assessment result that includes a risk description, severity, and time urgency. The structured risk assessment result categorizes risk levels into four levels: alert, warning, emergency, and danger. Each level corresponds to a different risk score range and handling recommendations. The risk description is presented in natural language, facilitating understanding and decision-making by the remote safety officer.

[0034] The remote driving method 100 continues with step 112, generating targeted early warning information based on the structured risk assessment results. The targeted early warning information is presented in an appropriate manner according to the risk level, including different intensities of warning methods such as visual markers on the interface, voice prompts, interface flashing, and multimodal all-around warnings. In some implementations, the early warning information combines augmented reality technology to visually display the location and type of risk.

[0035] Finally, remote driving method 100 reaches step 114, presenting warning information to the remote safety operator through a human-machine interface. The human-machine interface uses a three-screen professional display configuration: the central screen displays the main driving view, while the left and right screens display auxiliary views and augmented reality risk warning interfaces, respectively. The presentation of the warning information adaptively adjusts according to the remote safety operator's current cognitive load.

[0036] Based on the remote driving control method 100 described in this embodiment of the invention, a deep understanding of traffic scenes is achieved through a visual language model. This allows for the identification of the behavioral characteristics and intentions of traffic participants, demonstrating stronger scene understanding capabilities compared to traditional object detection methods. A multi-dimensional risk assessment algorithm quantifies risks from multiple dimensions, including collision probability, time urgency, severity of consequences, difficulty of avoidance, and assessment certainty, generating structured risk assessment results that more accurately determine risk levels. The targeted early warning information presentation method based on risk levels provides appropriate intensity warnings under different risk conditions, avoiding alarm fatigue caused by excessive warnings, while ensuring timely attention to high-risk scenarios. This reduces the cognitive burden on remote safety operators and improves the accuracy and efficiency of risk identification and response.

[0037] Referring to Figure 2, this embodiment of the invention also provides a remote driving control method 200 for scene understanding based on a visual language model. First, starting from step 202, preprocessing is performed on multiple video stream data. The preprocessing process includes multiple steps such as noise reduction, enhancement, and key region detection. In some embodiments, the preprocessing employs an adaptive sampling mechanism based on scene complexity, using higher frame rates and resolutions to collect data from high-risk areas. Multi-level caching and preloading strategies optimize the continuity of the video stream under network fluctuation conditions, ensuring the stability and integrity of data transmission.

[0038] As shown in Figure 2, the remote driving method 200 continues with step 204, performing basic risk screening using a lightweight visual language model. The lightweight visual language model employs a 0.5B parameter scale model architecture and is deployed on the in-vehicle terminal to achieve real-time risk interception with a response time of less than 50ms. Dynamic enhancement technology based on an attention mechanism improves the visual quality of potential risk areas, enabling the lightweight model to accurately identify obvious risks and activate emergency warning mechanisms.

[0039] The remote driving method 200 then proceeds to step 206, transmitting the preprocessed video data and basic screening results to the cloud. The cloud-edge collaborative architecture performs initial risk screening by deploying a lightweight visual language model on the vehicle, transmitting only critical data to the cloud for in-depth analysis, significantly reducing network transmission latency. In some implementations, the transmission process utilizes a 5G communication module to support dual-carrier redundant links, ensuring communication reliability.

[0040] Referring to Figure 2, the remote driving method 200 continues with step 208, performing deep scene understanding analysis through a complete visual language model deployed in the cloud. The complete visual language model employs a dedicated model with 13B parameters, optimized for driving scenarios. During the pre-training phase, the model incorporates 5 million detailed traffic scene images to achieve domain knowledge injection. The traffic rule knowledge injection mechanism enhances the visual language model's understanding of traffic regulations and safety standards.

[0041] As shown in Figure 2, deep scene understanding analysis includes a three-level analysis process. In step 210 of the remote driving method 200, basic-level analysis is performed to identify road elements and their attributes. This basic-level analysis identifies traffic elements such as vehicles, pedestrians, and signs on the road, and extracts basic attribute information such as the position, size, and color of each element. In some implementations, the basic-level analysis supports multimodal input adaptation that simultaneously processes RGB images, depth maps, and optical flow information.

[0042] The remote driving method 200 moves to step 212 to perform a relational analysis of the spatial relationships and interaction patterns between elements. This relational analysis constructs a relationship graph among traffic participants, analyzing the spatial positional relationships, intersections of movement trajectories, and potential interaction patterns between elements. The multi-scale visual language model application method simultaneously performs frame-level, short-time series, and long-time series analysis in the temporal dimension, and combines global scene understanding with local detail analysis in the spatial dimension.

[0043] Referring to Figure 2, the remote driving method 200 reaches step 214, where semantic-level analysis is performed to understand the overall semantics of the scene and the implications of potential risks. Semantic-level analysis involves multi-layered semantic analysis, from object recognition and behavior understanding to intent inference, to achieve a deep understanding of the traffic scene. In some implementations, semantic-level analysis can understand subtle clues such as the hesitant posture of pedestrians crossing the road and judge the abnormal driving trajectory of vehicles in front, and perform causal reasoning based on pre-trained world knowledge to predict possible evolutions and potential risks of the scene.

[0044] In some implementations, multi-scale visual language model analysis methods progressively deepen semantic understanding from basic road element recognition to complex risk semantics. The traffic scene-specific visual language model inference framework employs a two-stage inference strategy: first, rapid risk screening is performed, followed by in-depth analysis of potential risk areas. An attention-guided computational resource allocation mechanism allocates more computational power to high-risk areas, while a caching mechanism reuses historical analysis results to reduce redundant computation, achieving efficient scene understanding processing.

[0045] Referring to Figure 3, this embodiment of the invention also provides a remote driving control method 300 for implementing a multi-dimensional risk assessment algorithm and generating targeted early warnings. The remote driving control method 300 begins at step 302 by calculating collision probability parameters. These parameters are quantified by analyzing the motion trajectories, velocity vectors, and predicted intersection points of traffic participants, and are expressed as a percentage from 0% to 100% indicating the probability of a potential collision. In some implementations, the collision probability calculation combines a physical motion model and a traffic participant behavior prediction model, considering factors such as vehicle braking distance and pedestrian reaction time.

[0046] As shown in Figure 3, the remote driving control method 300 proceeds to step 304, quantifying the time urgency parameter. The time urgency parameter quantifies the time window from risk identification to potential collision occurrence with second-level precision, providing the remote safety operator with a decision-making time assessment. The time urgency calculation considers the current vehicle speed, the spatial distance to the obstacle, and the expected trajectory change rate. In some implementations, the time urgency parameter is corrected in conjunction with network transmission latency to ensure the accuracy of the warning time window.

[0047] The remote driving method 300 continues with step 306, assessing the severity parameters of the consequences and the difficulty parameters of hazard avoidance. The severity parameters quantify the damage level of a potential accident using a 1-10 scale, comprehensively considering factors such as collision speed, participant type, and collision angle. The difficulty parameters also use a 1-10 scale to assess the complexity of performing hazard avoidance maneuvers, considering environmental factors such as road conditions, traffic density, and weather conditions. In some implementations, the assessment of the severity parameters of the consequences and the difficulty parameters of hazard avoidance is trained and optimized based on a large amount of historical accident data and driving scenario data.

[0048] Referring to Figure 3, the remote driving method 300 continues with step 308, calculating the evaluation deterministic parameters. The evaluation deterministic parameters represent the confidence level of the risk assessment result as a percentage from 0-100%, reflecting the confidence level of the visual language model in understanding the current scene. The evaluation deterministic calculation considers multiple factors such as image quality, scene complexity, and the consistency of model output. The Bayesian network uncertainty quantification method precisely expresses the confidence level of the risk assessment, describing the uncertainty range of the assessment result through a probability distribution.

[0049] As shown in Figure 3, the remote driving method 300 then executes step 310, calculating a comprehensive risk score based on five parameters. The comprehensive risk score employs a weighted fusion algorithm, comprehensively calculating parameters across five dimensions: collision probability, time urgency, severity of consequences, difficulty of avoidance, and assessment certainty. In some implementations, the weighting coefficients are adaptively adjusted according to different driving scenarios; urban road scenarios emphasize time urgency, while highway scenarios emphasize the severity of consequences. A multi-step predictive inference mechanism simulates various possible paths the scenario may take and their risk probabilities, providing a forward-looking assessment basis for the comprehensive risk score.

[0050] The remote driving control method 300 reaches decision point 312 and determines whether the risk level is dangerous. Risk level quantification divides the comprehensive risk score into four levels: alert, warning, emergency, and dangerous. Each level corresponds to a specific score range and warning trigger threshold. The alert level corresponds to a comprehensive risk score of 0-25 points, the warning level to 26-50 points, the emergency level to 51-75 points, and the dangerous level to 76-100 points. A dynamic threshold adjustment mechanism adaptively adjusts the warning trigger conditions based on the driving environment, vehicle speed, and safety operator status, lowering the trigger threshold in complex traffic environments to improve warning sensitivity.

[0051] Referring to Figure 3, when the risk level is determined to be hazardous, the remote driving method 300 executes step 314, employing a multimodal all-around early warning approach. This multimodal all-around early warning combines various warning signals such as visual markers, voice alarms, interface flashing, and tactile feedback to ensure that the remote safety operator can promptly perceive high-risk situations. In some implementations, the hazardous level warning triggers an automatic pre-braking mechanism, providing the remote safety operator with more reaction time.

[0052] When the risk level is determined to be non-hazardous, the remote control method 300 executes step 316, selecting the appropriate warning method based on the alert, warning, or emergency level. The alert level uses visual markers on the interface, employing color changes or icons to alert the remote safety operator to potential risks. The warning level uses a combination of visual markers and voice prompts, with the voice prompts describing the risk type and location in concise and clear language. The emergency level uses multiple warning methods, including visual markers, voice alarms, and interface flashing, with the interface flashing using high-contrast colors to attract the safety operator's attention.

[0053] As shown in Figure 3, the targeted early warning information generation process incorporates augmented reality technology to display the risk location and type on the remote operation interface. Augmented reality visualization technology overlays the risk area onto the real-time video screen as a semi-transparent colored overlay, with red indicating high-risk areas and yellow indicating medium-risk areas. In some implementations, the overall intersection risk heat map generation technology provides an intuitive display of risk distribution in complex intersection scenarios, with a green route overlay displaying the recommended safest travel path.

[0054] In some implementations, the warning intensity adjustment mechanism adaptively optimizes based on the remote safety officer's current cognitive load. Cognitive load assessment is quantified by monitoring indicators such as the safety officer's attention distribution, response time, and operation frequency. When a high cognitive load is detected, the warning intensity is automatically increased, using more prominent visual cues and more frequent voice reminders. When the cognitive load is low, the warning intensity is appropriately reduced to avoid excessively interfering with the safety officer's normal monitoring work.

[0055] Furthermore, referring to Figure 4, this embodiment of the invention also provides a remote driving control method 400 for identifying traffic participants and analyzing their behavioral characteristics. The remote driving control method 400 first executes step 402 to identify the position and motion state of vehicles, pedestrians, and non-motorized vehicles. Position identification uses a visual language model to analyze multiple video streams in real time, extracting the precise coordinate information of each traffic participant in three-dimensional space. Motion state identification includes the calculation of dynamic parameters such as velocity vectors, acceleration changes, and motion direction. In some embodiments, motion state identification combines optical flow analysis and temporal frame difference technology to achieve sensitive detection of minute motion changes.

[0056] The remote driving method 400 then proceeds to step 404, analyzing the attention direction and behavioral intentions of each traffic participant. Attention direction analysis is determined by recognizing visual cues such as pedestrian head orientation, vehicle turn signal status, and driver gaze direction. Behavioral intention analysis infers from changes in the posture of traffic participants, movement trajectory patterns, and environmental interactions. In some implementations, behavioral intention analysis employs a sequence-based deep learning model, capable of recognizing subtle behavioral features such as pedestrians' preparatory movements before crossing the road and slight deviations by vehicles before changing lanes.

[0057] The remote driving method 400 moves to step 406 to predict the possible movement trajectories of each participant. The trajectory prediction algorithm combines the current motion state, historical behavior patterns, and environmental constraints to generate multiple possible future movement paths. Each predicted trajectory includes a probability weight, reflecting the likelihood of that trajectory being realized. In some implementations, trajectory prediction uses a Monte Carlo method for random sampling, generating a diverse set of trajectories covering both normal and abnormal behaviors. A physical constraint model ensures that the predicted trajectories conform to kinematic laws and traffic rule restrictions.

[0058] Referring to Figure 4, the method for constructing a relationship graph among traffic participants establishes connections by analyzing the spatial relationships, intersections of movement trajectories, and potential interaction patterns of each participant. The relationship graph adopts a directed graph structure, where nodes represent traffic participants, edges represent mutual influence relationships, and edge weights reflect the intensity of influence. In some implementations, a dynamic update mechanism for the relationship graph adjusts the attributes of nodes and edges in real time according to changes in the scenario, ensuring that the relationship graph accurately reflects the current traffic situation.

[0059] The technique for calculating the similarity between the current scenario and historical high-risk scenarios based on historical driving data employs a multi-dimensional feature matching algorithm. The feature vector includes dimensions such as the distribution of traffic participant types, spatial layout patterns, motion state characteristics, and environmental condition parameters. Similarity calculation uses a weighted combination of cosine similarity and Euclidean distance to generate a similarity score within the range of 0-1. In some implementations, the historical high-risk scenario database contains over 100,000 labeled scenarios, covering risk cases across various road types, weather conditions, and traffic situations.

[0060] Referring again to Figure 4, the remote driving method 400 continues with step 408, monitoring the remote safety operator's response to the warning information. Response behavior monitoring is achieved by recording multi-dimensional information such as the safety operator's operation timestamps, mouse click locations, keyboard input, and eye-tracking data. In some implementations, response behavior is categorized into three basic types: accepting the warning and performing the suggested operation, ignoring the warning and continuing the original operation, and actively intervening in vehicle control. Operation behavior records use millisecond-level precision timestamps to ensure the accuracy of behavior sequence analysis.

[0061] The remote control method 400 then proceeds to step 410, analyzing the accuracy and effectiveness of the early warning. Early warning accuracy is assessed by comparing the predicted early warning results with actual events, calculating indicators such as true positive rate, false positive rate, and missed detection rate. Early warning effectiveness is quantified by analyzing indicators such as safety officer response time, operation success rate, and accident avoidance rate. In some implementations, a sliding time window mechanism is used to evaluate the early warning effect, dynamically analyzing recent early warning performance to identify trends in early warning performance.

[0062] The remote driving method 400 reaches decision point 412 and determines whether a new risk scenario has been detected. New risk scenario identification is based on a similarity comparison between the scenario feature vector and a historical scenario database. When the similarity is below a preset threshold and a risk event actually occurs, it is determined to be a new risk scenario. In some implementations, the determination of new risk scenarios combines expert annotation and automatic identification to ensure the accuracy and completeness of the identification results.

[0063] When a new risk scenario is discovered, the remote driving method 400 executes step 414, storing the new risk scenario in a scenario library for optimization. The scenario storage includes complete information such as video data, feature vectors, risk tags, and handling results. The scenario library adopts a hierarchical storage architecture, with frequently accessed typical scenarios stored in a cache and infrequently accessed rare scenarios stored in a large-capacity storage device. In some implementations, the scenario library incremental update mechanism periodically performs cluster analysis on newly added scenarios to identify new risk patterns and update the risk identification algorithm.

[0064] As shown in Figure 4, when no new risk scenarios are detected, the remote driving method 400 executes step 416 to adjust the risk assessment parameters based on the analysis results. Parameter adjustment includes optimization across multiple dimensions, such as risk score weighting coefficients, warning trigger thresholds, and similarity matching thresholds. The adjustment strategy is based on statistical results of warning accuracy, using a gradient descent algorithm to find the direction for parameter optimization. In some implementations, parameter adjustment employs A / B testing to compare the warning effects of different parameter configurations and select the parameter combination with superior performance.

[0065] In some implementations, monitoring the remote safety operator's response to warning information includes recording detailed actions taken by the safety operator to accept, ignore, or intervene in the warning. Acceptance records include actions such as clicking the confirmation button, executing suggested actions, and adjusting vehicle control parameters. Ignoring records include actions such as failing to respond to the warning after a timeout, closing the warning window, and continuing the existing operation. Intervention records include active control actions such as emergency braking, steering maneuvers, and speed adjustments. The operation records are in a structured data format to facilitate subsequent statistical analysis and pattern recognition.

[0066] The response time for warnings at different risk levels is statistically analyzed by examining the time interval between the warning being triggered and the safety officer's response. The average response time for alert-level warnings is 3-5 seconds, for warning-level warnings it is 1-3 seconds, for emergency-level warnings it is 0.5-1 second, and for hazard-level warnings it is less than 0.5 seconds. In some implementations, response time statistics are combined with analysis of individual safety officer differences to establish personalized response time baselines for abnormal behavior detection and fatigue assessment.

[0067] The cognitive load of safety officers is assessed by monitoring indicators such as attention deficit, operational error rate, and response delay. Cognitive load assessment employs a multi-dimensional scoring model, comprehensively considering factors such as working hours, number of monitored vehicles, warning frequency, and environmental complexity. In some implementations, cognitive load assessment is combined with physiological signal monitoring, using biomarkers such as heart rate variability, eye movement patterns, and electroencephalogram (EEG) signals to provide objective assessment data.

[0068] Establishing a database of safety officer operational behavior patterns is achieved by collecting and analyzing a large amount of safety officer operational data. The database includes information such as operational sequence patterns, decision-making preference characteristics, response time distribution, and error type statistics. In some implementations, behavior pattern recognition employs machine learning algorithms to automatically discover safety officers' operational habits and behavioral patterns, providing data support for personalized early warning strategies. The database utilizes privacy protection technologies to ensure the security of safety officers' personal information.

[0069] In some implementations, a dual verification mechanism combining a visual language model and traditional computer vision algorithms enhances the reliability of risk assessment through parallel processing and result comparison. The visual language model handles semantic understanding and complex scene analysis, while the traditional computer vision algorithms handle object detection and motion tracking. When the judgments from both algorithms are consistent, the confidence level of the risk assessment is increased. When the judgments differ, an advanced security protocol is activated, employing a conservative decision-making model to prioritize security.

[0070] The dynamic frequency adjustment strategy for scene complexity assesses scene complexity based on factors such as the number of traffic participants, the rate of change in their motion states, and environmental conditions. Simple scenarios, including straight-line driving and situations with few traffic participants, use an analysis frequency of 10Hz. Complex scenarios, including intersections, multi-vehicle convergence, and severe weather, use an analysis frequency of 30Hz. In some implementations, an adaptive algorithm is employed to dynamically adjust the analysis frequency based on real-time scene changes, balancing computational resource consumption and analysis accuracy requirements.

[0071] The risk-scoring priority scheduling algorithm prioritizes computational tasks based on a comprehensive risk score. High-risk scenarios receive higher processing priority and more computing resources. The scheduling algorithm employs a preemptive scheduling strategy, pausing the execution of low-priority tasks when a higher-risk scenario is detected. In some implementations, priority scheduling is combined with a load balancing mechanism to avoid overloading a single server node and ensure the stability of overall processing performance.

[0072] Dynamic load balancing resource expansion methods achieve automatic scaling by monitoring metrics such as CPU utilization, memory usage, and network bandwidth of the server cluster. When the system load exceeds a preset threshold, backup server nodes are automatically activated to increase processing capacity. The load balancing algorithm employs a weighted round-robin strategy, assigning different weights based on server performance differences. In some implementations, resource expansion supports elastic scaling, automatically adjusting the number of servers according to changes in business needs, achieving a balance between cost optimization and performance assurance.

[0073] In summary, referring to Figure 5, one implementation scenario of the remote driving control method described in this invention includes: achieving an end-to-end high-risk scene identification and early warning process through the coordinated operation of an in-vehicle terminal, cloud analysis, and a remote operation terminal. At the in-vehicle terminal, 11 high-definition cameras form a 360-degree panoramic view, providing comprehensive visual monitoring of the environment surrounding the remote vehicle. The multiple video streams acquired by the cameras are obtained in real-time at 1080p resolution and 30fps frame rate, ensuring high-quality visual input data for subsequent processing stages.

[0074] NVIDIA Orin series automotive compute units are responsible for data processing and lightweight model inference tasks on the vehicle. The automotive compute units perform preprocessing operations such as denoising, enhancement, and key region detection on the acquired multi-channel video streams to improve video data quality and highlight potentially risky areas. In some implementations, the automotive compute units deploy a lightweight visual language model with 0.5B parameters to achieve preliminary risk screening with a response time of less than 50ms, instantly identifying obvious risk scenarios and initiating emergency warning mechanisms.

[0075] As shown in Figure 5, the high-reliability 5G communication module integrates dual-carrier redundant links to ensure reliable communication between the vehicle-mounted device and the cloud. The dual-carrier redundant links connect simultaneously to the 5G networks of two different carriers. When the primary link fails or signal quality deteriorates, it automatically switches to the backup link, ensuring the continuity and stability of data transmission. Processed video data and preliminary analysis results are uploaded to the cloud via the 5G network. The transmission process employs data compression and priority scheduling strategies to optimize network bandwidth utilization efficiency.

[0076] The cloud-based analytics utilizes an NVIDIA A100×4 GPU server cluster to deploy the complete visual language model, providing high-performance deep scene understanding and risk assessment capabilities. The server cluster is equipped with low-latency solid-state storage for scene caching and historical data comparative analysis. In some implementations, the complete visual language model employs a dedicated architecture with 13B parameters, optimized for driving scenarios, and incorporates 5 million detailed traffic scene images during the pre-training phase.

[0077] Referring to Figure 5, the cloud processing workflow fuses and performs temporal analysis on the received multi-view videos to construct a unified scene representation. Multi-view fusion technology integrates video streams from different cameras into a complete environmental perception image, while temporal analysis tracks the dynamic changes of various elements within the scene. Cloud analysis achieves three layers of scene understanding: basic, relational, and semantic. The basic layer identifies road elements and their attributes; the relational layer analyzes the spatial relationships and interaction patterns between elements; and the semantic layer understands the overall semantic meaning and potential risks of the scene.

[0078] The risk assessment process employs a quantitative analysis across five dimensions: collision probability, time urgency, severity of consequences, difficulty of avoidance, and assessment certainty. Collision probability is expressed as a percentage from 0-100%, the time urgency is quantified with second-level precision for the decision-making window, the severity of consequences and difficulty of avoidance are assessed using a 1-10 scale, and assessment certainty reflects the credibility of the risk assessment results. The five parameters are weighted and fused using a algorithm to calculate a comprehensive risk score, generating tiered early warning information at four levels: alert, warning, emergency, and danger.

[0079] As shown in Figure 5, a single server node in the cloud supports concurrent processing of 8-12 vehicles. Real-time monitoring of multiple vehicles is ensured through multi-threaded parallel processing and dynamic resource allocation mechanisms. The load balancing algorithm allocates computing resources based on the scenario complexity and risk level of each vehicle, with high-risk scenarios receiving higher processing priority. In some implementations, the server cluster supports elastic scaling, automatically adjusting the number of nodes according to changes in business load to ensure that processing capacity matches actual needs.

[0080] The remote control terminal employs a professional three-screen display configuration. The central screen displays the primary driving view, while the left and right screens display auxiliary views and an augmented reality risk warning interface, respectively. This three-screen layout provides the remote safety operator with comprehensive visual information. The central screen focuses on the field of vision in the primary driving direction, while the auxiliary screens extend the visual coverage to the sides and rear. The augmented reality risk warning interface overlays risk areas as a colored overlay onto the real-time video feed, intuitively displaying the location, type, and severity of the risk.

[0081] Referring to Figure 5, the warning presentation employs differentiated multimodal methods based on the risk level. The alert level uses visual markers on the interface, attracting the safety officer's attention through color changes or icon displays. The warning level uses a combination of visual markers and voice prompts, with the voice prompts describing the risk content in concise and clear language. The emergency level uses multiple warning methods, including visual markers, voice alarms, and interface flashing, with the flashing using high-contrast colors to enhance visual impact. The hazard level employs a multimodal, all-around warning, combining visual, auditory, and tactile feedback, and triggering an automatic pre-braking mechanism to give the safety officer more reaction time.

[0082] Operator status monitoring assesses the cognitive load level of remote safety operators in real time, quantifying it through monitoring indicators such as attention distribution, response time, and operation frequency. The cognitive load assessment results are used to adjust the warning intensity; when a high cognitive load is detected, the warning intensity is automatically increased, employing more prominent visual cues and more frequent voice reminders. In some implementations, operator status monitoring combines eye-tracking and physiological signal monitoring technologies to provide a more accurate cognitive status assessment.

[0083] As shown in Figure 5, safety officer response behavior records are fed back to the cloud model as feedback data, forming a closed-loop mechanism for continuous learning and parameter optimization. Response behavior records include the safety officer's acceptance, ignoring, or intervention actions in response to warnings, along with corresponding operation timestamps and result evaluations. Feedback data is used to analyze the accuracy and effectiveness of warnings and identify areas for optimization in the warning algorithm. The parameter optimization process adjusts the risk assessment weight coefficients, warning trigger thresholds, and similarity matching parameters based on the feedback analysis results.

[0084] The continuous learning mechanism expands the scenario library and optimizes the algorithm model by collecting new risk scenarios and safety operator experience. New risk scenario identification is based on the similarity comparison between scenario feature vectors and the historical scenario library. When an unknown risk pattern is discovered, it is stored in the scenario library for subsequent training. In some implementations, continuous learning employs incremental learning algorithms, incorporating new scenario understanding capabilities without affecting existing knowledge, thereby achieving iterative improvements in the ability to identify high-risk scenarios in remote driving.

[0085] The end-to-end processing workflow optimizes the entire chain from video data acquisition to early warning information presentation. Onboard preprocessing reduces data transmission volume, cloud-based deep analysis provides accurate risk assessment, and intelligent presentation on remote terminals ensures timely and effective early warning delivery. Each stage works collaboratively through asynchronous parallel processing and priority scheduling mechanisms, achieving millisecond-level end-to-end response performance under limited computing resources, providing technical assurance for the safety and reliability of remote driving control.

[0086] This invention also discloses a remote overhead system.

[0087] Referring to Figure 6, a remote driving control system 500 includes multiple functional modules to achieve intelligent monitoring and risk warning of remote vehicles. The remote driving control system 500 adopts a distributed architecture design, achieving end-to-end high-risk scenario identification and early warning processing capabilities through the collaborative operation of on-board terminals, cloud analysis, and remote operation terminals. In some embodiments, the remote driving control system 500 supports simultaneous monitoring of multiple remote vehicles, and a single operator can manage the safety monitoring tasks of 6-8 vehicles.

[0088] As shown in Figure 6, the remote driving control system 500 includes a data acquisition module 502, used to acquire multi-channel video stream data from a remote vehicle. The data acquisition module 502 integrates 11 high-definition cameras, forming a 360-degree panoramic view coverage, providing comprehensive visual monitoring capabilities for the environment surrounding the remote vehicle. The cameras use 1080p resolution and a 30fps frame rate for real-time data acquisition, ensuring high-quality and timely video stream data. In some embodiments, the data acquisition module 502 is equipped with an NVIDIA Orin series in-vehicle computing unit, responsible for the initial processing and data preprocessing tasks of the video stream.

[0089] The onboard computing unit performs preprocessing operations such as noise reduction, enhancement, and key area detection to improve video data quality and highlight potentially risky areas. The data acquisition module 502 integrates a highly reliable 5G communication module, supporting dual-carrier redundant link configurations to ensure reliable communication between the onboard unit and the cloud. The dual-carrier redundant link connects simultaneously to two different carriers' 5G networks; when the primary link fails or signal quality degrades, it automatically switches to the backup link, ensuring the continuity and stability of data transmission.

[0090] Referring to Figure 6, the remote driving control system 500 also includes a scene understanding module 504, used for scene understanding of multi-channel video stream data based on a visual language model, identifying traffic participants and their behavioral characteristics in the traffic scene. The scene understanding module 504 is deployed in a cloud-based analytics subsystem, employing an NVIDIA A100×4 GPU server cluster to provide high-performance deep scene understanding capabilities. The server cluster is equipped with low-latency solid-state storage for scene caching and historical data comparison analysis. In some implementations, the scene understanding module 504 uses a complete visual language model with 13-bit parameters, specifically optimized for driving scenarios.

[0091] The scene understanding module 504 performs multi-level scene analysis and processing, including basic-level analysis to identify road elements and their attributes, relation-level analysis to analyze spatial relationships and interaction patterns between elements, and semantic-level analysis to understand the overall semantics and potential risk implications of the scene. Traffic participant identification covers road elements such as vehicles, pedestrians, and non-motorized vehicles, and behavioral feature analysis includes multi-dimensional feature extraction such as location information, movement state, attention direction, and behavioral intent. In some implementations, the scene understanding module 504 supports concurrent analysis tasks for 8-12 vehicles on a single server node, ensuring real-time monitoring of multiple vehicles through multi-threaded parallel processing and dynamic resource allocation mechanisms.

[0092] As shown in Figure 6, the remote driving control system 500 also includes a risk assessment module 506, which quantifies the risk level of identified potential risk scenarios based on the scenario understanding results using a multi-dimensional risk assessment algorithm, generating a structured risk assessment result that includes risk description, severity, and time urgency. The risk assessment module 506 employs a five-dimensional parameter assessment system, comprehensively calculating parameters such as collision probability, time urgency, severity of consequences, difficulty of avoidance, and assessment certainty. Collision probability is expressed as a percentage from 0-100%, time urgency is quantified with second-level precision, and the severity of consequences and difficulty of avoidance are assessed using a 1-10 level rating system.

[0093] The risk assessment module 506 calculates a comprehensive risk score using a weighted fusion algorithm, classifying risk levels into four categories: alert, warning, emergency, and danger. The structured risk assessment results include detailed risk descriptions, presented in natural language as the risk type, location, and recommended handling procedures. In some implementations, the risk assessment module 506 employs a dynamic threshold adjustment mechanism, adaptively adjusting warning triggering conditions based on driving environment, vehicle speed, and safety operator status, thereby improving warning sensitivity in complex traffic environments.

[0094] The remote driving control system 500 also includes a warning presentation module 508, which generates targeted warning information based on structured risk assessment results and presents it to the remote safety operator through a human-machine interface. The warning presentation module 508 is deployed on the remote operation terminal and uses a three-screen professional display system configuration. The central screen displays the main driving view, while the left and right screens display auxiliary views and augmented reality risk warning interfaces, respectively. This three-screen display layout provides the remote safety operator with comprehensive visual information, expands the visual coverage, and highlights risk areas.

[0095] The early warning presentation module 508 selects differentiated early warning presentation methods based on the risk level. For the alert level, visual markers are used on the interface; for the warning level, a combination of visual markers and voice prompts is used; for the emergency level, multiple warnings are used, including visual markers, voice alarms, and interface flashing; and for the danger level, a multimodal, all-around early warning is used, triggering an automatic pre-braking mechanism. Augmented reality technology overlays the risk area as a colored overlay onto the real-time video screen, intuitively displaying the risk location, type, and severity. In some implementations, the early warning presentation module 508 integrates operator status monitoring, assessing the cognitive load level of remote safety personnel in real time and adjusting the early warning intensity based on their cognitive state.

[0096] As shown in Figures 5 and 6, the modules establish a complete processing chain through data transfer and collaborative working mechanisms. The data acquisition module 502 transmits the acquired multi-channel video stream data to the scene understanding module 504. The scene understanding module 504 performs in-depth analysis of the video data and transmits the scene understanding results to the risk assessment module 506. The risk assessment module 506 performs a quantitative risk assessment based on the scene analysis results, generates a structured risk assessment result, and transmits it to the early warning presentation module 508. The early warning presentation module 508 generates targeted early warning information based on the risk assessment results and presents it to the remote safety officer through a human-computer interaction interface.

[0097] In some implementations, the modules employ an asynchronous parallel processing mechanism. The data acquisition module 502 continuously collects video stream data, while the scene understanding module 504 and risk assessment module 506 perform analysis and processing in parallel. The early warning presentation module 508 updates the early warning display in real time. A priority scheduling algorithm allocates computing resources based on risk level, with high-risk scenarios receiving higher processing priority and faster response times. A load balancing mechanism ensures efficient resource utilization of the server cluster and supports elastic scaling to adapt to changes in business load. The remote driving control system 500, through modular design and a collaborative working mechanism, achieves end-to-end intelligent processing capabilities from data acquisition to early warning presentation.

[0098] This invention also discloses a readable storage medium.

[0099] A computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the remote driving method described in any of the above embodiments. The computer-readable storage medium may include any entity or device capable of carrying a computer program, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc. The computer program includes computer program code. The computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable storage medium may include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc.

[0100] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0101] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a system including a processing module or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0102] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A remote driving control method, characterized in that, include: Acquire multi-channel video stream data from remote vehicles; Based on a visual language model, scene understanding is performed on the multi-channel video stream data to identify traffic participants and their behavioral characteristics in the traffic scene. Based on the scenario understanding results, the risk level of the identified potential risk scenarios is quantified through a multi-dimensional risk assessment algorithm, generating a structured risk assessment result that includes risk description, severity, and time urgency. Based on the structured risk assessment results, targeted early warning information is generated and presented to remote safety officers through a human-computer interaction interface.

2. The remote driving control method according to claim 1, characterized in that, The scene understanding based on the visual language model includes: preprocessing the multi-channel video stream data, including denoising, enhancement, and key region detection; performing basic risk screening using a lightweight visual language model; transmitting the preprocessed video data and basic screening results to the cloud; and performing deep scene understanding analysis using a complete visual language model deployed in the cloud.

3. The remote driving control method according to claim 2, characterized in that, The deep scene understanding analysis includes: basic-level analysis, which identifies road elements and their attributes; relation-level analysis, which analyzes the spatial relationships and interaction patterns between elements; and semantic-level analysis, which understands the overall semantics of the scene and the implications of potential risks.

4. The remote driving control method according to claim 1, characterized in that, The multi-dimensional risk assessment algorithm includes: calculating collision probability parameters; quantifying time urgency parameters; assessing the severity of consequences parameters; determining the difficulty of avoidance parameters; and calculating and assessing certainty parameters.

5. The remote driving control method according to claim 4, characterized in that, The risk level quantification includes: calculating a comprehensive risk score based on the five parameters; classifying the risk level into four levels—alert, warning, emergency, and danger—based on the comprehensive risk score; and configuring a corresponding warning trigger threshold for each risk level.

6. The remote driving control method according to claim 1, characterized in that, The generation of targeted early warning information includes: selecting the early warning presentation method according to the risk level, wherein the reminder level uses visual markers on the interface, the warning level uses visual markers and voice prompts, the emergency level uses visual markers, voice alarms and interface flashing, and the danger level uses multimodal all-round early warning; displaying the risk location and type in the remote operation interface through augmented reality technology; and adjusting the early warning intensity according to the current cognitive load of the remote safety officer.

7. The remote driving control method according to claim 1, characterized in that, Also includes: Monitor the response behavior of remote security personnel to early warning information; Analyze the accuracy and effectiveness of early warning systems; Adjust the risk assessment parameters based on the analysis results; And new risk scenarios are stored in the scenario library for system optimization.

8. The remote driving control method according to claim 7, characterized in that, The monitoring of remote safety officers' response to early warning information includes: recording safety officers' acceptance, ignoring, or intervention actions in response to early warnings; statistically analyzing the response time for early warnings at different risk levels; assessing changes in the cognitive load of safety officers at different time periods; and establishing a database of safety officer operational behavior patterns.

9. The remote driving control method according to claim 1, characterized in that, The identification of traffic participants and their behavioral characteristics in traffic scenarios includes: identifying the position and movement status of vehicles, pedestrians, and non-motorized vehicles; analyzing the attention direction and behavioral intention of each traffic participant; predicting the possible movement trajectory of each participant; constructing a relationship graph among traffic participants; and calculating the similarity between the current scenario and historical high-risk scenarios based on historical driving data.

10. A remote driving control system, characterized in that, include: The data acquisition module is used to acquire multi-channel video stream data from remote vehicles; The scene understanding module is used to perform scene understanding on the multi-channel video stream data based on a visual language model, and to identify traffic participants and their behavioral characteristics in the traffic scene. The risk assessment module is used to quantify the risk level of the identified potential risk scenarios based on the scenario understanding results, and generate a structured risk assessment result that includes risk description, severity and time urgency. And an early warning presentation module, used to generate targeted early warning information based on the structured risk assessment results and present it to remote safety officers through a human-computer interaction interface.

11. A readable storage medium, characterized in that, The readable storage medium stores computer instructions that, when executed by a processor, implement the remote driving method as described in any one of claims 1-9.