Multi-vehicle cooperative command and intelligent scheduling method and system

By optimizing multi-vehicle collaborative command and dispatch through digital twin command view and reinforcement learning model, and combining edge computing and adaptive selection of communication links, the problems of data redundancy and security vulnerabilities in existing technologies are solved, and efficient and secure multi-vehicle collaborative dispatch is achieved.

CN121640728AInactive Publication Date: 2026-03-10BEIJING HIZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack adaptive learning capabilities in multi-vehicle collaborative command and dispatch, resulting in data redundancy, poor timeliness, inability to achieve efficient information exchange and accurate spatiotemporal resource allocation, and potential vulnerabilities in safety management measures.

Method used

By employing a digital twin command view, reinforcement learning models, and edge computing technology, combined with vehicle state matrices and adaptive selection of communication links, the scheduling model and safety boundaries are optimized. Tasks are allocated and edge preprocessing is performed through reinforcement learning models to achieve efficient data transmission and safe management.

Benefits of technology

It improves the intelligence, responsiveness, and safety of multi-vehicle collaborative operations, and enhances the system's robustness and data transmission efficiency in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640728A_ABST
    Figure CN121640728A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-vehicle cooperative command and intelligent scheduling method and system, and belongs to the technical field of intelligent traffic systems, and the method comprises the steps: obtaining registration request data, including vehicle identifiers, real-time position coordinates and sensor types; according to the real-time position coordinates and preset geographic information data, generating a digital twin command view, and outputting a vehicle state matrix; target vehicles and task sub-targets are distributed through a pre-configured reinforcement learning model, and a scheduling instruction set and a task state database are generated; controlling the target vehicle to collect sensor data and perform edge preprocessing to generate an edge preprocessing data set and a terminal state code; selecting a communication link to transmit the edge preprocessing data set to a command center, and updating a task state database; and dynamically optimizing a reward function and an electronic fence boundary of the reinforcement learning model. According to the invention, efficient command and accurate scheduling in multi-vehicle collaborative operation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation systems technology, and in particular to a method and system for multi-vehicle collaborative command and intelligent dispatch. Background Technology

[0002] With the acceleration of urbanization and the continuous expansion of transportation, traditional single-point vehicle management methods are no longer able to meet the needs of increasingly complex traffic organization patterns and diverse application scenarios. Especially in situations requiring multi-party collaboration, such as emergency rescue and security for large-scale events, how to achieve efficient information exchange, dynamic task coordination, and precise spatiotemporal resource allocation among multiple mobile platforms has become a key bottleneck restricting the improvement of overall system efficiency.

[0003] In existing technological systems, task assignment mechanisms largely rely on human experience to formulate rules and logic, lacking adaptive learning capabilities. This often results in low flexibility and robustness when facing complex and changing environments. Although some existing systems have introduced sensor fusion technology and local data processing units to alleviate the burden on backbone communication, the failure to fully integrate the actual operating status of the vehicle-mounted terminal and the quality changes of surrounding communication links for joint optimization leads to high redundancy and poor timeliness in uploaded data, failing to provide high-quality data support to higher-level command organizations. Furthermore, most current systems still maintain static configuration for sensitive area boundaries, failing to consider the impact of electromagnetic environment fluctuations and terrain evolution on actual coverage effectiveness, creating potential vulnerabilities in safety management measures. Summary of the Invention

[0004] To address the aforementioned technical issues, this application provides a method and system for multi-vehicle collaborative command and intelligent dispatch.

[0005] Firstly, this application provides a method for multi-vehicle collaborative command and intelligent dispatching, employing the following technical solution: A method for multi-vehicle collaborative command and intelligent dispatch, the method comprising: Obtain registration request data, including vehicle identifier, real-time location coordinates, and sensor type; Based on the real-time location coordinates and preset geographic information data, a digital twin command view is generated and a vehicle status matrix is ​​output; wherein, the digital twin command view includes a vehicle location layer, a signal coverage heat map layer, and an electronic fence boundary. In response to the input task instructions, based on the vehicle state matrix, electronic fence boundaries and sensor types, the target vehicle and task sub-objectives are allocated through a pre-configured reinforcement learning model, and a scheduling instruction set and task state database are generated. The target vehicle is controlled to collect sensor data according to the scheduling instruction set, and the sensor data is preprocessed at the edge to generate an edge preprocessing dataset and a terminal status code. Based on the terminal status code and real-time network quality data, a communication link is selected to transmit the edge preprocessing dataset to the command center and update the task status database. Based on the historical task status database and communication logs, the reward function and electronic fence boundary of the reinforcement learning model are dynamically optimized.

[0006] By adopting the above technical solutions, a digital twin command view, a vehicle status matrix, and a reinforcement learning-driven intelligent scheduling mechanism were constructed, achieving efficient command and precise scheduling in multi-vehicle collaborative operations. Combined with edge computing data preprocessing and adaptive communication link selection strategies, data transmission efficiency and system robustness were improved. Furthermore, through closed-loop feedback optimization of the task status database and communication logs, the scheduling model and safety boundary settings were continuously improved, significantly enhancing the system's intelligence level, responsiveness, and security in complex environments.

[0007] Secondly, this application provides a multi-vehicle collaborative command and intelligent dispatch system, which adopts the following technical solution: A multi-vehicle collaborative command and intelligent dispatch system, the system comprising: The data acquisition module is used to acquire registration request data, including vehicle identifier, real-time location coordinates, and sensor type; The vehicle status matrix output module is used to generate a digital twin command view based on the real-time location coordinates and preset geographic information data, and output the vehicle status matrix; wherein, the digital twin command view includes a vehicle location layer, a signal coverage heat map layer, and an electronic fence boundary. The task scheduling module is used to respond to input task instructions, and based on the vehicle state matrix, electronic fence boundaries and sensor types, allocate target vehicles and task sub-objectives through a pre-configured reinforcement learning model, and generate a scheduling instruction set and task state database. The edge preprocessing module is used to control the target vehicle to collect sensor data according to the scheduling instruction set, perform edge preprocessing on the sensor data, and generate an edge preprocessing dataset and a terminal status code. The communication transmission module is used to select a communication link to transmit the edge preprocessed dataset to the command center based on the terminal status code and real-time network quality data, and to update the task status database. The model optimization module is used to dynamically optimize the reward function and electronic fence boundary of the reinforcement learning model based on the historical task status database and communication logs.

[0008] Thirdly, this application provides a computer device, which adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to perform the steps of the method as described in the first aspect.

[0009] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description

[0010] Figure 1 This is a first flowchart illustrating a multi-vehicle collaborative command and intelligent dispatch method according to one embodiment of this application.

[0011] Figure 2 This is a second flowchart illustrating a multi-vehicle collaborative command and intelligent dispatch method according to one embodiment of this application.

[0012] Figure 3 This is a schematic diagram of the third process of a multi-vehicle collaborative command and intelligent dispatch method according to one embodiment of this application.

[0013] Figure 4 This is a schematic diagram of the fourth process of a multi-vehicle collaborative command and intelligent dispatch method according to one embodiment of this application.

[0014] Figure 5 This is a schematic diagram of the fifth process of a multi-vehicle collaborative command and intelligent dispatch method according to one embodiment of this application.

[0015] Figure 6 This is a schematic diagram of the sixth process of a multi-vehicle collaborative command and intelligent dispatch method according to one embodiment of this application.

[0016] Figure 7 This is a schematic diagram of the seventh process of a multi-vehicle collaborative command and intelligent dispatch method according to one embodiment of this application. Detailed Implementation

[0017] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0018] This application discloses a method for multi-vehicle collaborative command and intelligent scheduling.

[0019] Reference Figure 1A method for multi-vehicle collaborative command and intelligent dispatch, the method comprising: Step S101: Obtain registration request data, including vehicle identifier, real-time location coordinates, and sensor type; During the system startup phase, registration request data needs to be acquired. This involves collecting basic identity information and current operating status parameters of all mobile terminals (i.e., vehicles) participating in the collaborative operation.

[0020] Specifically, the vehicle identifier is used to uniquely identify each vehicle device connected to the system; the real-time location coordinates provide spatial positioning information of the vehicle's current geographical location, usually from high-precision positioning devices such as Global Navigation Satellite System (GNSS) or Differential GPS; and the sensor type indicates the information category attributes of different types of sensor modules carried by each vehicle, such as visible light cameras, infrared imagers, lidar (LiDAR), millimeter-wave radar, or environmental monitoring probes.

[0021] Step S102: Generate a digital twin command view based on real-time location coordinates and preset geographic information data, and output a vehicle status matrix; wherein, the digital twin command view includes a vehicle location layer, a signal coverage heat map layer, and an electronic fence boundary. Specifically, digital twins refer to a modeling method that maps the state of physical objects to a virtual space, thereby enabling real-time monitoring and predictive simulation. The vehicle location layer projects the previously obtained location coordinates onto a map plane, forming a dynamically changing point-based distribution map. The signal coverage heatmap layer, based on the performance indicators of wireless communication facilities pre-deployed on network nodes (such as base station transmit power, propagation loss models, etc.), combined with actual detected field strength samples, plots the network service quality assessment results at different locations within the area, with color intensity indicating signal strength. The setting of electronic fence boundaries delineates specific areas based on the heatmap; these areas may have restricted access or operational permissions due to severe electromagnetic shielding, significant terrain obstruction, or other security considerations. The comprehensive view formed in this context not only helps to intuitively understand the global resource distribution but also provides key reference dimensions for subsequent task planning.

[0022] Meanwhile, the vehicle state matrix obtained by weighting various elements is essentially a structured vector set. Its elements include, but are not limited to, state variables such as the distance between the vehicle's current position and the target point, whether it is near a restricted area, what type of sensing components it is equipped with, and the remaining battery / fuel level. These parameters will directly affect the selection tendency of candidate objects in the subsequent task matching process.

[0023] Step S103: In response to the input task command, based on the vehicle state matrix, electronic fence boundary and sensor type, the target vehicle and task sub-objective are allocated through a pre-configured reinforcement learning model, and a scheduling command set and task state database are generated. When a specific task instruction is received from an external source, the system begins to execute task decomposition and resource allocation operations based on a reinforcement learning model. In this part, the reinforcement learning algorithm plays a central role; its basic idea is to use a trial-and-error approach to explore and search for the optimal strategy.

[0024] In this embodiment, the Q-learning framework can be used as one of the model-free reinforcement learning paradigms. Its characteristic is that it can complete the learning iteration of the value function without prior knowledge of the complete environment transition probability matrix. The input mainly includes two aspects: first, the various index values ​​from the vehicle state matrix generated in the previous step, especially those features reflecting mobility and availability; second, the relevant descriptive information of the task to be processed, such as the coordinates of the task location, the types of data to be collected, and constraints such as deadlines.

[0025] It is worth noting that, to improve matching efficiency and robustness, an auxiliary factor called "task urgency" can be introduced. This factor can be quantified by calculating the percentage of overlap between the task area and known danger zones. A larger overlap indicates a higher risk, and a higher priority coefficient should be assigned accordingly. After the trained Q-network continuously adjusts the action-value estimates until convergence and stability, the final output includes two main parts: one part specifies which vehicle or vehicles are most suitable for each subtask, and the other part is a series of specific driving route suggestions. These route designs fully consider various possible disturbances and interference factors, and reserve approximately 10% length margin as a fault tolerance buffer to enhance the ability to cope with emergencies.

[0026] Step S104: Control the target vehicle to collect sensor data according to the scheduling instruction set, perform edge preprocessing on the sensor data, and generate edge preprocessing dataset and terminal status code; The process involves issuing a set of dispatch instructions according to a predetermined plan, triggering target vehicles to conduct on-site data collection and perform preliminary processing locally. Specifically, the controlled platform activates its various onboard detection instruments to conduct on-site observation activities based on the issued instructions. The resulting massive amounts of raw streaming media data often contain a large amount of redundant noise components, so it is necessary to compress and refine them to improve upload bandwidth utilization and reduce the load on the central server.

[0027] To this end, the system employs a series of advanced edge computing techniques for front-end preprocessing. For radio frequency signals, wavelet transform is used for denoising and filtering to effectively suppress background white noise, while extracting energy peaks within the frequency band of interest as representative statistical features. For video sequences, one frame is captured every second as a keyframe, and then compressed using the H.265 international standard protocol. This ensures the integrity of the visual content while significantly reducing file size.

[0028] In addition, all processed data is timestamped and linked to the original collection time for easy traceability and accountability later. After the entire process is complete, a standardized "edge preprocessing dataset" is generated for subsequent transmission. Simultaneously, a set of "terminal status codes" reflecting the terminal's health status is also generated. This code consists of several binary bits, each corresponding to a flag indicating whether a certain hardware functional module is functioning correctly, which can be used to determine if any abnormal faults exist.

[0029] Step S105: Based on the terminal status code and real-time network quality data, select the communication link to transmit the edge preprocessed dataset to the command center and update the task status database. In complex field environments where multiple heterogeneous networks may coexist, the rational allocation of various channel resources is particularly important. The system then determines which combination of methods should be used to send message packets based on the terminal status code.

[0030] For example, if the test results show that all functions are intact, the high-speed, low-latency fifth-generation cellular mobile communication system (5G NR) will be selected by default to carry the main service traffic; otherwise, if it is found that some components are damaged, which increases the risk of loss of critical data, the system will be immediately switched to a more reliable backup channel, such as the Ka / Ku band satellite communication link to transmit important summary reports. At the same time, the Mesh self-organizing network topology will be used to try to restore the original file fragment backup copy as a precaution.

[0031] Understandably, this flexible switching mechanism greatly enhances the overall system's survivability and resilience, maintaining a minimum level of effective communication even in the face of extreme weather or enemy interference attacks. Furthermore, after each successful reception and parsing of a new batch of data packets, the relevant records in the central database are promptly updated to maintain contextual consistency.

[0032] Step S106: Dynamically optimize the reward function and electronic fence boundary of the reinforcement learning model based on the historical task status database and communication logs.

[0033] The system also needs to execute a closed-loop feedback adjustment procedure to continuously improve its future performance. This part involves two evolutionary directions: First, the upgrading and improvement of the reinforcement learning model itself. Specifically, this involves extracting past problem cases and corresponding solutions from a long-accumulated historical task execution archive, and extracting new reward and punishment criteria that are more in line with actual situations, so that more accurate and reasonable choices can be made when facing similar situations in the future. Second, the revision project surrounding the adjustment of electronic fence boundaries. With the accumulation of more field survey experience, the original static division is obviously unable to meet the growing security protection needs. Therefore, redefining the restricted area with the help of periodically scanning and updating the signal field strength distribution image has become an inevitable trend. In particular, special attention should be paid to areas with repeated weak signal blind spots, and these areas should be included in the new control list to prevent accidental intrusions from happening again.

[0034] In the above implementation, a digital twin command view, a vehicle state matrix, and a reinforcement learning-driven intelligent scheduling mechanism are constructed to achieve efficient command and precise scheduling in multi-vehicle collaborative operations. By combining edge computing data preprocessing and adaptive communication link selection strategies, data transmission efficiency and system robustness are improved. Furthermore, through closed-loop feedback optimization of the task state database and communication logs, the scheduling model and safety boundary settings are continuously improved, significantly enhancing the system's intelligence level, responsiveness, and security in complex environments.

[0035] Reference Figure 2 As one implementation of step S102, the step of generating a digital twin command view and outputting a vehicle status matrix based on real-time location coordinates and preset geographic information data includes: Step S201: Obtain the real-time location coordinates of the vehicle and preset geographic information data including historical signal coverage strength; In the stage of acquiring the vehicle's real-time location coordinates and preset geographic information data, the system needs to access the vehicle's current latitude and longitude information provided by multiple mobile terminals (such as vehicle-mounted GPS modules, feedback from cellular communication base stations, etc.). This information constitutes the basic spatiotemporal anchor points for the subsequent analysis of the entire system.

[0036] Simultaneously, the system also needs to load a pre-prepared geographic information database, which includes, but is not limited to, road network topology (i.e., road connectivity and attribute descriptions) and historical signal coverage strength records within the region. Historical signal coverage strength refers to the average power level of radio signals received in different areas over a past period, usually expressed in dBm, reflecting the service accessibility and quality distribution of communication infrastructure. This type of data can be derived from historical statistical data from operators or simulation results from simulation platforms.

[0037] Step S202: Generate a vehicle position layer based on the real-time vehicle position coordinates; Among them, the vehicle location layer projects the real-time vehicle location coordinates obtained above onto the map plane to form a dynamically changing point distribution pattern. Step S203: Generate an initial signal heat map based on historical signal coverage intensity, and update the heat map value weight distribution through real-time vehicle location coordinates to obtain the updated initial signal heat map. The system first utilizes existing historical signal coverage data to construct a spatial distribution image on the map reflecting the long-term communication service quality, known as the "initial signal heatmap layer." This image is essentially a rasterized two-dimensional matrix, with each cell corresponding to a specific area on the ground, and its value representing the historical average signal strength level of that area. To adapt this heatmap to constantly changing real-world conditions, the system introduces a real-time update mechanism. Specifically, whenever a new vehicle reports its current location, it also uploads the instantaneous signal strength sample value collected at that location.

[0038] Subsequently, the system creates a circular influence region centered on the location of each vehicle and applies a Gaussian function to correct the existing thermal values ​​within this region. The core idea behind this step is that areas closer to the sampling point should be more influenced by the new observation data, while areas farther away from the sampling point should still mainly rely on the original statistical results, thus achieving a local adaptive adjustment effect.

[0039] Specifically, the process of updating the thermal value weight distribution includes: generating a Gaussian attenuation weight field with the real-time vehicle location coordinates as the center and a preset radius as the influence range; and weighting and fusing the historical thermal values ​​within the attenuation weight field with the real-time signal sample values, calculated using the following formula: ; In the above formula, α is the history decay factor, used to adjust the balance ratio between old data and new measurements; d i σ is the distance between the sampling point and the center of the grid; σ is the standard deviation of the Gaussian distribution, which determines how quickly the influence range decreases with distance.

[0040] Step S204: Based on the updated initial signal heat map, the area where the signal strength is lower than the preset strength threshold for a preset duration is marked as the electronic fence boundary, and a signal coverage heat map is obtained. Specifically, by scanning the new heatmap, which has undergone multiple iterations of optimization, pixel by pixel, closed or semi-closed areas exhibiting continuous weak signal phenomena are identified. It should be noted that the requirement of a duration consistently below a preset threshold not only emphasizes the insufficiency of signal intensity but also adds a time dimension requirement, meaning that only areas that fail to restore good communication conditions for an extended period are considered candidates.

[0041] Furthermore, to avoid false alarms caused by chance, a minimum area threshold needs to be set for screening to ensure that the ultimately demarcated target truly constitutes a problem area of ​​a certain scale. Once the above two constraints are met, a polygonal outline can be drawn at the corresponding location to mark the potential risk isolation zone. However, since relying solely on signal indicators to demarcate boundaries may ignore the influence of actual traffic flow, it is necessary to further cross-compare it with the road topology to exclude situations such as bridge underpasses and tunnel entrances that, although temporarily without signals, belong to normal driving paths, thereby improving the rationality and practicality of the electronic fence demarcation.

[0042] Step S205: Overlay the vehicle location layer, signal coverage heat map, and electronic fence boundary onto the digital map to generate a digital twin command view; This step aims to integrate and display various heterogeneous information on a single interface. The digital twin command view is a highly abstract human-computer interaction interface that allows dispatchers to intuitively view the overall distribution of the fleet, key monitoring areas, and potential safety hazards.

[0043] Specifically, the vehicle location layer provides basic object tracking functionality; the signal heatmap layer displays the macro-level service availability status; and the geofence boundary, as a warning marker, highlights areas with concentrated problems. These three elements are both independent and interconnected, collectively forming a complete scene recognition model. In terms of technical implementation, such composite graphics are often rendered using a GIS engine, supporting various interactive methods such as zooming and panning, allowing users to flexibly focus on key areas of interest for in-depth analysis.

[0044] Step S206: Calculate the vehicle state weight value based on the spatial relationship between the vehicle's real-time location coordinates and the electronic fence boundary; In the process of calculating the vehicle state weight value based on the spatial relationship between the vehicle's real-time location coordinates and the electronic fence boundary, the system attempts to quantify the degree of risk exposure faced by each vehicle.

[0045] To address this, two complementary evaluation factor combination strategies were adopted: Firstly, the straight-line Euclidean distance d between any vehicle and its nearest edge of the electronic fence was calculated, and the corresponding distance attenuation coefficient β=e was derived accordingly.-λd λ is the decay rate parameter, which determines how steep the decline in risk score occurs as the vehicle moves away from the danger zone. On the other hand, considering the significant differences in importance of different types of sensing devices in the emergency response process, a priority factor γ is added to the evaluation system. For example, special vehicles equipped with high-definition cameras or emergency relief material transport devices should obviously be given higher attention. The product of the two, ω = β·γ, is defined as the vehicle's comprehensive state weight value, which can be used in functional modules such as sorting and classification, resource allocation, and even triggering automatic alarms.

[0046] Step S207: Output a vehicle state matrix containing vehicle identifier, real-time location coordinates, and state weight values.

[0047] The vehicle status matrix is ​​actually a tabular data container. Row indices correspond to different vehicle IDs, and column fields include, but are not limited to, a unique identification code, the current latitude and longitude combination, and the calculated risk assessment value. This organizational format greatly facilitates advanced computational tasks such as retrieval, batch processing, and even machine learning modeling for upper-layer applications.

[0048] In the above implementation, relying on a multi-layered information fusion architecture, this solution cleverly combines several key technologies, including spatiotemporal data analysis, physical field modeling theory, spatial geometric calculations, and weight evaluation mechanisms, successfully establishing a closed-loop pathway from the acquisition of data from underlying sensors to the issuance of management commands from higher levels. Compared to traditional manual inspection modes based on human experience or simple alarm systems that rely solely on single-dimensional threshold judgments, this technical solution significantly enhances the ability to anticipate and handle emergencies in complex urban traffic environments.

[0049] Reference Figure 3 As one implementation of step S103, in response to the input task command, the steps of allocating target vehicles and task sub-objectives through a pre-configured reinforcement learning model based on the vehicle state matrix, electronic fence boundaries, and sensor types, and generating a scheduling instruction set and task state database include: Step S301: Parse the input task instructions and extract the task target coordinates and task type identifier; The system needs to interface with raw task information from user terminals or other upper-level control platforms. These tasks typically exist in structured or unstructured forms, such as JSON messages, XML documents, or API call parameters. During parsing, the key is to identify two core elements: first, clearly defining the geographical location of the task (i.e., the task target coordinates), generally represented using the WGS84 latitude and longitude system; and second, distinguishing the requirement attributes of different types of tasks (such as data collection, transportation, and monitoring). This step is not only the foundation for all subsequent operations but also determines whether the entire scheduling strategy can accurately pinpoint the essential characteristics of the task requirements.

[0050] Step S302: Calculate the task urgency evaluation value based on the spatial relationship between the task target coordinates and the electronic fence boundary; In this context, an electronic fence boundary refers to one or more pre-defined, artificially demarcated geographical areas, commonly used to restrict specific activity areas, monitor abnormal behavior, or delineate priority response zones. In this context, the relative position of the mission target's coordinates is closely related to the security and sensitivity of its environment.

[0051] To scientifically quantify the urgency of this spatial correlation, the system introduces a composite index: the task urgency evaluation value. Specifically, on the one hand, it uses GIS tools to obtain signal heatmap data around the target point, reflecting the communication load, event density, or resource density within the area; on the other hand, it calculates the proportion of overlapping area where the coordinate falls within multiple electronic fences. This measures the degree of spatial constraint, A total A represents the total area of ​​the mission region. overlap The area represents the overlap.

[0052] Subsequently, an urgency score E is generated through linear weighting to comprehensively assess the importance level of the current task. The specific formula is as follows: Where k1 and k2 are preset coefficients; this method avoids the traditional scheduling mode that simply relies on the nearest principle, and instead considers more complex real-world factors, making task sorting more reasonable and forward-looking.

[0053] Step S303: Input the task type identifier, vehicle state matrix, sensor type and task urgency evaluation value into the pre-configured reinforcement learning model, and output the matching mapping relationship between the target vehicle identifier and the task sub-objectives. The crucial inference stage involves inputting the task type identifier, vehicle state matrix, sensor type, and task urgency evaluation value from the vehicle state matrix into the reinforcement learning model, outputting the matching mapping relationship between the target vehicle and the task sub-objectives. The vehicle state matrix contains multi-dimensional state information for each candidate vehicle, including real-time operating status, battery level, load status, and current location. This information constitutes a vital component of the reinforcement learning environment state. Sensor types reflect the perception capabilities and functional adaptability of different vehicles, directly impacting their competence in performing specific tasks. The reinforcement learning model acts as an intelligent decision engine, maximizing the long-term reward function to achieve the optimal matching mapping relationship between the target vehicle and the task sub-objectives under complex multi-constraint conditions. This mapping relationship considers not only the current static matching degree but also incorporates the prediction and optimization of future state transitions.

[0054] Specifically, reinforcement learning algorithms (such as Q-learning) continuously update their action selection strategies based on historical experience, ensuring that each decision maximizes the long-term reward function as much as possible. Where w1, w2, and w3 are the corresponding weighting coefficients. This approach not only achieves effective allocation of vehicle resources but also balances task timeliness, route economy, and system stability, thereby significantly improving overall scheduling efficiency.

[0055] Step S304: Based on the matching mapping relationship and the road topology data in the preset geographic information data, generate an anti-offset path that includes a path point sequence and a fault tolerance buffer. In real-world transportation networks, due to their highly nonlinear and uncertain nature, traditional shortest path algorithms may struggle to handle the impact of factors such as GPS drift, temporary obstacles, and road closures. To address this, this application employs an enhanced path generation function. First, it uses the classic Dijkstra algorithm to quickly calculate the theoretically optimal route, then further refines it. Specifically, it sets several fault-tolerant checkpoints at fixed intervals (e.g., 10% of the total length) along the entire path, and constructs a circular buffer zone with a radius of 5 meters around each checkpoint. This design not only helps reduce the risk of misjudgments caused by sensor errors but also effectively improves path-following accuracy and disaster recovery capabilities. Furthermore, the choice of the circular buffer zone fully considers the average deviation range (approximately 3-5 meters) commonly found in civilian-grade GPS devices, ensuring that the actual driving trajectory remains within a safe and controllable range, greatly enhancing the system's robustness and practicality.

[0056] Step S305: Integrate the target vehicle identifier, task sub-objectives, and anti-offset path to generate a scheduling instruction set; The scheduling instruction set is essentially a well-structured package of operation commands, containing complete task context information, such as the specified vehicle number, a summary of the specific task content, recommended routes and their corresponding fault tolerance mechanisms, and deadline timestamps. This information, after standardized encoding, can be easily transmitted over the network or cached locally, allowing the front-end scheduling center to monitor the overall progress in real time and make corresponding adjustments. More importantly, this instruction set possesses excellent compatibility and scalability, supporting access from various terminal devices and interaction with heterogeneous platforms, facilitating seamless upgrades as the system scales up in the future.

[0057] Step S306: Create a task status data table and initialize the binding status and progress flag bits of the task sub-targets and target vehicle identifiers.

[0058] The creation of a task status data table initializes the binding status and progress flags between task sub-targets and target vehicles, marking the transition of the entire scheduling process from the planning phase to the execution and monitoring phase. As one of the core database objects within the system, the task status data table provides real-time feedback and closed-loop control. It primarily includes fields such as task ID, corresponding target vehicle identifier, current working status (pending execution / collecting data / completed), and remaining available time window. Continuously updating this status information not only allows for the timely detection of potential bottlenecks but also assists subsequent optimization modules in making more reasonable resource reallocation decisions. For example, if a task times out, the system can re-trigger a new round of scheduling based on the latest traffic conditions, thereby ensuring service continuity and user experience satisfaction.

[0059] The above implementation achieves fully automated control of the entire process, from task perception and intelligent allocation to reliable execution. The entire solution fully considers multiple constraints and uncertainties in real-world application scenarios. By introducing anti-offset path design and refined state management mechanisms, it significantly improves the intelligence level of vehicle scheduling and the reliability of system operation, ultimately achieving efficient, stable, and adaptive vehicle task scheduling technology in complex environments.

[0060] Reference Figure 4 As one implementation of step S104, the step of performing edge preprocessing on the sensor data includes: Step S401: Perform wavelet transform noise reduction on the original radio frequency signal in the sensor data, and extract the energy value of the preset frequency band as the first summary feature; Specifically, by utilizing the multi-resolution analysis characteristics of wavelet transform, the original radio frequency signal is jointly decomposed in the time and frequency domains, thereby effectively separating noise components and useful information from the signal. As a mathematical tool, wavelet transform can perform localized analysis of signals at different scales, making it particularly suitable for processing non-stationary and transient signals. By selecting appropriate wavelet basis functions and the number of decomposition levels, high-frequency noise can be specifically suppressed without losing the main characteristic information of the signal.

[0061] In the noise reduction process, soft or hard thresholding methods are typically used to process wavelet coefficients, removing small-amplitude coefficients that are mainly contributed by noise and retaining large coefficients that represent the main characteristics of the signal. Subsequently, the energy value of a preset frequency band is extracted as the first summary feature. This process essentially involves integrating or statistically calculating the energy of the denoised signal within a specific frequency band to form a quantitative indicator that reflects the activity level of the signal in that band. This energy feature has good stability and discriminative power, effectively characterizing the essential properties of radio frequency signals while significantly reducing data dimensionality, facilitating subsequent transmission and processing.

[0062] Step S402: Adaptively adjust the keyframe extraction probability of the raw video data in the sensor data, and encode the keyframes to generate the second summary features; Specifically, the probability or frequency of keyframe extraction can be dynamically adjusted by detecting the background light intensity. The fundamental principle of this adaptive mechanism lies in recognizing that the information density and quality of video data differ significantly under different lighting conditions. In strong light environments, image contrast is high and there is relatively little redundant information between adjacent frames, so the frame extraction frequency can be appropriately reduced to decrease the amount of data. However, in low light or environments with drastic changes in lighting, image quality is easily affected, requiring an increase in the frame extraction frequency to ensure that sufficient effective visual information is captured.

[0063] Specifically, background light intensity detection is typically achieved by analyzing the statistical distribution of pixel brightness values ​​in the video stream, potentially involving image processing techniques such as histogram analysis and average brightness calculation. Keyframe extraction, a core technology in video summarization, aims to select the most representative frame sequence from a continuous video stream, ensuring both information integrity and data compression. Extraction strategies may be based on a combination of algorithms, including inter-frame differencing, motion detection, and content change detection.

[0064] Furthermore, after obtaining the keyframes, they are H.265 encoded to generate a second digest feature. H.265, as a next-generation high-efficiency video coding standard, offers higher compression efficiency and better image quality preservation compared to the traditional H.264 standard. Implementing H.265 encoding on edge devices not only significantly reduces data transmission bandwidth requirements but also protects the privacy of the original video data to a certain extent. Core technology modules involved in the encoding process, such as motion estimation, intra-frame prediction, and transform quantization, have been optimized and adapted to the limited edge computing resources. The generated encoded data, serving as the second digest feature, retains the core information of the video content while existing in a highly compressed form, facilitating efficient transmission in edge environments with limited network bandwidth.

[0065] Step S403: Combine the first summary features and / or the second summary features with the collection timestamp to form a structured summary report.

[0066] Among these, the data acquisition timestamp serves as crucial metadata, providing a unified time reference for various sensor data and enabling effective correlation analysis and time-series modeling of sensor data from different spatiotemporal locations. The generation process of the structured summary report involves multiple technical steps, including data format standardization, field definition normalization, and data integrity verification.

[0067] By organizing and encapsulating feature data extracted from different types of sensors according to a predefined data structure, a summary report with a unified interface format is formed. This not only facilitates subsequent data storage and retrieval but also provides convenience for data parsing and processing by upper-layer application systems. This structured design also supports advanced functions such as incremental updates and partial retransmission, further enhancing the system's flexibility and reliability.

[0068] In the above embodiments, specialized processing strategies are employed to address the different characteristics of radio frequency signals and video data, enabling efficient preprocessing and feature extraction of multimodal sensor data at the edge. This significantly reduces data volume and transmission overhead while ensuring data quality, and the structured encapsulation lays a solid foundation for subsequent data fusion and intelligent analysis.

[0069] Reference Figure 5 As one implementation of step S105, the steps of selecting a communication link to transmit the edge preprocessed dataset to the command center based on the terminal status code and real-time network quality data, and updating the task status database, include: Step S501: Poll real-time network quality data to obtain channel quality indicators for at least two communication links; Continuous monitoring of the current available network performance is crucial. Since wireless connections in mobile environments often exhibit high uncertainty (such as signal attenuation due to terrain obstruction and increased bit error rates due to other electromagnetic interference), it is essential to periodically collect key performance parameters for various links. This includes statistical data on actual throughput, latency jitter, and packet loss probability exhibited by different physical layer interfaces such as 5G cellular networks, satellite communication links, and Mesh self-organizing networks. Polling is not a simple timed query mechanism but combines active probing with passive listening. It can estimate link response speed by sending test packets to measure round-trip time (RTT) and indirectly infer link stability using statistical counters reported by the underlying driver.

[0070] Step S502: Select the primary transmission link and backup link from the communication link set based on the anomaly level identifier of the terminal status code and the channel quality index. This step essentially involves implementing a two-factor driven link optimization algorithm. On one hand, the system's robustness depends on understanding the health status of front-end nodes. If a vehicle is in a fault warning state (e.g., low battery or offline camera), it should be prioritized as a low-priority node or protected with additional redundancy measures, even if its location has good external communication conditions. On the other hand, given the diverse link options, maximizing bandwidth utilization while meeting basic service quality requires introducing a scientific and reasonable weighting model for scoring and evaluation.

[0071] Specifically, the three dimensions of packet loss rate, average latency, and instantaneous bandwidth can be assigned different coefficients and then weighted and summed to obtain a comprehensive score for each link. This score is then used in conjunction with predefined priority mapping rules to complete the final path selection. For example, when the terminal is in good condition but the 5G link performs best, it will naturally be the first choice as the backbone channel. In extremely harsh environments, however, satellite links with stronger anti-interference capabilities may need to be sacrificed for some speed to take on the main responsibility, supplemented by Mesh relay to ensure message reachability within a local area.

[0072] Step S503: Split the edge preprocessing dataset into a structured summary report and raw data fragments; Given that different types of data have drastically different requirements for timeliness and completeness, it is necessary to classify them meticulously and adopt differentiated transmission strategies. Key summary reports refer to core intelligence fragments that carry the most important tactical significance. They are small in size but extremely valuable, often requiring only a few KB or even less to fully express target situation, threat level, or other urgent instructions. In contrast, raw data fragments encompass a large amount of uncompressed high-resolution image frame sequences, radar point cloud maps, audio and video streams, etc. While a single fragment may not have intuitive interpretive value, it is indispensable for post-event review and analysis. This separated architecture not only helps alleviate the uneven distribution of pressure on limited links in high-concurrency scenarios but also allows tasks of different priorities to maintain their respective Service Level Agreements (SLAs) while sharing infrastructure.

[0073] Step S504: Transmit the structured summary report to the command center in real time through the main transmission link, and at the same time transmit the raw data fragments asynchronously through the backup link to obtain the transmission result; For the transmission of critical summary reports, given the emphasis on extremely high real-time response, the connectionless UDP protocol is the most suitable carrier. Compared to the connection-oriented TCP protocol, UDP eliminates the three-way handshake process for establishing a virtual circuit, thus significantly shortening the end-to-end response time, making it particularly suitable for pushing alarm messages that are sudden and frequent. To further enhance its transmission reliability, auxiliary metadata such as CRC check fields and sequence numbers are appended to the header of each UDP packet. Once the receiver detects an error, it can quickly request a retransmission without waiting for a timeout mechanism to be triggered.

[0074] As for the backup and upload of raw data fragments, a certain degree of time lag is tolerable, making it more suitable to use a byte-stream-oriented TCP protocol for ordered batch forwarding. Furthermore, an asynchronous working mechanism is specifically designed, meaning that even if the main link is blocked, it will not affect the queuing and writing process of data blocks waiting to be sent in the background buffer pool, truly achieving decoupling between the front-end and back-end.

[0075] Step S505: Update the task progress field and communication quality log in the task status database according to the transmission results.

[0076] Each successful delivery of a summary report is recorded, and the progress flag in the corresponding task record is updated synchronously (e.g., changing from "Ready" to "Data Returned"). Simultaneously, the system accurately records various performance parameters collected during the session into a dedicated log table, including but not limited to the type of link used, the total latency in milliseconds, and the number of packet losses. This approach benefits operations personnel by allowing them to trace historical events and identify potential problems, while also providing valuable experience samples for future version iterations to support machine learning training, thereby driving the entire scheduling engine towards greater intelligence and autonomy.

[0077] In the above implementation, a two-factor linkage judgment model is introduced, and multiple technical means such as the "urgent first, then non-urgent" data decomposition mechanism and the master-slave separation concurrent scheduling strategy are used to effectively improve the overall survivability, scalability and user experience satisfaction of the system.

[0078] Reference Figure 6 As one implementation of step S106, the step of dynamically optimizing the reward function and electronic fence boundary of the reinforcement learning model based on the historical task state database and communication logs includes: Step S601: Extract the classification results of the causes of task execution delay and the associated electronic fence identifiers from the historical task status database; Task execution delay refers to the phenomenon where the actual completion time exceeds the expected planned time, which may be caused by various factors such as path planning errors, resource conflicts, external interference, or equipment failure. An electronic fence, on the other hand, is a virtual geographical boundary used to limit the operational range of moving objects.

[0079] At this stage, by filtering and aggregating the task status logs accumulated during the system's long-term operation, it is possible to accurately locate which tasks experienced delays in specific areas and establish a spatial-event mapping relationship based on the corresponding geofence numbers. For example, if a vehicle frequently experiences response delays after entering a geofence area with ID F-032, that area will become the focus of the next stage of analysis. This process not only relies on structured database field definitions (such as task ID, start time, end time, and region), but also requires an efficient data retrieval mechanism to support large-scale concurrent query requirements.

[0080] Step S602: Analyze the correlation between transmission failure events in the communication log and the spatial distribution of the electronic fence boundary; Since modern intelligent transportation systems commonly use wireless communication as the primary means of interaction, signal stability directly impacts task execution efficiency. In-depth analysis of historical communication logs can identify the locations and timestamps of anomalies such as increased packet loss rates and frequent connection drops, projecting these onto the corresponding electronic fence coordinate system. Subsequently, spatial statistical tools (such as kernel density estimation and hotspot analysis) are used to quantify the spatial clustering of communication anomalies in different areas, thereby determining whether significant regional differences exist.

[0081] Step S603: Based on the spatial distribution correlation and the classification results of the causes of task execution delay, generate the weight adjustment parameters of the reward function of the reinforcement learning model, and update the reward function coefficients of the reinforcement learning model. Specifically, this application proposes a weight adjustment mechanism based on causal attribution, which assigns different attention levels to various delay types (path planning failure, communication interruption, sensor anomaly) and differentiates their weights by statistically analyzing their occurrence frequency within each electronic fence. For example, if a certain type of delay repeatedly occurs in a specific area, it indicates that the area is more sensitive to this type of problem, and its weight in the total reward expression should be increased accordingly to encourage the model to make behavioral decisions that are more conducive to avoiding such risks in the future.

[0082] In this embodiment of the application, the weight adjustment parameters are calculated as follows: ; In the above formula, η is the learning rate, c is the delay class index, and f c The frequency of occurrence of each type of delay event within the electronic fence area is used to classify the root causes of task execution delays into three categories: path planning failure, communication interruption, and sensor anomaly.

[0083] Subsequently, the resulting weight increments Δw c These coefficients are sequentially added to the original w1, w2, w3 to form a new reward function expression r = w1·A + w2·B - w3·C, where each term represents a different evaluation dimension (such as task completion, energy consumption, safe distance, etc.). It should be noted that these new coefficients do not take effect immediately, but are loaded all at once before the start of the next complete scheduling round to ensure the consistency and convergence of the training process.

[0084] Step S604: Based on the spatiotemporal distribution characteristics of historical signal strength sampling values ​​at the electronic fence boundary, calculate the electronic fence boundary expansion vector and reconstruct the electronic fence boundary polygon coordinate set.

[0085] Specifically, electronic fences are usually pre-set as closed polygons, but in reality, due to factors such as terrain undulations and obstacles, some edge areas may have persistent low signal strength areas, causing terminals deployed inside to be unable to upload data or receive instructions normally.

[0086] To address this issue, the original boundary shape needs to be locally expanded. To this end, the original polygon is first divided into several small grid cells of equal area. Then, for each cell, a signal intensity sample sequence is collected over a time window, and its mean μ and variance σ² are estimated accordingly. The expansion operation is triggered only when both meet a preset threshold condition (i.e., σ² > θ1 and μ < θ2). This indicates that the current region exhibits both significant signal fluctuations and a low overall signal strength, suggesting room for improvement. The resulting expansion vector points towards the centroid of the grid, and its magnitude is negatively correlated with the average signal strength, reflecting the principle of expanding the weaker the signal.

[0087] Furthermore, the expansion vectors generated by each of the aforementioned meshes act uniformly on all vertices of the original polygon, causing them to shift outward by a certain distance. Considering that multiple adjacent meshes may jointly push the stretching of the same boundary segment, a convex hull algorithm (such as the Graham scan method) is also introduced to eliminate possible cross-folding phenomena, ensuring that the final graphic still maintains good geometric properties (such as simple connectivity and no self-intersections). In addition, a continuous boundary fusion strategy is adopted to allow the expansion effect between adjacent meshes to transition smoothly and prevent the generation of jagged, abrupt contours.

[0088] In the above implementation, dynamic reward function optimization and adaptive reconstruction of electronic fence spatial structure are realized under the reinforcement learning framework, forming a complete closed-loop control loop from data analysis to policy feedback to environmental perception and optimization, which is especially suitable for application scenarios with high dynamics and high uncertainty.

[0089] Reference Figure 7 As a further implementation of step S502, the step of selecting the primary transmission link and the backup link from the communication link set based on the anomaly level identifier of the terminal status code and the channel quality index further includes: Step S701: Parse the exception level identifier of the terminal status code; After completing local processing, the terminal device abstracts the quality information of the raw sensor data into a 3-bit binary code-based anomaly level identifier. This identifier carries metadata describing whether the current data meets the expected completeness and accuracy. For example, "000" represents data integrity (DATA_INTACT), indicating that all acquisition and verification processes are normal; while "001" corresponds to data verification failure (DATA_CORRUPT), possibly due to noise interference or hardware failure causing errors in critical fields. This status code is not a simple Boolean flag, but a structured field embedded in the frame header or other fixed locations according to a preset protocol format, facilitating the central node to quickly determine the overall reliability of the data to be transmitted.

[0090] Step S702: If the anomaly level identifier indicates that the data is complete, select the 5G or satellite link with the best channel quality index as the main transmission link. Among them, the optimal channel quality index refers to a quantitative scoring system derived by comprehensively considering multiple physical layer parameters. Specifically, for 5G links, RSRP (Reference Signal Received Power) is mainly used to measure the base station signal strength, and a threshold of -110dBm or above is usually set as usable range; while for satellite links, C / N0 (carrier-to-noise ratio density) is relied upon to reflect the ratio of the received useful signal to the background noise, and the typical criterion is greater than 45dB-Hz.

[0091] In addition, end-to-end latency (RTT) can be introduced as a key factor affecting user experience, especially important in low-latency application scenarios such as emergency command. Therefore, the final link score is obtained by weighting the three factors: score = 0.6 * RSRP_norm + 0.3 * (1 / RTT), where RSRP_norm represents the normalized signal power value. The reason for not simply using the strongest signal link is that signal strength alone cannot fully reflect the actual communication performance, especially in the presence of large jitter or asymmetric paths, which may lead to misjudgment.

[0092] Step S703: If the anomaly level identifier indicates a data anomaly, then the satellite link is used as the primary transmission link to transmit the structured summary report, and the Mesh network is used as the backup link to transmit raw data fragments. In cases where an anomaly level identifier indicates data anomaly, the system enters a more complex dual-channel parallel operating mechanism. In this step, the actions of "transmitting structured summary reports via satellite link" and "transmitting raw data fragments via mesh network fragmentation" together constitute a multi-layered redundancy protection strategy.

[0093] Specifically, the former leverages the wide-area coverage of high-orbit satellites to prioritize the uploading of compressed and refined core intelligence content under bandwidth constraints. To improve transmission stability, the UDP-Lite protocol can replace the UDP portion of the traditional TCP / IP stack, allowing for a certain degree of error tolerance without interrupting the connection. Additionally, Reed-Solomon error correction codes are appended to the IP packet header to enhance anti-interference performance. The latter utilizes the temporary topology formed by neighboring nodes within a vehicular ad hoc network (VANET) to move fragmented raw data. This action, based on a time window division mechanism, divides the continuously collected data stream into several independent units for separate transmission. Combined with the OLSR (Optimized Link State Routing Protocol), it automatically finds the optimal forwarding path. If a segment fails to deliver, a NACK mechanism is triggered to re-request until the three-attempt limit is reached.

[0094] Understandably, the two complement each other effectively in terms of spatial and temporal dimensions: satellite links, although slower, are almost unrestricted by geography and are suitable for transmitting critical, small-scale information; while mesh networks, though constrained by propagation distance, can provide higher instantaneous throughput and are suitable for transmitting large amounts of detailed multimedia materials.

[0095] Step S704: Record the packet loss rate and transmission delay of the primary transmission link / backup link into the communication quality log.

[0096] By meticulously monitoring each interaction event, historical data on the operational status of each link can be continuously accumulated. The packet loss rate, reflecting the proportion of data loss due to various reasons, is not only a direct indicator of link robustness but also provides input samples for subsequent algorithm optimization. Furthermore, using hardware-level timestamps to accurately capture transmission latency effectively eliminates software-level errors, ensuring that the measured latency values ​​are closer to reality. These statistical data are uniformly packaged into a structured document in JSON format and stored, facilitating both program reading and analysis as well as manual review and traceability.

[0097] In the above implementation, terminal status code parsing is used to realize the perception of data integrity on the edge side. Based on this, dynamic scheduling and fault-tolerant control are carried out in combination with various communication link characteristics, thereby achieving efficient and reliable data backhaul, which improves the data delivery success rate and timeliness of emergency response applications in a distributed sensor network environment.

[0098] This application also discloses a multi-vehicle collaborative command and intelligent dispatch system.

[0099] A multi-vehicle collaborative command and intelligent dispatch system, specifically comprising: The data acquisition module is used to acquire registration request data, including vehicle identifier, real-time location coordinates, and sensor type; The vehicle status matrix output module is used to generate a digital twin command view based on real-time location coordinates and preset geographic information data, and output the vehicle status matrix; wherein, the digital twin command view includes a vehicle location layer, a signal coverage heat map layer, and an electronic fence boundary. The task scheduling module is used to respond to input task instructions, and based on the vehicle state matrix, electronic fence boundaries and sensor types, allocate target vehicles and task sub-objectives through a pre-configured reinforcement learning model, and generate a scheduling instruction set and task state database. The edge preprocessing module is used to control the target vehicle to collect sensor data according to the scheduling instruction set, perform edge preprocessing on the sensor data, and generate an edge preprocessing dataset and terminal status code. The communication transmission module is used to select a communication link to transmit the edge preprocessed dataset to the command center based on the terminal status code and real-time network quality data, and to update the task status database. The model optimization module is used to dynamically optimize the reward function and electronic fence boundary of the reinforcement learning model based on the historical task status database and communication logs.

[0100] The multi-vehicle collaborative command and intelligent dispatch system of this application embodiment can implement any of the above methods, and the specific working process of each module in the system can refer to the corresponding process in the above method embodiment.

[0101] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0102] This application also discloses a computer device.

[0103] Computer equipment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a multi-vehicle collaborative command and intelligent scheduling method as described above.

[0104] This application also discloses a computer-readable storage medium.

[0105] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the methods of a multi-vehicle cooperative command and intelligent dispatching method.

[0106] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0107] It should be noted that the computer device and storage medium in the embodiments of this application are respectively electronic devices and storage media applying the above-described multi-vehicle collaborative command and intelligent dispatch method. Therefore, all embodiments of the above-described multi-vehicle collaborative command and intelligent dispatch method are applicable to the computer device and storage medium, and can achieve the same or similar beneficial effects. For the computer device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple; relevant details can be found in the descriptions of the method embodiments.

[0108] In this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0109] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0110] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A multi-vehicle cooperative command and intelligent dispatching method, characterized in that, The method comprises: acquiring registration request data, including a vehicle identifier, real-time position coordinates, and sensor types; generating a digital twin command view according to the real-time position coordinates and preset geographic information data, and outputting a vehicle state matrix; wherein the digital twin command view comprises a vehicle position layer, a signal coverage heat map layer, and an electronic fence boundary; in response to an input task instruction, assigning target vehicles and task sub-targets based on the vehicle state matrix, the electronic fence boundary, and the sensor types, generating a scheduling instruction set and a task status database through a pre-configured reinforcement learning model; controlling target vehicles to collect sensor data according to the scheduling instruction set, performing edge preprocessing on the sensor data, generating an edge preprocessed data set and a terminal state code; based on the terminal state code and real-time network quality data, selecting a communication link to transmit the edge preprocessed data set to a command center, and updating the task status database; dynamically optimizing the reward function of the reinforcement learning model and the electronic fence boundary according to the historical task status database and the communication log.

2. The method of claim 1, wherein, The step of generating a digital twin command view according to the real-time position coordinates and preset geographic information data, and outputting a vehicle state matrix comprises: acquiring vehicle real-time position coordinates and preset geographic information data including historical signal coverage strength; generating a vehicle position layer according to vehicle real-time position coordinates; generating an initial signal heat map layer based on the historical signal coverage strength, updating the heat value weight distribution through the vehicle real-time position coordinates to obtain an updated initial signal heat map layer; based on the updated initial signal heat map layer, marking the area where the signal strength is lower than the preset strength threshold for a preset time length as an electronic fence boundary to obtain a signal coverage heat map; superimposing the vehicle position layer, the signal coverage heat map, and the electronic fence boundary in a digital map to generate a digital twin command view; calculating a vehicle state weight value based on the spatial relationship between the vehicle real-time position coordinates and the electronic fence boundary; outputting a vehicle state matrix containing a vehicle identifier, real-time position coordinates, and a state weight value.

3. The method of claim 2, wherein, The step of, in response to an input task instruction, assigning target vehicles and task sub-targets based on the vehicle state matrix, the electronic fence boundary, and the sensor types, generating a scheduling instruction set and a task status database through a pre-configured reinforcement learning model comprises: parsing the input task instruction to extract task target coordinates and a task type identifier; calculating a task urgency evaluation value based on the spatial position relationship between the task target coordinates and the electronic fence boundary; inputting the task type identifier, the vehicle state matrix, the sensor type, and the task urgency evaluation value into the pre-configured reinforcement learning model to output a matching mapping relationship between the target vehicle identifier and the task sub-target; generating an anti-deviation path containing a path point sequence and a fault tolerance buffer zone according to the matching mapping relationship and road topology data in the preset geographic information data; integrating the target vehicle identifier, the task sub-target, and the anti-deviation path to generate a scheduling instruction set; A task state data table is created to initialize the binding state of the task sub-targets and the target vehicle identifiers and the progress identification bit.

4. The method of claim 1, wherein, The step of performing edge preprocessing on the sensor data comprises: Wavelet transform is performed on the original radio frequency signals in the sensor data to reduce noise, and the energy values of a preset frequency band are extracted as first summary features; The key frame extraction probability of the original video data in the sensor data is adaptively adjusted, and the key frames are encoded to generate second summary features; The first summary features and / or second summary features are combined with the collection time stamps to form a structured summary report.

5. The method of claim 4, wherein, Based on the terminal state code and real-time network quality data, the step of selecting a communication link to transmit the edge preprocessed data set to the command center and updating the task state database comprises: Polling real-time network quality data to obtain channel quality indicators of at least two communication links; According to the abnormal level identifier of the terminal state code and the channel quality indicator, a main transmission link and a backup link are selected from the communication link set; The edge preprocessed data set is split into a structured summary report and original data fragments; The structured summary report is transmitted in real time to the command center through the main transmission link, and the original data fragments are transmitted asynchronously through the backup link to obtain a transmission result; According to the transmission result, the task progress field and the communication quality log in the task state database are updated.

6. The method of claim 1, wherein, According to the historical task state database and the communication log, the step of dynamically optimizing the reward function of the reinforcement learning model and the boundary of the electronic fence comprises: Extracting task execution delay reason classification results and associated electronic fence identifiers from the historical task state database; Analyzing the spatial distribution correlation between transmission failure events in the communication log and the electronic fence boundary; Based on the spatial distribution correlation and the task execution delay reason classification result, a weight adjustment parameter of the reinforcement learning model reward function is generated, and the reward function coefficient of the reinforcement learning model is updated; According to the spatio-temporal distribution characteristics of the electronic fence boundary based on the historical signal strength sampling values, an electronic fence boundary expansion vector is calculated, and the electronic fence boundary polygon coordinate set is reconstructed.

7. The method of claim 5, wherein, The step of selecting a main transmission link and a backup link from the communication link set according to the abnormal level identifier of the terminal state code and the channel quality indicator further comprises: Analyzing the abnormal level identifier of the terminal state code; If the abnormal level identifier indicates data integrity, select the 5G or satellite link with the best channel quality indicator as the main transmission link; If the abnormal level identifier indicates data anomaly, select the satellite link as the main transmission link to transmit the structured summary report, and select the Mesh network as the backup link to transmit the original data fragments; Record the packet loss rate and transmission delay of the main transmission link / backup link to the communication quality log.

8. A multi-vehicle cooperative command and intelligent dispatching system, characterized in that, The system comprises: A data acquisition module is configured to acquire registration request data, including a vehicle identifier, a real-time position coordinate, and a sensor type; The vehicle state matrix output module is configured to generate a digital twin command view and output a vehicle state matrix according to the real-time position coordinates and preset geographic information data; the digital twin command view includes a vehicle position layer, a signal coverage heat map layer, and an electronic fence boundary; The task scheduling module is configured to respond to an input task instruction, assign a target vehicle and a task sub-target based on the vehicle state matrix, the electronic fence boundary, and a sensor type, generate a scheduling instruction set and a task state database, and distribute the scheduling instruction set to the target vehicle through a preconfigured reinforcement learning model. The edge preprocessing module is configured to control the target vehicle to collect sensor data according to the scheduling instruction set, perform edge preprocessing on the sensor data, and generate an edge preprocessing data set and a terminal state code. The communication transmission module is configured to select a communication link to transmit the edge preprocessing data set to a command center based on the terminal state code and real-time network quality data, and update the task state database. The model optimization module is configured to dynamically optimize a reward function of the reinforcement learning model and the electronic fence boundary according to a historical task state database and a communication log.

9. A computer device, characterized by: A computer program product is provided, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the program to implement the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program product is provided, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the program to implement the method of any one of claims 1 to 7.