External light route driven end cloud multi-modal perception method and system and storage medium
By using an external light router to drive an edge-cloud multimodal perception method, control tokens are generated by the external light router to achieve edge-cloud collaborative perception. This solves the problem of low accuracy of inference results from edge devices and improves the accuracy and consistency of inference under latency constraints.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AEROSPACE AGE LOW AERIAL TECHNOLOGY CO LTD
- Filing Date
- 2026-05-21
- Publication Date
- 2026-06-19
Smart Images

Figure CN122247921A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an edge-cloud multimodal sensing method, system, and storage medium driven by an external lightweight router. Background Technology
[0002] Currently, with the development of scenarios such as aerial equipment (e.g., drones) inspection, urban governance, video surveillance, mobile terminal intelligent analysis, industrial visual quality inspection, and vehicle-mounted assisted perception, edge devices are increasingly undertaking complex visual perception tasks. To address this, related technologies typically deploy lightweight models on edge devices to achieve fast inference with low latency and low cost, resulting in relatively low accuracy of the inference results. Therefore, there is currently no satisfactory solution for improving the accuracy of inference results while meeting latency requirements. Summary of the Invention In view of this, embodiments of the present invention provide an edge-cloud multimodal perception method, system, and storage medium driven by an external light router to solve the problem of low accuracy of inference results caused by related technologies. Specifically, embodiments of the present invention can determine the target control token through task semantic information, and realize edge-cloud multimodal perception driven by an external light router according to the path pattern indicated by the target control token. This allows the construction of a unified application-layer control object through the control token, and then the realization of specific control semantics for edge-cloud collaborative perception through the unified application-layer control object. This enables more precise control of edge-cloud multimodal perception (i.e., edge-cloud collaborative perception) to meet constraints such as latency. Furthermore, edge-cloud multimodal perception driven by an external light router can determine highly accurate target inference results. Thus, embodiments of the present invention can effectively improve the accuracy of inference results while meeting latency requirements. Moreover, embodiments of the present invention can transform the target control token into an execution control product that can be directly consumed and followed by the execution side, thereby effectively avoiding deviations between execution behavior and scheduling intent, and improving the consistency, interpretability, and controllability of results in the edge-cloud collaborative process.
[0003] According to one aspect of the present invention, an edge-cloud multimodal perception method driven by an external light router is provided. The method is applied to a target decision control side within an edge-cloud multimodal perception system driven by an external light router, wherein the target decision control side is an edge device or a cloud within the system. The method includes: Obtain routing decision indication data corresponding to the target processing task, and obtain the task semantic information of the target processing task; Based on the routing decision indication data, runtime context indication data is determined; and based on the runtime context indication data and the task semantic information, an initial control token is generated; wherein, the initial control token is generated through an external light router, which is located outside the multimodal large model; Based on the initial control token, a target control token is determined. A control token includes at least one of at least a preset path pattern, which includes at least two of the following: a fast path, a slow path, and a cascaded composite path. Based on the target control token, an execution control product is determined. The fast path is executed through the end-side device, the slow path is executed through the cloud, and the cascaded composite path is executed jointly by the end-side device and the cloud. Based on the path pattern indicated by the target control token, the target execution side of the target processing task is determined; wherein, the target execution side includes the end-side device and / or the cloud; Based on the execution control output, the target execution side determines the target inference result of the visual data to be processed under the target processing task.
[0004] According to another aspect of the present invention, an edge-cloud multimodal perception system driven by an external light router is provided. The external light router-driven edge-cloud multimodal perception system includes an edge device and a cloud, wherein the edge device or the cloud is the target decision control side of the external light router-driven edge-cloud multimodal perception system; wherein, The target decision control side is used to obtain routing decision indication data corresponding to the target processing task, and to obtain the task semantic information of the target processing task; The target decision control side is also used to determine the running context indication data based on the routing decision indication data; and to generate an initial control token based on the running context indication data and the task semantic information; wherein the initial control token is generated through an external light router, which is located outside the multimodal large model; The target decision control side is further configured to determine a target control token based on the initial control token. A control token includes at least one path mode from at least one preset path mode, and the at least one preset path mode includes at least two of the following: fast path, slow path, and cascaded composite path. Based on the target control token, the execution control product is determined. The fast path is executed through the end-side device, the slow path is executed through the cloud, and the cascaded composite path is executed jointly by the end-side device and the cloud. The target decision control side is further configured to determine the target execution side of the target processing task based on the path pattern indicated by the target control token; wherein the target execution side includes the end-side device and / or the cloud. The target decision control side is also used to enable the target execution side to determine the target reasoning result of the visual data to be processed under the target processing task based on the execution control output.
[0005] According to another aspect of the present invention, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods mentioned above.
[0006] This invention can acquire routing decision indication data corresponding to a target processing task, as well as task semantic information of the target processing task. Based on the routing decision indication data, runtime context indication data can be determined. An initial control token can be generated based on the runtime context indication data and task semantic information. The initial control token can be generated via an external light router located outside the multimodal large model. Then, based on the initial control token, a target control token can be determined. Each control token includes at least one of at least a preset path pattern, which includes at least two of the following: fast path, slow path, and cascaded composite path. Based on the target control token, the execution control output can be determined. The fast path is executed via the edge device, the slow path is executed via the cloud, and the cascaded composite path is executed jointly by the edge device and the cloud. Correspondingly, based on the path pattern indicated by the target control token, the target execution side of the target processing task can be determined. The target execution side includes the edge device and / or the cloud. Based on this, the target execution side can determine the target inference result of the visual data to be processed under the target processing task based on the execution control output. As can be seen, embodiments of the present invention can determine the target control token through task semantic information, and realize edge-cloud multimodal perception driven by an external light router according to the path pattern indicated by the target control token. That is, a unified application layer control object can be constructed through the control token, and then the specific control semantics of edge-cloud collaborative perception can be realized through the unified application layer control object, so as to more accurately control edge-cloud multimodal perception (i.e., edge-cloud collaborative perception) to meet constraints such as latency. Furthermore, the edge-cloud multimodal perception driven by an external light router can determine the target inference result with higher accuracy. That is, embodiments of the present invention can effectively improve the accuracy of inference results while meeting latency requirements. Moreover, embodiments of the present invention can transform the target control token into an execution control product that can be directly consumed and followed by the execution side, thereby effectively avoiding the deviation between execution behavior and scheduling intention, and improving the consistency, interpretability and controllability of results in the edge-cloud collaborative process. Attached Figure Description
[0007] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart illustrating an edge-cloud multimodal perception method driven by an external light router according to an exemplary embodiment of the present invention is shown. Figure 2 A flowchart illustrating another edge-cloud multimodal perception method driven by an external light router according to an exemplary embodiment of the present invention is shown. Figure 3 A schematic diagram of a control token according to an exemplary embodiment of the present invention is shown; Figure 4 A schematic diagram of a closed-loop optimization according to an exemplary embodiment of the present invention is shown; Figure 5 A flowchart illustrating another edge-cloud multimodal perception method driven by an external light router according to an exemplary embodiment of the present invention is shown. Figure 6 A schematic block diagram of an edge-cloud multimodal sensing system driven by an external light router is shown according to an exemplary embodiment of the present invention. Detailed Implementation
[0008] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0009] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0010] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0011] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0012] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0013] It should be noted that the embodiments of the present invention can provide an edge-cloud multimodal perception system driven by an external light router. The edge-cloud multimodal perception system driven by an external light router can be used to execute the edge-cloud multimodal perception method driven by an external light router. That is to say, the execution subject of the edge-cloud multimodal perception method driven by an external light router provided in the embodiments of the present invention can be the edge-cloud multimodal perception system driven by an external light router, such as the target decision control side included in the edge-cloud multimodal perception system driven by an external light router. The edge-cloud multimodal perception system driven by an external light router can include an edge device (also referred to as edge) and a cloud. The target decision control side can be an edge device or a cloud. The embodiments of the present invention do not limit this. Optionally, the edge side can refer to the near-end execution side located closer to the data source relative to the cloud. For example, it can serve as a near-end processing side for visual data access, fast path task execution (also known as fast path task), candidate attention region generation, local implementation of control tokens, and collaborative interaction with the cloud. It can include device terminals and / or edge terminals. The cloud can refer to cloud-side nodes or cloud service sides that provide remote inference, complex semantic understanding, slow path task execution (also known as slow path task), etc. For example, it can undertake tasks such as multimodal large model inference, complex scene understanding, slow path task execution, global evidence processing, and policy updates.
[0014] Optionally, the device terminal can be a terminal device located close to the data source, directly undertaking data acquisition and / or local processing functions, such as an airborne terminal, smart camera, vehicle-mounted terminal, mobile terminal, etc. In some embodiments, it can perform functions such as visual data acquisition, basic preprocessing, local fast path inference, or candidate region of interest generation, and can serve as one implementation form on the edge side. Correspondingly, the edge terminal can be a computing node deployed near the site, undertaking near-end inference, control orchestration, or data relay functions, such as an edge box, edge gateway, edge server, etc., and can serve as the near-end deployment location in this embodiment of the invention. For example, the edge terminal can undertake the deployment and operation of modules such as an external lightweight router, fast path lightweight task head, budget executor, and control token adapter; or, the cloud can also undertake the deployment and operation of modules such as an external lightweight router, etc.; this embodiment of the invention does not limit this.
[0015] Based on this, the edge device may include one or more electronic devices, which may be terminals (i.e., clients) or servers; optionally, the terminals mentioned herein may include, but are not limited to: smartphones, tablets, laptops, desktop computers, airborne terminals, smart cameras, vehicle terminals, etc.; the servers mentioned herein may be independent physical servers, or server clusters or distributed systems composed of multiple physical servers, etc. Optionally, edge-cloud collaboration may refer to the collaborative processing between the edge and the cloud (also known as the cloud side) around visual data processing, path selection, control and governance, result verification, and strategy optimization; based on this, the overall operating mode of the embodiments of the present invention can be constituted, enabling fast paths (also known as rapid paths) and slow paths (also known as slow-speed paths) to achieve on-demand collaboration and closed-loop optimization under various constraints, etc.
[0016] Optionally, the edge-cloud multimodal perception method driven by an external light router provided in this embodiment of the invention can be applied to any inspection and processing scenario (also referred to as an edge-cloud multimodal perception scenario driven by an external light router), and this embodiment of the invention does not limit it. For example, the inspection and processing scenario can be an urban governance scenario, an emergency management scenario, a traffic management scenario, etc.
[0017] Based on the above description, this embodiment of the invention proposes an edge-cloud multimodal perception method driven by an external light router. This edge-cloud multimodal perception method driven by an external light router can be executed by the target decision control side included in the aforementioned edge-cloud multimodal perception system driven by an external light router. That is, the edge-cloud multimodal perception method driven by an external light router can be applied to the target decision control side included in the edge-cloud multimodal perception system driven by an external light router, where the target decision control side is the edge device or the cloud in the edge-cloud multimodal perception system driven by an external light router. Figure 1 As shown, the edge-cloud multimodal awareness method driven by the external light router may include the following steps S101-S105: S101, obtain the routing decision indication data corresponding to the target processing task, and obtain the task semantic information of the target processing task.
[0018] Optionally, the target processing task can be any task, and this embodiment of the invention does not limit it. Optionally, the routing decision indication data corresponding to the target processing task may include the perceived visual data and / or the edge-side rapid perception results of the perceived visual data under the target processing task, etc., and this embodiment of the invention does not limit it. Optionally, the perceived visual data under the target processing task may be the visual data to be processed under the target processing task, or it may be the historical visual data under the target processing task (i.e., it may be the video data collected before the visual data to be processed), or it may be the preset visual data under the target processing task, etc., and this embodiment of the invention does not limit it. Optionally, the preset visual data may be set according to experience or actual needs, and this embodiment of the invention does not limit it. Optionally, a visual data may be a single image, or a continuous video frame, or it may be a set of keyframes, etc., and this embodiment of the invention does not limit it.
[0019] Optionally, the aforementioned edge-side rapid perception results may include, but are not limited to, at least one of the following: candidate regions of interest (such as regions where targets are detected) in the perceived visual data, preliminary category detection results (such as target category detection results), confidence level, uncertainty measure, and difficulty score, etc.; this embodiment of the invention does not limit these. Optionally, the edge-side rapid perception results of the perceived visual data may be obtained by processing the perceived visual data through an edge-side target perception model (which may include, but is not limited to, at least one of the following: lightweight detection model, lightweight classification model, instance segmentation model, OCR (Optical Character Recognition) model, lightweight tracking model, and lightweight verification model, etc., such as any lightweight large model, etc.), such as inputting the perceived visual data into the target perception model to output edge-side rapid perception results through the target perception model, etc.; it should be noted that this embodiment of the invention does not limit the specific model architecture of the target perception model. Optionally, the confidence level in the edge-side fast perception result can be the average of the confidence levels of each target (i.e., the detected target, such as the target indicated by the candidate region of interest) in the perceived visual data, or it can include the confidence level of each target, or the confidence level of each image (such as a video frame) in the perceived visual data (the confidence level of an image can be the average of the confidence levels of each target in the corresponding image), etc.; this embodiment of the invention does not limit this. Optionally, the uncertainty measure in the edge-side fast perception result can be the average of the uncertainty measures of each target in the perceived visual data (i.e., the confidence level of the perceived visual data), or it can include the uncertainty measure of each target, or it can include the uncertainty measure of each image in the perceived visual data (the uncertainty measure of an image can be the average of the uncertainty measures of each target in the corresponding image), etc.; this embodiment of the invention does not limit this. Optionally, the difficulty score in the edge-side rapid perception result can be the difficulty score of the perceived visual data. For example, it can be determined based on the confidence and / or uncertainty measure in the edge-side rapid perception result, or it can be directly output by the target perception model. The specific method of determining the difficulty score is not limited in the embodiments of the present invention. For example, the higher the confidence, the lower the difficulty score, the higher the uncertainty measure, the higher the difficulty score, and so on.
[0020] Optionally, the target decision control side may store routing decision indication data in its own storage space, and the routing decision indication data can be obtained from its own storage space; or, a routing decision indication data download link can be obtained to download the routing decision indication data, thereby achieving the acquisition of routing decision indication data; or, when the target decision control side is in the cloud, it can also receive routing decision indication data sent by the end-side device to achieve the acquisition of routing decision indication data, etc.; the embodiments of the present invention do not limit this. Optionally, the visual data to be processed can be collected by the device terminal, or it can be determined from the task execution instruction for the visual data to be processed under the target processing task. The task execution instruction can carry the visual data to be processed, and the task execution instruction can be detected based on the task execution operation. The visual data to be processed indicated by the task execution instruction can be the visual data to be processed indicated by the task execution operation, etc.; the embodiments of the present invention do not limit this. It should be noted that the embodiments of the present invention do not limit the specific implementation of the task execution operation.
[0021] Optionally, the target decision control side may store task semantic information of the target processing task in its own storage space. In this case, the task semantic information can be obtained from its own storage space. Alternatively, when a task semantic information setting operation for the target processing task is detected, the task semantic information indicated by the task semantic information setting operation can be used as the task semantic information of the target processing task to obtain the task semantic information of the target processing task, etc. This embodiment of the invention does not limit this aspect. It should be noted that this embodiment of the invention does not limit the specific implementation method of the task semantic information setting operation. For example, the task semantic information may come from a preset task template, task identifier, natural language prompt, or business template, etc. It should be noted that this embodiment of the invention does not limit the specific content of the task semantic information.
[0022] Optionally, the edge device and / or cloud may include a data access and preprocessing module to acquire image or video frames and perform preprocessing such as decoding, frame extraction, scale normalization, and quality screening, while associating task semantic information. For example, the edge device may acquire routing decision indication data through the data access and preprocessing module, or the edge device may acquire visual data to be processed through the data access and preprocessing module, etc.; the embodiments of the present invention do not limit this.
[0023] S102, based on the routing decision indication data, determine the running context indication data; and based on the running context indication data and task semantic information, generate the initial control token.
[0024] The initial control token can be generated via an external light router, which can be located outside the multimodal large model, such as outside the target multimodal large model in the cloud, etc. In this embodiment of the invention, the target decision control side may include an external light router, which can generate the initial control token based on runtime context indication data and task semantic information. Optionally, the aforementioned runtime context indication data can serve as an auxiliary basis for control token generation and / or budget governance.
[0025] Optionally, the runtime context indication data may include, but is not limited to, at least one of the following: edge-side rapid perception results, system resource status (including the resource status of the edge device and / or the resource status of the cloud), network status (such as the network status between the edge device and the cloud), cache status (including the cache status of the edge device and / or the cache status of the cloud), energy consumption status (including the energy consumption status of the edge device and / or the energy consumption status of the remote end), temperature status (including the temperature status of the edge device and / or the temperature status of the cloud), historical result consistency information, and queue length (such as the task execution queue length of the target processing task), etc.; this embodiment of the invention does not limit this. Optionally, historical result consistency information may be category detection consistency information of the same target (such as the probability of the same category), target bounding box detection consistency information of the same target (such as the probability of the target bounding box being consistent), or historical inference result consistency information, etc.; this embodiment of the invention does not limit this.
[0026] Optionally, when determining the operating context indication data based on the routing decision indication data, if the routing decision indication data includes perceived visual data, the target decision control side can call the target perception model to determine the edge-side rapid perception result of the perceived visual data, thereby determining the operating context indication data based on the edge-side rapid perception result; or, if the routing decision indication data includes the edge-side rapid perception result of the perceived visual data, then the edge-side rapid perception result can be determined from the routing decision indication data, thereby determining the operating context indication data based on the edge-side rapid perception result, and so on; this embodiment of the invention does not limit this. Optionally, when the target decision control side is the cloud, in the process of calling the target perception model to determine the edge-side rapid perception result of the perceived visual data, the cloud can call the same perception model as the target perception model on the edge to determine the edge-side rapid perception result of the perceived visual data; or, the cloud can send the perceived visual data to the edge device, then the edge device can call the target perception model to determine the edge-side rapid perception result of the perceived visual data and return the edge-side rapid perception result to the cloud, and so on; this embodiment of the invention does not limit this.
[0027] Optionally, when determining the runtime context indication data based on the edge-side rapid sensing results, the target decision control side can obtain the current runtime environment indication data and use the edge-side rapid sensing results and the current runtime environment indication data to determine the runtime context indication data; or, the edge-side rapid sensing results can be used as the runtime context indication data, etc.; this embodiment of the invention does not limit this. For example, when determining the runtime context indication data using the edge-side rapid sensing results and the current runtime environment indication data, the edge-side rapid sensing results and the current runtime environment indication data can be added to the runtime context indication data, thus concatenating the edge-side rapid sensing results and the current runtime environment indication data to determine the runtime context indication data; or, the runtime context template can be filled (i.e., field filling) according to the edge-side rapid sensing results and the current runtime environment indication data to obtain the runtime context indication data, etc.; this embodiment of the invention does not limit this. Optionally, the runtime context template can be set according to experience or actual needs, this embodiment of the invention does not limit this.
[0028] Optionally, the current operating environment indication data may include, but is not limited to, at least one of the following: system resource status, network status, cache status, energy consumption status, temperature status, historical result consistency information, and queue length, etc., and the embodiments of the present invention do not limit this.
[0029] Optionally, the aforementioned runtime context indication data and task semantic information can be used as decision inputs for an external light router (also referred to as an external light router module). In this embodiment of the invention, the external light router can be a routing decision module (such as a lightweight routing decision module) located outside the multimodal large model. It can be used to generate control tokens (i.e., structured control tokens) by combining task semantic information and runtime context information (i.e., runtime context indication data), thus being responsible for completing control decisions before actual inference execution. Optionally, the external light router can be composed of one or more of the following: a rule system, a state machine, a lightweight learnable model, a small Transformer (such as a neural network structure based on a self-attention mechanism), a small language model, and a small multimodal model, etc. This embodiment of the invention does not limit the specific construction of the external light router, that is, it does not limit the rule system, lightweight learnable model, etc. For example, a rule system or state machine can generate control tokens using a preset token generation template; a lightweight learnable model can output scores for each path mode, budget risk scores, and suggested values for control fields, which are then used by a template generator to form control tokens, and so on; optionally, the preset token generation template and template generator can be set according to experience or actual needs, and this embodiment of the invention does not limit this. It should be understood that since the target decision control side is an edge device or the cloud, and the edge device includes a device terminal and / or an edge terminal, the external lightweight router can be deployed at the edge terminal, or at the device terminal or cloud control plane, etc., but all are located outside the multimodal large model, etc.
[0030] Optionally, the external lightweight router may include, but is not limited to, at least one of the following: a rule / threshold routing submodule (also known as a rule routing submodule, which can be used to provide one or more of Service Level Agreement (SLA) guardrails (also known as rule guardrails), safety paths (such as fast paths), and out-of-bounds fallback conditions (such as timeout fallback), a fusion routing submodule (also known as a lightweight learnable routing submodule, which can be used to perform feature modeling (i.e., fusion analysis) on task semantic information and runtime context information, and output one or more of the following: scores for each path pattern, budget risk scores, and control field suggestion information, such as fusion analysis through a lightweight fusion analysis model), an arbitration and pruning submodule (also known as an adjustment submodule, which can be used to arbitrate (such as pruning or replacement) when a pending control token conflicts with a rule guardrail, or to shrink the proposed control parameters in advance when it is predicted that the initial control token will trigger a budget overrun), and a memory retrieval submodule (which can be used to recall policy knowledge entries in similar states from the routing memory and serve as a priori source for initializing or adjusting control token fields), etc.; the embodiments of the present invention do not limit this. Optionally, the proposed control parameters can be set based on experience or actual needs, and this embodiment of the invention does not limit this; for example, the proposed control parameters may include path patterns, such as reverting the path pattern to a fast path, etc. Optionally, the lightweight fusion analysis model may include, but is not limited to, at least one of the following: lightweight learnable model, small Transformer, small language model, small multimodal routing model, etc., and this embodiment of the invention does not limit this.
[0031] Optionally, the rule guardrail may include, but is not limited to, at least one of the following: maximum end-to-end latency rule, minimum execution success rate rule, result stability rule, maximum cloud call frequency rule, maximum bandwidth usage rule, maximum end-side temperature limit rule, end-side power consumption limit rule, and abnormal rollback rule, etc., and this embodiment of the invention does not limit this; Optionally, one task may correspond to one rule guardrail, and the rule guardrails (i.e., the corresponding rule guardrails) under different tasks may be the same or different, and this embodiment of the invention does not limit this; Optionally, the rule guardrail under a task may be set according to experience or according to actual needs, and this embodiment of the invention does not limit this. Optionally, the out-of-bounds rollback condition may be set according to experience or according to actual needs, and this embodiment of the invention does not limit this. The control field suggestion information can be used to guide the suggested adjustments to the fields involved in the control token, such as guiding how to trim or replace which fields, etc.
[0032] Optionally, the control token (also referred to as a control Token or control object) can be a protocol-based data object (i.e., a protocol-based control object) that follows a predefined and version-evolving structure. It can uniformly carry control semantics such as runtime level, path pattern, region of interest, keyframe strategy, control vector, payload policy (also referred to as payload_policy, payload constraint, payload form, etc.), and version information, serving as the core control object throughout decision-making, governance, execution, and auditing. Optionally, the control token supports version evolution.
[0033] Optionally, a control token may include, but is not limited to, at least one of the following: a header, a body, and a trailer, which are not limited in this embodiment of the invention. The header may be a pre-entry field section of the control token, used to describe the basic identity, version, and runtime context of the control token, providing a unified entry point for subsequent parsing, routing, governance, and tracing; the body may be a core field section of the control token, used to carry the control semantics that this embodiment of the invention actually applies to the execution chain, and is the main object of budget governance, semantic compilation, and path execution; the trailer may be a post-entry field section of the control token, used to record the summary, parent-child relationship, and integrity information of the control token, supporting derived token chain tracing, evidence solidification, and auditability.
[0034] Optionally, the header of a control token may include, but is not limited to, at least one of the following: schema version (also represented as schema_version, which can be used to identify the schema version followed by the control token), token identifier (also represented as token_id, which can be used to uniquely identify the control token), and profile identifier (also represented as profile_id, which can be used to identify the profile), etc., which are not limited in this embodiment of the invention. Optionally, a profile can be used to indicate at least one of the following: a running permission, computing power level, and response policy, etc., which are not limited in this embodiment of the invention. The schema can be used to define the data structure conventions for the control token's field composition, field semantics, hierarchical structure, serialization rules, and version rules, thereby providing a structural foundation for unified parsing, unified governance, unified compilation, and unified auditing of control tokens, ensuring that all modules understand the control token consistently.
[0035] Optionally, the subject of a control token may include, but is not limited to, at least one of the following: path mode (also represented as path_mode), region of interest definition (also represented as roi_spec, which can be used to indicate the range of the region of interest (i.e., the region of interest)), keyframe policy (also represented as keyframe_policy, such as collecting one keyframe every 5 seconds or selecting 2 keyframes from a visual data), control vector (also represented as control_vector), payload policy (also represented as payload_policy, such as uploading pixels or intermediate features of the region of interest), and model version information (also represented as model_versions, which can be used to indicate the model version used for inference), etc.; the embodiments of the present invention do not limit this. The Region of Interest (ROI) is a local area in visual data that requires focused processing, analysis, or review. It can be used to limit the scope of interest for fast and slow paths, reduce processing overhead for irrelevant areas, and improve the targeting of local review and budget governance. The control vector is a set of parameters consisting of at least one manageable control parameter (i.e., including parameter information of each manageable control parameter in at least one manageable control parameter). It can be used to uniformly express control information affecting path selection, execution scope, payload form, and execution configuration. The execution scope includes, but is not limited to, at least one of the following: ROI retention range (i.e., the scope of the region of interest), keyframe range, whole frame / local execution range (which can be used to indicate whether to execute the whole frame or locally, etc.), whether slow path review is allowed, etc. The execution configuration includes, but is not limited to, at least one of the following: input resolution, frame sampling frequency, compression level, model / operator version, execution graph configuration, and post-processing threshold (i.e., parameter information), etc. For example, at least one manageable controllable parameter may include, but is not limited to, at least one of the following: slow path trigger threshold, upper limit of the number of ROIs, upper limit of the area of ROIs, ROI expansion boundary (also known as expansion coefficient, such as how many times the region of interest is expanded), frame sampling frequency, Top-K (first K, K is a positive integer) retention number, input resolution, compression level, load strategy, model / operator version, execution graph configuration, and post-processing threshold, etc., which are not limited in this embodiment of the invention. Optionally, intermediate features may include, but are not limited to, at least one of the following: edge features of the region of interest, texture grayscale features, attention weight features, multi-head self-attention mapping features, feature map channel dimension features, intermediate representation after positional encoding fusion, spatial structure features, pixel statistical features, etc., which are not limited in this embodiment of the invention.
[0036] Optionally, the tail of a control token may include, but is not limited to, at least one of the following: control token digest (also referred to as token_hash), parent token digest (also referred to as parent_token_hash), control token and derived control token chain digest (also referred to as token_hash_chain, also known as control token chain digest, i.e., the digest of the control token chain), optional signature, etc., but the embodiments of the present invention do not limit this. The control token digest can be a digest value calculated by normalizing and serializing the fields of the control token excluding the control token digest itself and the optional signature field. For example, it can be first sorted, formatted, separated, and encoded according to preset normalization rules (to achieve normalization), then concatenated into an ordered string according to a preset concatenation order (to achieve serialization), and finally fused representation calculation to obtain a unique digest value, etc. The parent token digest can be the token_hash of the parent control token that generated the current control token, which can be used to represent the version evolution relationship between the original control token (i.e., the initial control token), the derived control token, and the target control token. The control token and derived control token chain digest can be a sequence of original control token digests and derived control token digests arranged in the generation order, etc., which can be used to reconstruct the budget governance and token derivation process, etc. Optionally, the preset normalization rules and preset concatenation order can be set according to experience or actual needs, and the embodiments of the present invention do not limit this. It should be noted that the embodiments of the present invention do not limit the method of determining any digest value; for example, any hash algorithm can be used for fused representation calculation to obtain a unique digest value (which is the hash value at this time), etc.
[0037] Optionally, a policy knowledge entry may include, but is not limited to, at least one of the following: state feature summary (state_features), control token template (token_template), effect statistics (effect_stats), evidence span (evidence_span), and version (version), etc., and the embodiments of the present invention do not limit this. Optionally, `state_features` may include, but is not limited to, at least one of the following: task semantic summary, task priority, number of candidate ROIs, fast path confidence statistics, uncertainty statistics, current bandwidth range, edge load range, temperature range, historical consistency features, and cache hit status; `effect_stats` can be used to record historical execution effect statistics, which may include, but is not limited to, at least one of the following: accuracy, recall, false positive rate, average latency, P95 latency (i.e., 95th percentile latency), cloud call rate, uplink bytes, budget out-of-bounds rate, failure rate, and manual review pass rate; `evidence_span` can be used to describe the range of evidence samples on which the policy knowledge entry is based, which may include, but is not limited to, at least one of the following: sample quantity, time window, task type, device range, and scenario range; `version` can be used to identify the version of the policy knowledge entry or the overall routing memory.
[0038] S103, based on the initial control token, determine the target control token. A control token includes at least one of the preset path modes. The preset path mode includes at least two of the following: fast path, slow path, and cascaded composite path. Based on the target control token, determine the execution control artifact. The fast path is executed through the end device, the slow path is executed through the cloud, and the cascaded composite path is executed jointly by the end device and the cloud.
[0039] In this context, a fast path can also be represented as FAST, a slow path as SLOW, and a cascaded composite path as a cascaded composite path. Based on this, a path mode can be any of at least one preset path mode. At least one preset path mode can include, but is not limited to, at least two of the following: fast path, slow path, and cascaded composite path (also known as cascaded verification path), etc. This embodiment of the invention does not limit these. For example, at least one preset path mode can also include, but is not limited to, at least one of the following: multi-level slow paths (such as calling different cloud-based large models to execute tasks) and task-specific paths (such as one task corresponding to one task-specific path, the task-specific path corresponding to one task can be set according to experience or actual needs), etc.
[0040] Optionally, a fast path can be used to execute fast path tasks, i.e., when the task mode is fast path, a fast path task can be executed based on the execution control product; a slow path can be used to execute slow path tasks, i.e., when the task mode is slow path, a slow path task can be executed based on the execution control product; a cascaded composite path can be used to execute cascaded composite path tasks, i.e., when the task mode is cascaded composite path, a cascaded composite path task can be executed based on the execution control product, and so on. Optionally, a fast path (i.e., a rapid path) can be a low-cost inference path executed by one or more lightweight task heads, such as tasks that can complete detection, segmentation, optical character recognition, classification, or lightweight inference, and can be used to quickly produce candidate results, uncertainty information, and local priors; a slow path (i.e., a slow path) can be a high-capacity inference path executed by a multimodal large model, and can be used for complex semantic understanding, difficult sample verification, or interpretation generation; a cascaded composite path can be a combination path that first executes a fast path, and then decides whether to call a slow path for verification based on a preset verification trigger condition, and can be used to balance processing efficiency and result accuracy under controlled resource budget. Optionally, the preset review trigger conditions can be set according to experience or actual needs, and this embodiment of the invention does not limit this. Optionally, the edge device may include a fast path execution module to execute the fast path; and / or, the cloud may include a slow path execution module to execute the slow path, etc. Optionally, the lightweight task head of the fast path may be one or more of the following: target detection head, instance segmentation head, OCR (Optical Character Recognition) head, classification head, lightweight tracking head, or a combination thereof; the target multimodal large model of the slow path may be a privately deployed model, a public cloud service, or an edge cloud service, etc.; this embodiment of the invention does not limit this.
[0041] Optionally, when determining the execution control product based on the target control token, the target control token can be semantically compiled using a control token adapter to obtain the execution control product. The execution control product includes control channel parameters and / or control fragments, enabling it to be consumed by the target execution side. The control fragment may include, but is not limited to, at least one of the following: region of interest, keyframe strategy, payload strategy, and verification intent (which can be used to control the model to verify specified content). The control fragment can be injected into the slow path inference link to constrain the inference of the target multimodal large model (e.g., constraining the processing range and processing method). That is, the control fragment can be inserted or spliced as a structured fragment into the input sequence of the target multimodal large model to explicitly constrain the slow path inference process (i.e., the slow path inference link). For example, the control fragment may also include task boundaries, region of interest descriptions, keyframe ranges, verification intents, output format requirements, prohibited processing regions, or confidence interpretation requirements, used to constrain the region of interest, determine the target, and define the output format in the input link of the slow path multimodal large model. It should be noted that the specific implementation of semantic compilation in the embodiments of the present invention is not limited. For example, semantic compilation can be performed through preset mapping rules, templated transformation rules, rule engines, lightweight compilers or semantic encoding models, or it can be performed through contextual semantic association, projection transformation, etc.
[0042] Optionally, control channel parameters can be used to explicitly constrain one or more of the following at the system interface layer or orchestration layer: execution range, load type, path mode, model version, and post-processing parameters. For example, control channel parameters may include, but are not limited to, at least one of the following: ROI coordinates, ROI mask reference, input resolution, load type, model version, execution graph identifier, post-processing threshold, maximum number of ROIs, slow path trigger threshold, whether slow path review is allowed, keyframe strategy, and load strategy, etc., which are not limited in this embodiment of the invention. Optionally, control channel parameters can be passed to the target execution side as HyperText Transfer Protocol Header (HTTP Header), general Remote Procedure Call Metadata (gRPC Metadata), request parameters, interface parameters, or bypass metadata, etc.
[0043] Optionally, the target decision control side may also include a control token adapter (also referred to as a Token Adapter or Control Token Adapter Module), which can semantically compile the target control token to obtain execution control artifacts. Accordingly, the control token adapter can semantically compile the target control token into control channel parameters and / or control fragments that can be directly consumed by the target execution side.
[0044] Optionally, the control token adapter can also generate a compilation digest (also represented as `compile_hash`) of the execution control product after semantic compilation is completed. The compilation digest serves as an identifier for the digested representation of the execution control product; it can be used to verify the consistency of the execution control product and can serve as one of the key indexes for evidence chain writing, result reconciliation, and subsequent reproduction and location, such as serving as an index for the execution control product, etc. In this embodiment of the invention, the execution control product and the compilation digest can provide direct constraints for subsequent path execution. It should be noted that this embodiment of the invention does not limit the specific generation method of the compilation digest, that is, it does not limit the specific representation form of the compilation digest. For example, the compilation digest can be a digest value calculated by a hash algorithm or message digest algorithm after normalizing and serializing the control channel parameters and / or control fragments output by the control token adapter, used to identify a semantic compilation result and verify the consistency of the execution control product. It should be noted that this embodiment of the invention does not limit the hash algorithm, message digest algorithm, and serialization method used to generate the compilation digest.
[0045] Optionally, the target decision control side can trigger the execution of obtaining the routing decision indication data corresponding to the target processing task and obtaining the task semantic information of the target processing task at preset token update intervals, thereby determining the target control token corresponding to the target processing task. This means the target control token can be updated at preset token update intervals. Alternatively, the execution of obtaining the routing decision indication data corresponding to the target processing task can be triggered when a routing memory update is detected, thereby updating the target control token. Alternatively, when the routing decision indication data is visual data to be processed, the target control token can be determined each time based on the current visual data to be processed, etc. This embodiment of the invention does not limit this. Optionally, the preset token update interval can be set according to experience or actual needs, and this embodiment of the invention does not limit this.
[0046] Based on this, embodiments of the present invention can enable routing intent to be further transmitted from the scheduling decision layer (such as the target decision control side) to the execution layer (such as the target execution side) through protocolized and versionable control tokens, as well as the field-level governance of control tokens by the budget executor and the semantic compilation of control token adapters, and form explicit constraints on the execution link in the form of control channel parameters and / or control fragments.
[0047] S104, based on the path pattern indicated by the target control token, determine the target execution side of the target processing task; wherein, the target execution side includes the end device and / or the cloud.
[0048] In this embodiment of the invention, when the path mode indicated by the target control token is a fast path, it can be determined that the target execution side includes the end-side device; when the path mode indicated by the target control token is a slow path, it can be determined that the target execution side includes the cloud; when the path mode indicated by the target control token is a cascaded composite path, it can be determined that the target execution side includes both the end-side device and the cloud.
[0049] S105, based on the execution control product, enables the target execution side to determine the target inference result of the visual data to be processed under the target processing task.
[0050] Optionally, the target execution side may also send the target inference results to the target monitoring platform. The target monitoring platform can be any monitoring platform, and this embodiment of the invention does not limit it.
[0051] Optionally, the target inference result may include, but is not limited to, the task inference result (which may include the inference result obtained by inferring the task inference target for the target processing task) and / or at least one additional inference information, etc., and this embodiment of the invention does not limit this. Optionally, the task inference target of the target processing task may be set according to experience or according to actual needs, and this embodiment of the invention does not limit this; for example, the task inference target of the target processing task may include, but is not limited to, whether traffic congestion occurs. Optionally, at least one additional inference information may include, but is not limited to, at least one of the following: confidence level (such as the confidence level of each keyframe), uncertainty measure (such as the uncertainty measure of each keyframe), difficulty score (such as that which can be obtained by comprehensive analysis based on confidence level and uncertainty measure), category similarity distribution (which can be used to distinguish the degree of difference between different categories, such as that which can be analyzed in all keyframes of the visual data to be processed), cross-frame consistency information (such as cross-frame consistency between the same target), out-of-distribution (OOD) signs, etc., and this embodiment of the invention does not limit this. Optionally, unlike the additional information for inference, the execution process status information may include the actual path mode, fast path time, slow path time, end-to-cloud transmission time, model version, load size, triggering conditions, exception codes, failure reasons, and result fusion methods. The execution process status information can be written into the evidence chain for subsequent reproduction, reconciliation, and accountability verification.
[0052] In one implementation, when the path mode indicated by the target control token is a fast path, the edge device can determine the target inference result of the visual data to be processed under the target processing task according to the execution control product; that is, the edge device can determine the target inference result of the visual data to be processed under the target processing task according to the execution control product. Optionally, when the target decision control side is in the cloud, the target decision control side can also send the execution control product to the edge device. Optionally, the edge device can obtain the visual data to be processed under the target processing task, and call the edge inference model according to the execution control product to determine the target inference result of the visual data to be processed. For example, it can determine the input data of the edge inference model based on the visual data to be processed according to the constraints indicated by the execution control product, and input the input data of the edge inference model into the edge inference model to output the target inference result through the edge inference model, etc.; the embodiments of the present invention do not limit this. Optionally, the edge-side inference model may be the same as or different from the target perception model; this embodiment of the invention does not limit this. It should be noted that this embodiment of the invention does not limit the specific model structure of the edge-side inference model. For example, the edge-side inference model may include, but is not limited to, at least one of the following: a lightweight detection model, an instance segmentation model, an OCR model, a classification model, a lightweight tracking model, a lightweight verification model, etc. For example, the target inference result may include the task inference result and at least one additional inference information.
[0053] In another implementation, when the path mode indicated by the target control token is a slow path, the cloud can determine the target inference result of the visual data to be processed under the target processing task according to the execution control product; that is, the cloud can determine the target inference result of the visual data to be processed under the target processing task according to the execution control product. Optionally, the cloud can obtain the visual data to be processed (such as visual data to be processed sent from the local device or the receiving end device), and can call the target multimodal large model according to the execution control product to determine the target inference result of the visual data to be processed. For example, the target multimodal large model input data can be determined based on the execution control product and the visual data to be processed, and the target multimodal large model input data can be input into the target multimodal large model to output the target inference result through the target multimodal large model, etc. For example, the image to be processed (also called the payload to be processed) can be constructed according to the keyframe strategy and ROI strategy (such as the region of interest range) indicated by the execution control product, thereby adding the image to be processed to the target multimodal large model input data, and / or the task inference prompt template can also be added to the target multimodal large model input data, etc. Optionally, the target inference result may include the task inference result and / or at least one additional inference information. Optionally, the task inference prompt template may be set according to experience or actual needs, and this embodiment of the invention does not limit this. Optionally, when the target decision control side is an edge device, the edge device may also send the execution control product to the cloud.
[0054] In another implementation, when the path mode indicated by the target control token is a cascaded composite path, the edge device can determine the initial inference result of the visual data to be processed under the target processing task according to the execution control product, and send the data to be uploaded corresponding to the initial inference result to the cloud. The data to be uploaded may include the initial inference result; optionally, the data to be uploaded may also include the verification load determined according to the execution control product. Accordingly, the cloud can then determine whether the initial inference result meets the verification conditions based on preset verification trigger conditions. If the verification conditions are met, the initial inference result is verified through the slow path to obtain the target inference result. Based on this, the initial inference result and the verification inference result obtained through the slow path can be fused, that is, the initial inference result can be adjusted through the verification inference result to obtain the target inference result. Optionally, when the target decision control side is the cloud, the cloud can also send the execution control product or the control channel parameters in the execution control product to the edge device, and so on. Optionally, the initial inference result may include the initial task inference result and / or at least one inference supplementary information corresponding to the initial task inference result; optionally, the verification payload may include, but is not limited to, at least one of the following: ROI cropping image / pixel, ROI mask or coordinates, keyframe, alignment metadata, fast path result summary (i.e., initial inference result summary), control segment, etc., and the embodiments of the present invention do not limit this. For example, the preset review trigger conditions may include, but are not limited to, at least one of the following: confidence level (such as the mean of confidence levels among keyframes, where the confidence level of a keyframe may be the mean of confidence levels of each target in the corresponding keyframe) is lower than the confidence level threshold; uncertainty measure (such as the mean of uncertainty measures among keyframes) is higher than the uncertainty measure threshold; conflict with historical results; the difference between Top-1 (i.e., the category with the highest category prediction probability value) and Top-2 (i.e., the category with the second highest category prediction probability value) is less than the category difference threshold; poor cross-frame consistency (such as the difference between the category or target box of the same target being greater than the consistency difference threshold); conflict with temporal tracking results; detection of out-of-distribution signs; and hitting a specified high-risk strategy knowledge entry (i.e., the target strategy knowledge entry used to determine the control token is a specified high-risk strategy knowledge entry), etc.; the embodiments of the present invention do not limit this. Optionally, the confidence level threshold, uncertainty measure threshold, category difference threshold, consistency difference threshold, and specified high-risk strategy knowledge entry can all be set according to experience or actual needs, and the embodiments of the present invention do not limit this. Optionally, the verification conditions can be determined to be met when one or more of the preset verification trigger conditions are met.
[0055] Optionally, when verifying the initial inference result through the slow path to obtain the target inference result, the target multimodal large model can be called according to the data to be uploaded to determine the verification inference result of the visual data to be processed. Then, the verification inference result can be used to adjust the initial inference result to obtain the target inference result, and so on. For example, the image data to be verified can be determined from the data to be uploaded (such as, but not limited to, at least one of the following: ROI cropped image / pixel, keyframe, fast path result summary, control segment, keyframe with confidence below the confidence threshold, keyframe with uncertainty measure above the uncertainty measure threshold, etc.), and the target multimodal large model can be called to infer the image data to be verified to obtain the verification inference path, and so on. Optionally, the priority of the verification inference result can be higher than the priority of the initial inference result, so that the verification inference result can be used to adjust the initial inference result to obtain the target inference result, thereby realizing the verification of the initial inference result.
[0056] Based on this, embodiments of the present invention can execute the corresponding inference path based on the path pattern indicated by the target control token and in combination with the execution control product. For example, when the path pattern is a fast path, a fast path task can be executed; when the path pattern is a slow path, a slow path task can be executed; when the path pattern is a cascaded composite path, the fast path can be executed first, and then a decision can be made based on preset review trigger conditions to determine whether to call the slow path for review. The slow path review result can correct, supplement, or merge the fast path inference result (i.e., the initial inference result) according to preset review inference correction rules, thereby obtaining the target inference result. Optionally, the preset review inference correction rules can be set according to experience or actual needs, and embodiments of the present invention do not limit this. It can be seen that the fast path, slow path, and cascaded composite path in embodiments of the present invention can be uniformly driven by the control token.
[0057] This invention can acquire routing decision indication data corresponding to a target processing task, as well as task semantic information of the target processing task. Based on the routing decision indication data, runtime context indication data can be determined. An initial control token can be generated based on the runtime context indication data and task semantic information. The initial control token can be generated via an external light router located outside the multimodal large model. Then, based on the initial control token, a target control token can be determined. Each control token includes at least one of at least a preset path pattern, which includes at least two of the following: fast path, slow path, and cascaded composite path. Based on the target control token, the execution control output can be determined. The fast path is executed via the edge device, the slow path is executed via the cloud, and the cascaded composite path is executed jointly by the edge device and the cloud. Correspondingly, based on the path pattern indicated by the target control token, the target execution side of the target processing task can be determined. The target execution side includes the edge device and / or the cloud. Based on this, the target execution side can determine the target inference result of the visual data to be processed under the target processing task based on the execution control output. As can be seen, embodiments of the present invention can determine the target control token through task semantic information, and realize edge-cloud multimodal perception driven by an external light router according to the path pattern indicated by the target control token. That is, a unified application layer control object can be constructed through the control token, and then the specific control semantics of edge-cloud collaborative perception can be realized through the unified application layer control object, so as to more accurately control edge-cloud multimodal perception (i.e., edge-cloud collaborative perception) to meet constraints such as latency. Furthermore, the edge-cloud multimodal perception driven by an external light router can determine the target inference result with higher accuracy. That is, embodiments of the present invention can effectively improve the accuracy of inference results while meeting latency requirements. Moreover, embodiments of the present invention can transform the target control token into an execution control product that can be directly consumed and followed by the execution side, thereby effectively avoiding the deviation between execution behavior and scheduling intention, and improving the consistency, interpretability and controllability of results in the edge-cloud collaborative process.
[0058] Based on the above description, this embodiment of the invention also proposes a more specific edge-cloud multimodal perception method driven by an external light router. Accordingly, this edge-cloud multimodal perception method driven by an external light router can be executed by the target decision control side included in the aforementioned edge-cloud multimodal perception system driven by an external light router. The target decision control side can be an edge device or the cloud within the edge-cloud multimodal perception system driven by an external light router, etc. Please refer to... Figure 2 The edge-cloud multimodal awareness method driven by the external light router may include the following steps S201-S207: S201, obtain the routing decision indication data corresponding to the target processing task, and obtain the task semantic information of the target processing task.
[0059] S202, based on the routing decision indication data, determine the running context indication data; and fuse and encode the running context indication data and task semantic information to obtain fused encoded state features.
[0060] Optionally, the target decision control side can invoke the target encoding model to perform fusion encoding on the runtime context indication data and task semantic information to obtain fusion encoded state features. Optionally, the target encoding model can be any encoding model, and this embodiment of the invention does not limit it, that is, this embodiment of the invention does not limit the model structure of the target encoding model. Optionally, fusion encoding can be performed through an external light router.
[0061] S203, Based on the fused coding state features, determine the target policy knowledge entries from the routing memory; wherein, the routing memory includes a set of policy knowledge entries.
[0062] Based on this, target strategy knowledge items can be determined from the set of strategy knowledge items. Optionally, the number of target strategy knowledge items can be one or more, and this embodiment of the invention does not limit this.
[0063] Optionally, the target decision control side can perform similarity retrieval based on the fused encoded state features and the state feature summary in the routing memory, thereby recalling the top Q policy knowledge entries with the highest similarity as target policy knowledge entries, where Q can be a positive integer.
[0064] Optionally, when the state feature summary is a feature located in the same vector space as the fused encoded state feature, the similarity between the fused encoded state feature and the state feature summaries of each policy knowledge item in the policy knowledge item set can be directly calculated to obtain the similarity between the fused encoded state feature and the state feature summaries of each policy knowledge item (such as cosine similarity or the reciprocal of Euclidean distance, etc.); or, when the state feature summary is not a feature located in the same vector space as the fused encoded state feature (such as the state feature summary being text or having a different dimension from the fused encoded state feature), the target encoding model can be called to encode the state feature summaries of each policy knowledge item to obtain the encoded features of the state feature summaries of each policy knowledge item, thereby calculating the similarity between the fused encoded state feature and the encoded features of the state feature summaries of each policy knowledge item in the policy knowledge item set, and so on; the embodiments of the present invention do not limit this.
[0065] S204, Generate the initial control token according to the target strategy knowledge entries.
[0066] Optionally, the target decision control side can generate pending control tokens for the target processing task according to the target strategy knowledge entries. In one implementation, fields such as path mode, load strategy, maximum number of attention areas, running gear identifier, slow path trigger threshold, and execution resolution (also known as input resolution) in the pending control token can be initialized or adjusted according to the target strategy knowledge entries. In another implementation, a lightweight learnable routing submodule can be invoked to generate pending control tokens based on target strategy knowledge entries, running context indication data, and task semantic information. For example, target strategy knowledge entries, running context indication data, and task semantic information can be input into the lightweight fusion analysis model in the lightweight learnable routing submodule to output pending control tokens through the lightweight fusion analysis model (e.g., pending control tokens can be generated through the scores of each output path mode, budget risk scores, and control field suggestion information). In another implementation, an entry initialization control token can be generated according to the target strategy knowledge entries, and a lightweight learnable routing submodule can be called to generate field adjustment indication information (such as one or more of the following: scores of various path modes, budget risk scores, and control field suggestion information) based on runtime context indication data and task semantic information. Thus, the entry initialization control token can be adjusted using the field adjustment indication information to obtain a pending control token, etc.; the embodiments of the present invention do not limit this.
[0067] Then, the rule barriers under the target processing task can be determined, and the pending control token can be adjusted according to the rule barriers to obtain the initial control token. Optionally, when the pending control token conflicts with the rule barriers, the pending control token can be adjusted according to a preset rule barrier conflict adjustment strategy to obtain the initial control token; optionally, the preset rule barrier conflict adjustment strategy can be set according to experience or actual needs, and this embodiment of the invention does not limit this. For example, when the latency involved in the pending control token does not meet the maximum end-to-end latency rule in the rule barriers, and the path mode is a slow path, the preset rule barrier conflict adjustment strategy can indicate to modify the path mode to a fast path, or it can indicate to adjust the keyframe strategy, etc.; or, when the historical execution effect statistics corresponding to the pending control token do not meet the result stability rules in the rule barriers, such as accuracy, recall, false positive rate, P95 latency, or execution success rate not reaching the corresponding preset stability threshold, and the path mode is a fast path, the preset rule barrier conflict adjustment strategy can indicate to modify the path mode to a cascaded composite path, etc. Optionally, the preset stability threshold corresponding to an execution effect (such as accuracy, recall, etc.) can be set according to experience or actual needs, and the embodiments of the present invention do not limit this.
[0068] S205, based on the initial control token, determine the target control token, a control token including at least one path mode in at least one preset path mode, at least two of the following preset path modes: fast path, slow path and cascaded composite path; and based on the target control token, determine the execution control artifact.
[0069] In one implementation, when determining the target control token based on the initial control token, the initial control token can be used as the target control token.
[0070] In another implementation, when determining the target control token based on the initial control token, budget constraint indication data under the target processing task can be determined. This budget constraint indication data may include the budget values of each budget constraint indicator among at least one budget constraint metric. The at least one budget constraint metric may include, but is not limited to, at least one of the following: edge computing power, edge-to-cloud communication bandwidth, end-to-end latency, cloud call count, operating power consumption, and temperature rise. For example, the budget value for edge computing power may be the edge computing power budget, the budget value for edge-to-cloud communication bandwidth may be the edge-to-cloud communication bandwidth budget, the budget value for end-to-end latency may be the end-to-end latency budget, the budget value for cloud call count may be the cloud call count budget, the budget value for operating power consumption may be the operating power consumption budget, the budget value for temperature rise may be the temperature rise budget, and so on. Then, budget governance can be performed on the initial control token based on the budget constraint indication data to obtain the target control token, so that the target control token meets the budget indicated by the budget constraint indication data. That is, at this time, the target control token can be a valid control token after budget governance, and so on. Optionally, the target decision control side may also include a budget executor (also referred to as a Budget Enforcer or a budget executor module). In this case, budget governance can be performed through the budget executor. The budget executor can monitor budgets such as computing power, bandwidth, latency, and call count on the endpoint. When necessary, it can perform field-level trimming, replacement, or overwriting on the control token to form a derived control token (i.e., a control token that adjusts the initial control token). In other words, it can be used to monitor and / or enforce at least one budget (i.e., adjust corresponding fields, such as forced rollback). It should be understood that the target control token can be the same as the initial control token (if the initial control token does not exceed the budget, the target control token after budget governance is the same as the initial control token, i.e., the fields adjusted by budget governance can be empty), or it can be different (if the initial control token exceeds the budget, the fields adjusted by budget governance can be non-empty, i.e., the initial control token can be adjusted, thus using the derived control token as the target control token), etc. This embodiment of the invention does not limit this.
[0071] Optionally, when performing budget governance on the initial control token based on budget constraint indication data to obtain the target control token, the indicator requirement monitoring information for each budget constraint indicator can be determined (such as the indicator requirement monitoring information for edge computing power, which can be the predicted requirement value of the corresponding budget constraint indicator, such as the required edge computing power, etc.). The indicator requirement monitoring information of a budget constraint indicator can also be called the indicator requirement monitoring information of the initial control token under the corresponding budget constraint indicator. Based on the indicator requirement monitoring information and budget value of each budget constraint indicator, it can be determined whether the initial control token meets any one of at least one budget governance constraint condition. At least one budget governance constraint condition includes budget governance soft constraints and / or budget governance hard constraints. If the initial control token meets the budget governance soft constraint condition, the initial control token is used as the target control token, and the constraint alarm message corresponding to the target control token is output. If the initial control token meets the budget governance hard constraint condition, field-level governance can be performed on the initial control token to obtain the target control token, such as... Figure 3 As shown; that is, field-level governance can be performed on the initial control token to generate a derived control token, which can then be used as the target control token. If the initial control token does not satisfy at least one budget governance constraint (i.e., does not satisfy any budget governance constraint), then the initial control token is used as the target control token. Optionally, when performing field-level governance on the initial control token to obtain the target control token, field-level governance can be performed on the initial control token according to a preset governance strategy to obtain the target control token. Optionally, the constraint alarm prompt information may include, but is not limited to, at least one of the following: risk warnings and recommended degradation strategies, etc.; Optionally, the constraint alarm prompt information may be generated according to a preset alarm generation strategy, which may be set according to experience or actual needs, and this embodiment of the invention does not limit this.
[0072] Optionally, both soft and hard constraints for budget governance can be set based on experience or actual needs, and this embodiment of the invention does not limit this. Optionally, preset governance strategies can be set based on experience or actual needs, and this embodiment of the invention does not limit this; for example, preset governance strategies may instruct adjustments to keyframe strategies (such as changing the extraction of one frame every 2 frames to one frame every 5 frames, or changing the extraction of 5 frames per second to 2 frames per second, etc.) or reduction of input resolution when the edge computing power budget is not met, or to change the payload strategy from uploading pixels of interest region to intermediate features when bandwidth is insufficient, or to change the path mode from cascaded composite path to fast path when latency is tight, etc.
[0073] The aforementioned budget governance can also be referred to as field-level governance. Optionally, field-level governance actions (i.e., preset governance strategies) may include, but are not limited to, at least one of the following: reducing the processing frame rate or increasing the frame extraction interval (i.e., adjusting the keyframe strategy), reducing the number of ROIs, shrinking the ROI area or expanding its range, reducing the input resolution, increasing the compression level, switching the uplink payload from ROI pixels to intermediate features, switching the path mode from cascaded composite paths to fast paths, switching the path mode from slow paths to fast paths or temporarily suppressing slow path calls, switching to a lighter fast path model or execution graph; wherein, the execution graph can be used to represent the inference execution topology or model / operator orchestration relationship composed of preprocessing operators, fast path models, slow path models, postprocessing operators, result fusion rules, and abnormal fallback nodes, etc.; the embodiments of the present invention do not limit this. Optionally, when generating a derived control token, a parent-child association can be established with the initial control token through parent token digest (which can be the token_hash of the initial control token), parent token identifier (which can also be represented as parent_token_id, such as the token identifier of the initial control token), governance action (i.e., budget governance action) digest, generation time, or control token hash chain identifier (such as control token hash chain digest), thereby forming a traceable and auditable token version chain. For example, parent token indication information (also known as parent-child association fields, such as but not limited to parent token digest, parent token identifier, governance action digest, generation time, or control token hash chain identifier) can be added to the derived control token. This allows setting the corresponding fields in the derived control token through the parent token indication information, thus forming a traceable token version chain.
[0074] In one implementation, when determining whether the initial control token satisfies any of the at least one budget governance constraint based on the indicator demand monitoring information and budget values of each budget constraint indicator, it can be determined whether the initial control token satisfies the hard budget governance constraint (also known as the field-level governance trigger condition) based on the indicator demand monitoring information and budget values of each budget constraint indicator. For example, based on the indicator demand monitoring information of each budget constraint indicator (such as the actual demand monitored for the corresponding budget constraint indicator, such as required computing power) and the budget value, it can be determined whether there is a budget constraint indicator to be governed among at least one budget constraint indicator (a budget constraint indicator to be governed can refer to a budget constraint indicator whose indicator demand monitoring information does not meet the corresponding budget value, such as a budget constraint indicator whose budget has been exceeded or a budget constraint indicator predicted to exceed the budget). When there is a budget constraint indicator to be governed among at least one budget constraint indicator, it can be determined that the initial control token satisfies the hard budget governance constraint; when there is no budget constraint indicator to be governed among at least one budget constraint indicator, it can be determined that the initial control token does not satisfy the hard budget governance constraint, and so on. Optionally, the failure of any budget constraint indicator's requirement monitoring information to meet the indicator value of any budget constraint indicator can refer to: the requirement monitoring information of any budget constraint indicator exceeding the budget value of any budget constraint indicator; or the difference between the requirement monitoring information of any budget constraint indicator and the budget value of any budget constraint indicator being less than the budget difference threshold corresponding to any budget constraint indicator; or the ratio between the requirement monitoring information of any budget constraint indicator and the budget value of any budget constraint indicator being greater than the budget ratio threshold corresponding to any budget constraint indicator, etc. This embodiment of the invention does not limit this; for example, if the edge computing power requirement monitoring information is greater than the edge computing power budget, it can be determined that the edge computing power budget is not met, etc. Optionally, the budget difference threshold and budget ratio threshold corresponding to a budget constraint indicator can both be set based on experience or based on actual needs; this embodiment of the invention does not limit this.
[0075] In another implementation, the initial control token can be determined to meet the soft constraint conditions of budget governance based on the indicator requirement monitoring information and budget value of each budget constraint indicator. For example, for any one of the at least one budget constraint indicators, if the difference indicator value (such as difference or ratio) between the indicator requirement monitoring information and the indicator value of any budget constraint indicator falls within the alarm interval range corresponding to any budget constraint indicator, it can be determined that the initial control token meets the soft constraint conditions of budget governance; if the difference indicator value between the indicator requirement monitoring information and the indicator value of any budget constraint indicator does not fall within the alarm interval range corresponding to any budget constraint indicator, it can be determined that the initial control token does not meet the soft constraint conditions of budget governance. Optionally, the alarm interval range corresponding to any budget constraint indicator can be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, taking the difference indicator value as a ratio as an example, the alarm interval range corresponding to any budget constraint indicator can be [70%-90%], etc.
[0076] Based on this, the embodiments of the present invention can divide budget governance into soft constraint governance (i.e., budget governance when the soft constraint conditions of budget governance are met) and hard constraint governance (i.e., budget governance when the hard constraint conditions of budget governance are met). Soft constraint governance can generate constraint alarm prompt information without modifying the initial control token; hard constraint governance can modify the initial control token to obtain a derived control token, and so on.
[0077] Optionally, in other embodiments, the budget constraint indication data may also include rule guardrails, etc.; accordingly, the preset governance process may also include adjusting the control token according to the rule guardrails to finally obtain the target control token, etc.; the present invention does not limit the specific implementation of budget governance.
[0078] As can be seen, this invention, under multiple constraints such as edge computing power, edge-cloud bandwidth, end-to-end latency, and cloud call count, achieves on-demand collaboration, controllable degradation, and stable operation between lightweight task heads and multimodal large models by constructing a unified application-layer control object (i.e., an initial control token) and combining it with a budget executor to implement field-level governance and path adjustment on this control object. This enables the system to perform fine-grained adjustments to path patterns, areas of interest, keyframe strategies, and payload strategies based on specific control semantics in resource-constrained and fluctuating operating conditions, thereby solving the problems of coarse path switching and uncontrollable system behavior under resource-constrained conditions. Based on this, this invention can implement field-level governance around the control token through a budget executor, forming a traceable chain of derived control tokens.
[0079] S206, Based on the path pattern indicated by the target control token, determine the target execution side of the target processing task; wherein, the target execution side includes end-side devices and / or the cloud.
[0080] S207, based on the execution control product, enables the target execution side to determine the target inference result of the visual data to be processed under the target processing task.
[0081] Optionally, the target decision control side can also generate a target stable frame identifier (also represented as FrameKey) corresponding to the visual data to be processed; thereby, a target evidence sample corresponding to the visual data to be processed can be generated based on the target stable frame identifier; and the target evidence sample can be added to the evidence chain (also represented as Evidence Chain), the evidence chain can be used to determine the set of policy knowledge items, and the set of policy knowledge items can be used to determine the initial control token. Optionally, the target evidence sample may include, but is not limited to, at least one of the following: target stable frame identifier, control token digest corresponding to the visual data to be processed, control token hash chain digest (such as the control token hash chain digest corresponding to the target control token, which can be used to record the initial control token and derived control token digest sequence), compilation digest, running gear identifier, budget status snapshot (such as indicator requirement monitoring information and / or budget value of various budget limit indicators, etc.), model version information, load statistics (such as actual load amount, etc.), latency statistics (i.e., actual latency), and inference output digest (i.e., a digest of the target inference result, such as generated by a preset inference output digest generation strategy, or generated by generating model output through the digest, etc.), etc., the embodiments of the present invention do not limit this. Optionally, the target stable frame identifier may include the data frame identifier of the visual data to be processed (which can be used to indicate the visual data to be processed), or it may include the frame identifier of a single image in the visual data to be processed or the frame identifier of each key frame in the key frame set (e.g., the frame identifier of a key frame can be used to indicate the corresponding key frame), etc.; this embodiment of the invention does not specify this. For example, a stable frame identifier may be a stable identifier corresponding to the input image, video frame, or key frame set, which can be used for frame-level alignment, evidence indexing, reproduction location, and responsibility verification, etc. Optionally, the preset inference output summary generation strategy, the preset control token summary generation strategy, and the summary generation model can all be set according to experience or actual needs, and this embodiment of the invention does not limit this. Optionally, the evidence chain may be a structured evidence set indexed by the stable frame identifier, which can be used to record key information such as control, governance, compilation, injection, budget, version, payload, and output, and provide a factual basis for reproduction and closed-loop optimization. It should be noted that the data format of the control token and the evidence chain is not limited in this embodiment of the invention. For example, it can be in JSON format (JavaScript Object Notation), or it can be implemented using binary protocol objects, message queue objects, structured database records, etc. The control token digest can be a digest calculated after normalizing and serializing a single control token, and the control token hash chain digest can be a digest calculated from the original control token and the sequence of derived control token digests.
[0082] Optionally, a stable frame identifier can be generated using an Access Unit hash (also referred to as an Access Unit), a normalized pixel hash, a perceptual hash combination, or other digest methods capable of stably identifying the input object. Optionally, a synchronization method (also referred to as `sync_method`) can also be recorded, which can be used to describe the frame-level alignment between the target stable frame identifier and the original video stream, decoded frames, keyframe set, end-side processing results, or cloud-based verification results. The synchronization method may include, but is not limited to, at least one of the following: Access Unit alignment, timestamp alignment, frame sequence number alignment, pixel hash alignment, perceptual hash alignment, etc. Optionally, the target evidence sample may also include, but is not limited to, at least one of the following: an injection method digest (such as whether a control segment is injected, the injection location, the injection field, and the injection template version), an execution graph digest (such as an execution graph identifier, model / operator node version, node connection relationships, post-processing rules, and fallback node configuration), anomaly codes, and failure reasons, etc. This embodiment of the invention does not limit these aspects.
[0083] Optionally, a reproducibility signature (also represented as `repro_signature`, or a reproducible signature) of the target evidence sample can be further generated. The reproducibility signature can be generated from the target stable frame identifier, control token hash chain digest, etc. Optionally, the control token hash chain digest can be generated from at least one combination of the digest value calculated after normalizing and serializing the control token hash chain and its optional associated information, and the model / operator version identifier. Optionally, the reproducibility signature can be further enhanced by combining a compiled digest or an injection-based digest, etc. Optionally, the reproducibility signature can be generated when the target evidence sample is encapsulated to support accurate reproduction. It should be noted that the specific algorithm for generating the reproducibility signature is not limited in the embodiments of the present invention, such as not requiring a specific hash algorithm as a necessary limitation. It should be noted that the hash algorithm, signature algorithm, and serialization method involved in the embodiments of the present invention are not limited, that is, they can be set according to experience or actual needs, etc. Based on this, the target evidence sample may include, but is not limited to, at least one of the following: target stable frame identifier, control token digest corresponding to the visual data to be processed, control token hash chain digest, injection method digest, execution graph digest, exception code, failure reason, reproducible signature, compilation digest, runtime status identifier, budget status snapshot, model version information, load statistics, latency statistics, and inference output digest, etc.
[0084] Optionally, the target decision control side may also include an audit traceability module to generate a target stability frame identifier and record target evidence samples, such as adding target evidence samples to the evidence chain.
[0085] Optionally, the target decision control side may also include a caching and quick detection module and / or a policy update module, etc., which are not limited in this embodiment of the invention. The caching and quick detection module can provide candidate regions of interest (ROI), historical results, cache hit rate, and reusable intermediate feature information, which can assist in routing decisions and cost control. The policy update module can update routing policy parameters, runtime parameters, and memory increments based on the evidence chain, and supports canary deployment and rollback. That is, the routing memory can be updated by updating the set of policy knowledge items based on the evidence chain. Therefore, the routing memory is not the original log repository, but includes a set of policy knowledge items determined by the evidence chain.
[0086] Correspondingly, the target decision control side can determine the set of policy knowledge items based on the evidence chain to update the policy knowledge item set, thereby updating the routing memory. Optionally, the policy knowledge item set can be determined based on the evidence chain (i.e., the current evidence chain) every preset update interval, and the determined policy knowledge item set (i.e., the current policy knowledge item set) can be updated to the policy knowledge item set in the routing memory, thus updating the routing memory every preset update interval; or, the policy knowledge item set can be determined based on the evidence chain to update the routing memory each time the increase in the evidence sample in the evidence chain reaches a preset increase threshold, etc.; this embodiment of the invention does not limit this. Optionally, both the preset update interval and the preset increase threshold can be set according to experience or actual needs, and this embodiment of the invention does not limit this.
[0087] Optionally, when determining the set of policy knowledge items based on the chain of evidence, the chain of evidence can be aggregated or distilled to obtain the set of policy knowledge items. That is, the set of policy knowledge items can be obtained by aggregating or distilling the chain of evidence.
[0088] In one implementation, evidence samples in the evidence chain can be grouped and aggregated according to preset grouping rules to obtain multiple evidence sample groups. Then, the strategy knowledge item template can be filled according to each evidence sample group in the multiple evidence sample groups to obtain the strategy knowledge item corresponding to each evidence sample group. The strategy knowledge item corresponding to each evidence sample group is then added to the strategy knowledge item set to determine the strategy knowledge item set. Optionally, both the preset grouping rules and the strategy knowledge item template can be set according to experience or actual needs, and this embodiment of the invention does not limit this. For example, the preset grouping rules may include, but are not limited to, at least one of the following: task semantic summary (e.g., determined according to task identifier, business template, Prompt, target category set, output requirements, or task priority), bandwidth bucketing, resource status bucketing, and uncertainty metric bucketing, etc.
[0089] It should be noted that the specific implementation method of filling in the data is not limited in the embodiments of the present invention. For example, when filling the strategy knowledge item template according to each of the multiple evidence sample groups to obtain the strategy knowledge items corresponding to each evidence sample group, for any evidence sample group in the multiple evidence sample groups and any field in the strategy knowledge item template, all field contents of any field can be determined from any evidence sample group, and the field contents with the largest number of identical field contents can be filled as the field contents of any field in the strategy knowledge item template, thereby obtaining the strategy knowledge item corresponding to any evidence sample group; or, for any evidence sample group in the multiple evidence sample groups, under the condition of satisfying the preset item generation constraints, the strategy knowledge item with the best target effect index under any evidence sample group can be extracted according to the strategy knowledge item template, and the extracted strategy knowledge item can be used as the strategy knowledge item corresponding to any evidence sample group, etc.; the embodiments of the present invention do not limit this. Optionally, the preset item generation constraints and target performance indicators can be set according to experience or actual needs, and the embodiments of the present invention do not limit this; for example, the preset item generation constraints may include, but are not limited to, at least one of the following: accuracy, latency and cost constraints, etc., and the target performance indicators may include, but are not limited to, at least one of the following: accuracy and latency, etc.
[0090] In another implementation, a set of policy knowledge items can be determined based on a chain of evidence using a policy knowledge item learning model (such as any reinforcement learning model). For example, the chain of evidence and / or policy knowledge item templates can be input into the policy knowledge item learning model to output a set of policy knowledge items, etc.; this embodiment of the invention does not limit this. Based on this, this embodiment of the invention can also aggregate or generalize the chain of evidence using the policy knowledge item learning model. Optionally, when reinforcement learning is used, one or more combinations of accuracy, latency, cost, and energy consumption can be used as reward signals, etc.; this embodiment of the invention does not limit this.
[0091] As can be seen, the updates to the routing strategy and routing memory in this embodiment of the invention can employ one or more of the following methods: supervised learning, knowledge distillation, reinforcement learning, offline rule mining, or online incremental updates, to achieve closed-loop optimization. For example, the initial control token can be determined through the updated routing memory. Optionally, the routing strategy may include, but is not limited to, at least one of the following: rules for determining control token fields such as path mode selection, execution level, maximum number of ROIs, slow path trigger threshold, input resolution, keyframe strategy, load strategy, model version, execution graph selection, and post-processing threshold.
[0092] Optionally, the updated routing memory can be updated through canary releases and rollbacks, meaning the updated routing memory supports both canary updates and rollbacks (e.g., Figure 4 (as shown), etc. It should be noted that the implementation method of the routing memory in this embodiment of the invention is not limited. For example, it can be implemented using a key-value library, a vector retrieval library, a rule library, or a combination thereof, etc.; and the update method of the routing memory in this embodiment of the invention is not limited. For example, the update method can be offline batch update, online incremental update, or gray-scale update, etc.
[0093] Based on this, embodiments of the present invention can, through stable frame identifiers, evidence chains, and routing memory (also referred to as Routing Memory), precipitate the control, governance, compilation, execution, and output elements in the edge-cloud collaboration process into traceable, reproducible, and closed-loop optimizable engineering assets. This enables the system to not only form a structured evidence set around a specific input frame, a specific control decision, a specific budget governance action, a specific execution injection method, and a specific actual model version, but also to extract, aggregate, and reuse strategy knowledge based on historical operational evidence, thereby solving the problems of difficulty in frame-level review, difficulty in responsibility verification, and difficulty in reusing historical experience. Accordingly, embodiments of the present invention enable the system to quickly restore the input, control, budget, and version status at the time of false alarms, missed alarms, link disputes, or operational anomalies, improving reproduction and reconciliation efficiency. Furthermore, embodiments of the present invention can continuously perform knowledge distillation and / or retrieval enhancement on the evidence chain, thereby continuously accumulating strategy knowledge and improving the stability and adaptability of subsequent routing decisions.
[0094] To further illustrate the embodiments of the present invention, we will take the example of an unmanned aerial vehicle (UAV) performing urban road inspections and identifying "fire lane obstruction". The UAV is equipped with a visible light camera to collect inspection video streams (i.e. visual data to be processed). The inspection video streams can be 1920×1080 resolution and 25 frames / second.
[0095] Based on this, such as Figure 5As shown, the visual data to be processed can be input into the data access and preprocessing module to realize the data access and preprocessing of the visual data to be processed. The data access and preprocessing module can perform decoding, frame extraction, scale normalization and quality screening on the inspection video stream, such as extracting video frames (i.e. key frame set) at 2 frames / second; at the same time, it receives task semantic information, which can be a business template such as "identify fire lane obstruction" or "identify illegally parked vehicles". Then, the caching and fast detection module can perform fast perception on the keyframe set, outputting information such as candidate areas of interest, preliminary category results, confidence levels, and uncertainty metrics (edge-side fast perception results, a type of runtime context information). For example, in a certain video frame (i.e., a certain keyframe), two suspected occupants are detected, one with a confidence level of 0.82 and the other with a confidence level of 0.41. Simultaneously, the edge-cloud multimodal perception system driven by the external lightweight router can collect current runtime environment indication data (another type of runtime context information), such as edge GPU (graphics processing unit) utilization of 56%, edge-cloud link bandwidth of 8 Mbps (Megabits per second), cache hit rate of 73%, and edge-side temperature of 62 degrees Celsius. Correspondingly, the external lightweight router module can jointly input the task semantic information and the aforementioned runtime context information, and combine it with rule guardrails, lightweight learnable routing results (i.e., the output of a lightweight learnable model), and / or routing memory retrieval results to generate an initial control token.
[0096] For example, suppose the initial control token may include the following: the run mode is SAFE, the path mode is a cascaded composite path, the region of interest is the two suspected target boxes mentioned above, the keyframe strategy is to select two keyframes from consecutive segments, the payload strategy is to upload pixels of the region of interest, and the control vector includes parameters such as an upper limit of 2 regions of interest, an input resolution of 960×540, and a slow path trigger threshold (such as a confidence threshold) of 0.45. Subsequently, the budget executor module can perform budget review (i.e., budget governance) on the initial control token. If the bandwidth is detected to drop below 2 Mbps, or the end-to-end latency budget becomes tight, the budget executor module can perform field-level governance on the initial control token, such as adjusting the upper limit of the region of interest from 2 to 1, switching the payload strategy from pixels of interest to intermediate features, or adjusting the path mode from a cascaded composite path to a fast path, thereby forming a derived control token. The derived control token can establish a correspondence with the initial control token through the parent token association field to form a derived control token chain; at this time, the derived control token can be used as the target control token. Correspondingly, if the initial control token does not meet the field-level governance triggering conditions, the initial control token can be used as the target control token.
[0097] Then, the control token adapter module can semantically compile the target control token to generate control channel parameters and / or control fragments, and generate a compilation summary. For example, control channel parameters can be used to constrain execution resolution, the scope of the region of interest, and the payload form; control fragments can be injected into the slow path input link to explicitly instruct the target multimodal large model to review the "fire lane occupancy" task and the specified region of interest. During the path execution phase, if the target control token indicates a fast path, the fast path execution module calls the edge-side inference model (such as a lightweight detection model) to perform local inference on the region of interest, such as outputting the target inference result within 120ms. The target inference result may include the "suspected fire lane occupancy" result and the corresponding confidence level, etc.; if the target control token indicates a slow path or a cascaded composite path, the slow path execution module calls the cloud-based multimodal large model (i.e., the target multimodal large model) to perform complex semantic understanding on the keyframe set and the region of interest. When the confidence level of the fast path output is lower than 0.45, or the uncertainty metric is higher than the corresponding uncertainty metric threshold, the slow path review is triggered. The slow path can further determine whether the area of interest is indeed a fire lane, whether there are vehicles occupying it, and whether an alarm description needs to be output. The verification results will cover or merge the fast path results to form the final inspection conclusion (i.e., the target reasoning result).
[0098] After execution, the audit traceability module generates a target stable frame identifier for the currently processed object (i.e., the visual data to be processed). Using the target stable frame identifier as an index, it writes information such as the control token hash chain (e.g., a chain composed of the initial control token and derived control token digests in sequence) or the control token hash chain digest (e.g., a digest of the entire chain again), compilation digest, budget snapshot, path pattern, model version, load statistics, latency statistics, and output digest into the target evidence sample, and then into the evidence chain. For example, it can record the target stable frame identifier, bandwidth of 1.8 Mbps (belonging to runtime context information), edge GPU utilization of 56%, fast path latency of 120ms, slow path review latency of 780ms, and the final judgment result of "fire lane obstruction". If subsequent misjudgments, disputes, or responsibility verification needs arise, the control decisions, budget governance actions, and execution paths at that time can be restored based on the evidence chain.
[0099] Furthermore, if the system repeatedly encounters similar scenarios of "medium bandwidth + high load at the edge + low-confidence target on the fast path" in multiple inspection tasks, the routing memory can be updated based on evidence samples. For example, "task semantic summary, bandwidth status, resource status, and fast path uncertainty" can be used as state features to aggregate historical evidence samples and extract the best-performing combination of control parameters. For instance, when the bandwidth is between 1 and 3 Mbps and the fast path confidence is below 0.5, cascaded composite paths are preferred, the maximum number of attention areas is set to 1, the input resolution is set to 640×360, and the load strategy prioritizes the use of intermediate features. Thus, this embodiment of the invention can realize a complete closed loop from input acquisition, control token generation, budget governance, semantic compilation, path execution, evidence consolidation to policy optimization in UAV inspection scenarios.
[0100] In summary, the embodiments of the present invention can achieve a complete closed loop of "control object - governance action - execution behavior - evidence solidification - strategy optimization". Furthermore, the embodiments of the present invention can call the target multimodal large model in the cloud on demand under budget constraints, reducing unnecessary cloud calls and communication burdens, and alleviating overall system latency pressure.
[0101] This invention embodiment can, after obtaining the routing decision indication data corresponding to the target processing task and the task semantic information of the target processing task, determine the running context indication data based on the routing decision indication data; and fuse and encode the running context indication data and the task semantic information to obtain fused encoded state features. Then, based on the fused encoded state features, the target policy knowledge entries can be determined from the routing memory; and an initial control token can be generated according to the target policy knowledge entries. Further, based on the initial control token, a target control token can be determined; and based on the target control token, the execution control product can be determined. Based on this, based on the path pattern indicated by the target control token, the target execution side of the target processing task can be determined; and based on the execution control product, the target execution side can determine the target inference result of the visual data to be processed under the target processing task. It can be seen that this invention embodiment can realize a governable, traceable, reproducible, and closed-loop optimized edge-cloud multimodal perception technology solution (i.e., an external lightweight route-driven edge-cloud multimodal perception method) through evidence chains, etc., which can realize unified control objects, execution-side deterministic constraints, frame-level evidence-based tracing, and evidence-based continuous optimization.
[0102] Based on the description of the relevant embodiments of the edge-cloud multimodal perception method driven by the external light router described above, this invention also proposes an edge-cloud multimodal perception system driven by an external light router; as follows... Figure 6As shown, the edge-cloud multimodal perception system driven by an external light router includes a target decision control side 601 and a target execution side 602. The target decision control side 601 is either an edge device or a cloud device within the edge-cloud multimodal perception system driven by an external light router. This edge-cloud multimodal perception system driven by an external light router can execute... Figure 1 or Figure 2 The external light router-driven edge-cloud multimodal perception method shown above, i.e., the edge-cloud multimodal perception system driven by the external light router, can run the above-mentioned units: The target decision control side 601 is used to obtain routing decision indication data corresponding to the target processing task, and to obtain the task semantic information of the target processing task; The target decision control side 601 is also used to determine the running context indication data based on the routing decision indication data; and to generate an initial control token based on the running context indication data and the task semantic information; wherein the initial control token is generated by an external light router, which is located outside the multimodal large model; The target decision control side 601 is further configured to determine a target control token based on the initial control token. A control token includes at least one path mode from at least one preset path mode. The at least one preset path mode includes at least two of the following: fast path, slow path, and cascaded composite path. Based on the target control token, the execution control product is determined. The fast path is executed through the end-side device, the slow path is executed through the cloud, and the cascaded composite path is executed jointly by the end-side device and the cloud. The target decision control side 601 is further configured to determine the target execution side 602 of the target processing task based on the path pattern indicated by the target control token; wherein, the target execution side 602 includes the end-side device and / or the cloud. The target decision control side 601 is also used to enable the target execution side 602 to determine the target reasoning result of the visual data to be processed under the target processing task based on the execution control product.
[0103] In one implementation, when the target decision control side 601 generates an initial control token based on the runtime context indication data and the task semantic information, it can specifically be used for: The runtime context indication data and the task semantic information are fused and encoded to obtain fused encoded state features; Based on the fused coding state features, target policy knowledge entries are determined from the routing memory; wherein, the routing memory includes a set of policy knowledge entries; Generate an initial control token according to the target strategy knowledge entries.
[0104] In another implementation, when the target decision control side 601 generates the initial control token according to the target strategy knowledge entries, it can specifically be used for: Generate a pending control token for the target processing task according to the target strategy knowledge entries; Determine the rule guardrails under the target processing task, and adjust the pending control tokens according to the rule guardrails to obtain the initial control token.
[0105] In another implementation, when determining the target control token based on the initial control token, the target decision control side 601 may specifically be used for: Determine the budget constraint indication data under the target processing task; wherein, the budget constraint indication data includes the budget value of each budget constraint indicator in at least one budget constraint indicator, and the at least one budget constraint indicator includes at least one of the following: edge computing power, edge-cloud communication bandwidth, end-to-end latency, cloud call count, operating power consumption, and temperature rise; Budget governance is performed on the initial control token based on the budget constraint indication data to obtain a target control token, so that the target control token satisfies the budget indicated by the budget constraint indication data.
[0106] In another implementation, when the target decision control side 601 performs budget governance on the initial control token based on the budget constraint indication data to obtain the target control token, it can specifically be used for: Determine the indicator requirement monitoring information for each of the aforementioned budget constraint indicators; Based on the indicator requirement monitoring information and budget value of each budget constraint indicator, it is determined whether the initial control token satisfies any one of the budget governance constraints, wherein the at least one budget governance constraint includes budget governance soft constraints and / or budget governance hard constraints. If the initial control token satisfies the budget governance soft constraint condition, then the initial control token is used as the target control token, and the constraint alarm message corresponding to the target control token is output. If the initial control token satisfies the budget governance hard constraint, then field-level governance is performed on the initial control token to obtain the target control token; If the initial control token does not satisfy the at least one budget governance constraint, then the initial control token will be used as the target control token.
[0107] In another implementation, when determining the execution control product based on the target control token, the target decision control side 601 can specifically be used for: The target control token is semantically compiled by the control token adapter to obtain the execution control product, which includes control channel parameters and / or control fragments, so that the execution control product can be consumed by the target execution side. The control segment includes at least one of the following: a region of interest, a keyframe strategy, a payload strategy, and a verification intent. The control segment supports injection into the slow path inference link to constrain the inference of the target inference multimodal large model.
[0108] In another implementation, when the target decision control side 601 determines the target inference result of the visual data to be processed under the target processing task based on the execution control output, it can be specifically used for: When the path pattern indicated by the target control token is the fast path, the end device determines the target inference result of the visual data to be processed under the target processing task according to the execution control output. When the path pattern indicated by the target control token is the slow path, the cloud determines the target inference result of the visual data to be processed under the target processing task according to the execution control output. When the path pattern indicated by the target control token is the cascaded composite path, the end device determines the initial inference result of the visual data to be processed under the target processing task according to the execution control product, and sends the data to be uploaded corresponding to the initial inference result to the cloud, the data to be uploaded including the initial inference result; then the cloud determines whether the initial inference result meets the review conditions based on the preset review trigger conditions. If the review conditions are met, the initial inference result is reviewed through the slow path to obtain the target inference result.
[0109] In another implementation, the target decision control side 601 can also be used for: Generate the target stable frame identifier corresponding to the visual data to be processed; Based on the target stable frame identifier, a target evidence sample corresponding to the visual data to be processed is generated; and the target evidence sample is added to the evidence chain, the evidence chain supporting the determination of a set of policy knowledge entries, and the set of policy knowledge entries supporting the determination of the initial control token. The target evidence sample includes at least one of the following: the target stable frame identifier, the control token digest corresponding to the visual data to be processed, the control token hash chain digest, the injection method digest, the execution graph digest, the exception code, the failure reason, the reproducible signature, the compilation digest, the running gear identifier, the budget status snapshot, the model version information, the load statistics, the latency statistics, and the inference output digest.
[0110] According to one embodiment of the present invention, Figure 6 In the externally powered lightweight router-driven edge-cloud multimodal sensing system shown, each device can be individually or entirely merged into one or more other units, or one or more of these devices can be further divided into multiple functionally smaller devices. This achieves the same operation without affecting the technical effects of the embodiments of the present invention. In practical applications, the function of one device can also be implemented by multiple devices, or the function of multiple devices can be implemented by one device. In other embodiments of the present invention, any externally powered lightweight router-driven edge-cloud multimodal sensing system may also include other devices. In practical applications, these functions can also be implemented with the assistance of other devices, and can be implemented collaboratively by multiple devices.
[0111] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0112] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0113] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A multimodal sensing method for edge-cloud driven by an external lightweight router, characterized in that, The externally powered lightweight router-driven edge-cloud multimodal perception method is applied to the target decision control side of the externally powered lightweight router-driven edge-cloud multimodal perception system. The target decision control side is an edge device or cloud in the externally powered lightweight router-driven edge-cloud multimodal perception system. The method includes: Obtain routing decision indication data corresponding to the target processing task, and obtain the task semantic information of the target processing task; Based on the routing decision indication data, runtime context indication data is determined; and based on the runtime context indication data and the task semantic information, an initial control token is generated; wherein, the initial control token is generated through an external light router, which is located outside the multimodal large model; Based on the initial control token, a target control token is determined. A control token includes at least one of at least a preset path pattern, which includes at least two of the following: a fast path, a slow path, and a cascaded composite path. Based on the target control token, an execution control product is determined. The fast path is executed through the end-side device, the slow path is executed through the cloud, and the cascaded composite path is executed jointly by the end-side device and the cloud. Based on the path pattern indicated by the target control token, the target execution side of the target processing task is determined; wherein, the target execution side includes the end-side device and / or the cloud; Based on the execution control output, the target execution side determines the target inference result of the visual data to be processed under the target processing task.
2. The method according to claim 1, characterized in that, The process of generating an initial control token based on the runtime context indication data and the task semantic information includes: The runtime context indication data and the task semantic information are fused and encoded to obtain fused encoded state features; Based on the fused coding state features, target policy knowledge entries are determined from the routing memory; wherein, the routing memory includes a set of policy knowledge entries; Generate an initial control token according to the target strategy knowledge entries.
3. The method according to claim 2, characterized in that, The step of generating an initial control token according to the target strategy knowledge entries includes: Generate a pending control token for the target processing task according to the target strategy knowledge entries; Determine the rule guardrails under the target processing task, and adjust the pending control tokens according to the rule guardrails to obtain the initial control token.
4. The method according to any one of claims 1-3, characterized in that, The process of determining the target control token based on the initial control token includes: Determine the budget constraint indication data under the target processing task; wherein, the budget constraint indication data includes the budget value of each budget constraint indicator in at least one budget constraint indicator, and the at least one budget constraint indicator includes at least one of the following: edge computing power, edge-cloud communication bandwidth, end-to-end latency, cloud call count, operating power consumption, and temperature rise; Budget governance is performed on the initial control token based on the budget constraint indication data to obtain a target control token, so that the target control token satisfies the budget indicated by the budget constraint indication data.
5. The method according to claim 4, characterized in that, The process of performing budget governance on the initial control token based on the budget constraint indication data to obtain the target control token includes: Determine the indicator requirement monitoring information for each of the aforementioned budget constraint indicators; Based on the indicator requirement monitoring information and budget value of each budget constraint indicator, it is determined whether the initial control token satisfies any one of the budget governance constraints, wherein the at least one budget governance constraint includes budget governance soft constraints and / or budget governance hard constraints. If the initial control token satisfies the budget governance soft constraint condition, then the initial control token is used as the target control token, and the constraint alarm message corresponding to the target control token is output. If the initial control token satisfies the budget governance hard constraint, then field-level governance is performed on the initial control token to obtain the target control token; If the initial control token does not satisfy the at least one budget governance constraint, then the initial control token will be used as the target control token.
6. The method according to any one of claims 1-3, characterized in that, The determination of execution control artifacts based on the target control token includes: The target control token is semantically compiled by the control token adapter to obtain the execution control product, which includes control channel parameters and / or control fragments, so that the execution control product can be consumed by the target execution side. The control segment includes at least one of the following: a region of interest, a keyframe strategy, a payload strategy, and a verification intent. The control segment supports injection into the slow path inference link to constrain the inference of the target inference multimodal large model.
7. The method according to any one of claims 1-3, characterized in that, The step of determining the target inference result of the visual data to be processed under the target processing task based on the execution control output includes: When the path pattern indicated by the target control token is the fast path, the end device determines the target inference result of the visual data to be processed under the target processing task according to the execution control output. When the path pattern indicated by the target control token is the slow path, the cloud determines the target inference result of the visual data to be processed under the target processing task according to the execution control output. When the path pattern indicated by the target control token is the cascaded composite path, the end device determines the initial inference result of the visual data to be processed under the target processing task according to the execution control product, and sends the data to be uploaded corresponding to the initial inference result to the cloud, the data to be uploaded including the initial inference result; then the cloud determines whether the initial inference result meets the review conditions based on the preset review trigger conditions. If the review conditions are met, the initial inference result is reviewed through the slow path to obtain the target inference result.
8. The method according to any one of claims 1-3, characterized in that, The method further includes: Generate the target stable frame identifier corresponding to the visual data to be processed; Based on the target stable frame identifier, a target evidence sample corresponding to the visual data to be processed is generated; and the target evidence sample is added to the evidence chain, the evidence chain supporting the determination of a set of policy knowledge entries, and the set of policy knowledge entries supporting the determination of the initial control token. The target evidence sample includes at least one of the following: the target stable frame identifier, the control token digest corresponding to the visual data to be processed, the control token hash chain digest, the injection method digest, the execution graph digest, the exception code, the failure reason, the reproducible signature, the compilation digest, the running gear identifier, the budget status snapshot, the model version information, the load statistics, the latency statistics, and the inference output digest.
9. An edge-cloud multimodal sensing system driven by an external lightweight router, characterized in that, The externally powered lightweight router-driven edge-cloud multimodal perception system includes an edge device and a cloud platform, wherein the edge device or the cloud platform serves as the target decision-making and control side of the externally powered lightweight router-driven edge-cloud multimodal perception system; wherein... The target decision control side is used to obtain routing decision indication data corresponding to the target processing task, and to obtain the task semantic information of the target processing task; The target decision control side is also used to determine the running context indication data based on the routing decision indication data; and to generate an initial control token based on the running context indication data and the task semantic information; wherein the initial control token is generated through an external light router, which is located outside the multimodal large model; The target decision control side is further configured to determine a target control token based on the initial control token. A control token includes at least one path mode from at least one preset path mode, and the at least one preset path mode includes at least two of the following: fast path, slow path, and cascaded composite path. Based on the target control token, the execution control product is determined. The fast path is executed through the end-side device, the slow path is executed through the cloud, and the cascaded composite path is executed jointly by the end-side device and the cloud. The target decision control side is further configured to determine the target execution side of the target processing task based on the path pattern indicated by the target control token; wherein the target execution side includes the end-side device and / or the cloud. The target decision control side is also used to enable the target execution side to determine the target reasoning result of the visual data to be processed under the target processing task based on the execution control output.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.