End-cloud cooperative target perception method and system, and storage medium

By using an edge-cloud collaborative target perception method, target attention heatmaps and sets of attention regions are generated using edge devices, and key frames are extracted. This solves the problem of low target perception accuracy in bandwidth-constrained and latency-sensitive scenarios, and achieves the effect of improving accuracy and reducing redundant data transmission under low latency.

CN122265928BActive Publication Date: 2026-08-04AEROSPACE AGE LOW AERIAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AEROSPACE AGE LOW AERIAL TECHNOLOGY CO LTD
Filing Date
2026-05-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In bandwidth-constrained and latency-sensitive scenarios, existing technologies cannot effectively improve the accuracy of target perception.

Method used

By using an edge-cloud collaborative target perception method, a target attention heatmap is generated using edge devices to determine the set of attention regions and extract key frames, thereby reducing the amount of data uploaded and achieving efficient cloud computing.

Benefits of technology

While ensuring low latency, improve the accuracy of target perception, reduce redundant data transmission and invalid cloud computing, and reduce overall system latency and bandwidth usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265928B_ABST
    Figure CN122265928B_ABST
Patent Text Reader

Abstract

The application provides an end-cloud cooperative target perception method and system and a storage medium. The method comprises: obtaining video data to be processed and task semantic information under a target monitoring task on an end-side device; generating a target attention heat map of each video frame to be processed based on the task semantic information; determining a set of attention regions of each video frame to be processed based on the target attention heat map of each video frame to be processed; determining a set of key frames from the video data to be processed based on the set of attention regions of each video frame to be processed; determining a cloud-side data packet to be processed based on the attention region indication data of each key frame in the set of key frames; uploading the cloud-side data packet to be processed to the cloud side to perform target perception based on the cloud-side data packet to be processed on the cloud side, and obtaining a target attention region perception result. The embodiment of the application can improve the accuracy of target perception while ensuring low latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an edge-cloud collaborative target perception method, system, and storage medium. Background Technology

[0002] Currently, with the widespread adoption of sensing terminals such as cameras and unmanned aerial vehicles (UAVs), video analytics tasks are rapidly increasing in bandwidth-constrained and latency-sensitive scenarios. Related technologies typically employ lightweight models for inference to reduce the amount of data transmitted and latency, but this results in lower accuracy in target perception. Therefore, there is currently no satisfactory solution for improving the accuracy of target perception while maintaining low latency. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide an edge-cloud collaborative target perception method, system, and storage medium to solve the problems of low accuracy in target perception caused by related technologies. In other words, embodiments of the present invention can accurately determine the set of attention regions of each video frame to be processed through the target attention heatmap of each video frame to be processed, and then accurately determine the set of key frames. This allows the cloud-based data packets with small data volumes to be uploaded to the cloud, effectively reducing redundant data transmission and invalid cloud computation. It can reduce the overall system latency and bandwidth usage while ensuring the accuracy of target perception. In other words, embodiments of the present invention can effectively improve the accuracy of target perception while ensuring low latency.

[0004] According to one aspect of the present invention, an edge-cloud collaborative target perception method is provided, the edge-cloud collaborative target perception method being applied to an edge-cloud collaborative target perception system, the edge-cloud collaborative target perception system including an edge device and a cloud, the method comprising: The device acquires the video data to be processed under the target monitoring task, and also acquires the task semantic information of the target monitoring task. Based on the task semantic information, target attention heatmaps are generated for each video frame to be processed in the video data to be processed; and based on the target attention heatmaps of each video frame to be processed, the set of attention regions for each video frame to be processed is determined. Based on the set of regions of interest for each video frame to be processed, a set of keyframes is determined from the video data to be processed; and based on the region of interest indication data of each keyframe in the set of keyframes, a data packet to be processed in the cloud is determined; wherein, the region of interest indication data of a keyframe includes the region of interest indication data of each region of interest in the set of regions of interest of the corresponding keyframe, the region of interest indication data of a region of interest includes the region of interest image pixels and region metadata of the corresponding region of interest, and the region metadata of a region of interest includes the region location information of the corresponding region of interest; The data packet to be processed in the cloud is uploaded to the cloud so that target perception is performed on the cloud side based on the data packet to be processed in the cloud, and the target interest area perception result is obtained.

[0005] According to another aspect of the present invention, an edge-cloud collaborative target perception system is provided, the edge-cloud collaborative target perception system comprising an edge device and a cloud; wherein, The end-side device is used to acquire video data to be processed under the target monitoring task, and to acquire the task semantic information of the target monitoring task; The edge device is further configured to generate target attention heatmaps for each video frame to be processed in the video data to be processed based on the task semantic information; and to determine the set of attention regions for each video frame to be processed based on the target attention heatmaps for each video frame to be processed. The edge device is further configured to determine a set of keyframes from the video data to be processed based on the set of regions of interest of each video frame to be processed; and to determine a data packet to be processed in the cloud based on the region of interest indication data of each keyframe in the set of keyframes; wherein, the region of interest indication data of a keyframe includes the region of interest indication data of each region of interest in the set of regions of interest of the corresponding keyframe, the region of interest indication data of a region of interest includes the region of interest image pixels and region metadata of the corresponding region of interest, and the region metadata of a region of interest includes the region location information of the corresponding region of interest; The end-side device is also used to upload the data packets to be processed in the cloud to the cloud; The cloud is used to perform target perception based on the data packets to be processed in the cloud, and obtain the target area of ​​interest perception result.

[0006] According to another aspect of the present invention, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods mentioned above.

[0007] This invention allows for the acquisition of video data to be processed under a target monitoring task, as well as the task semantic information of the target monitoring task, on the edge device side. Then, based on the task semantic information, a target attention heatmap for each video frame to be processed can be generated. Based on the target attention heatmaps of each video frame, a set of regions of interest for each video frame can be determined. Further, based on the set of regions of interest for each video frame, a set of keyframes can be determined from the video data to be processed. Based on the region of interest indication data of each keyframe in the keyframe set, a cloud-based data packet to be processed can be determined. The region of interest indication data for a keyframe includes the region of interest indication data for each region in the corresponding keyframe's region of interest set; the region of interest indication data for a region includes the region of interest image pixels and region metadata; and the region metadata for a region of interest includes the region location information. Based on this, the cloud-based data packet to be processed can be uploaded to the cloud for target perception on the cloud side, obtaining the target region of interest perception result. As can be seen, the embodiments of the present invention can accurately determine the set of attention regions of each video frame to be processed by the target attention heatmap of each video frame to be processed, and then accurately determine the set of key frames. This allows the cloud-based data packets with small data volume to be processed to be uploaded to the cloud, thereby effectively reducing redundant data transmission and invalid cloud computing. It can reduce the overall system latency and bandwidth usage while ensuring the accuracy of target perception. In other words, the embodiments of the present invention can effectively improve the accuracy of target perception while ensuring low latency. Attached Figure Description

[0008] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart illustrating an edge-cloud collaborative target perception method according to an exemplary embodiment of the present invention is shown. Figure 2 A flowchart illustrating another edge-cloud collaborative target perception method according to an exemplary embodiment of the present invention is shown; Figure 3 A schematic block diagram of an edge-cloud collaborative target perception system according to an exemplary embodiment of the present invention is shown. Detailed Implementation

[0009] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.

[0010] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0011] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0012] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0013] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0014] It should be noted that the embodiments of the present invention can provide an edge-cloud collaborative target perception system, which can be used to execute the edge-cloud collaborative target perception method. That is, the executing entity of the edge-cloud collaborative target perception method provided in the embodiments of the present invention can be the edge-cloud collaborative target perception system. The edge-cloud collaborative target perception system can include an edge device (also referred to as the edge) and the cloud, which is not limited in the embodiments of the present invention. Optionally, the edge can refer to the local processing side relative to the cloud (also referred to as the cloud side), and can include device terminals and / or edge devices; the cloud side (such as a cloud server) can refer to a remote cloud computing processing side, such as a remote data center. Optionally, the edge-cloud collaboration involved in the embodiments of the present invention can include, but is not limited to, at least one of the following: collaborative processing between the edge and the cloud around ROI (Region of Interest) extraction, keyframe selection, data uploading, sparse inference, and feedback control, etc. Optionally, the edge device and the cloud can be connected via wireless or wired network communication.

[0015] Optionally, the device terminal can be a terminal device located close to the data source and directly undertaking the functions of data acquisition and / or local processing, such as airborne terminals, smart cameras, vehicle terminals, mobile terminals, etc., which can be used as one form of implementation on the edge side; correspondingly, the edge end can be a computing node deployed near the site and undertaking the functions of near-end inference, control orchestration, or data relay, such as edge boxes, edge gateways, edge servers, etc., which can be used as another form of implementation on the edge side.

[0016] Based on this, the end-side device may include one or more electronic devices, which may be a terminal (i.e., a client) or a server; optionally, the terminal mentioned herein may include, but is not limited to: smartphones, tablets, laptops, desktop computers, airborne terminals, smart cameras, vehicle terminals, etc.; the server mentioned herein may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, etc.

[0017] Optionally, the edge-cloud collaborative target perception method provided in this embodiment of the invention can be applied to any target perception scenario (also referred to as an edge-cloud collaborative target perception scenario), and this embodiment of the invention does not limit it. For example, the target perception scenario can be an urban governance scenario, an emergency management scenario, or a traffic infrastructure inspection scenario (such as traffic marking damage detection, etc.), and so on.

[0018] Based on the above description, this embodiment of the invention proposes an edge-cloud collaborative target perception method. This edge-cloud collaborative target perception method can be executed by the aforementioned edge-cloud collaborative target perception system, that is, the edge-cloud collaborative target perception method can be applied to the edge-cloud collaborative target perception system, which includes edge devices and the cloud. Figure 1 As shown, the edge-cloud collaborative target perception method may include the following steps S101-S104: S101, acquire the video data to be processed under the target monitoring task on the end device side, and acquire the task semantic information of the target monitoring task.

[0019] The process executed by the edge-cloud collaborative target perception system on the edge device side can also be called the process executed by the edge-cloud collaborative target perception system through the edge device; correspondingly, it can also be expressed as the corresponding process that the edge device in the edge-cloud collaborative target perception system can execute.

[0020] Optionally, the target monitoring task (also known as the target inspection task) can be any task, and the embodiments of the present invention do not limit it; for example, it can be a road occupation construction monitoring task, or a traffic marking damage detection task, etc.

[0021] Optionally, the edge device may store the video data to be processed (also known as the original video sequence) under the target monitoring task in its own storage space. In this case, the edge device can obtain the video data to be processed from its own storage space; or, it can obtain a download link for the video data to be processed and download the video data to be processed according to the download link to obtain the video data to be processed; or, when the edge device includes a device terminal (such as a drone), the video data to be processed can be collected through the device terminal to obtain the video to be processed, etc.; the embodiments of the present invention do not limit this.

[0022] Optionally, the edge device may store the task semantic information of the target monitoring task in its own storage space. In this case, the edge device may obtain the task semantic information from its own storage space; or, it may obtain a download link for the task semantic information of the target monitoring task and download the task semantic information using the download link; or, it may receive a task semantic setting instruction for the target monitoring task and use the task semantic information indicated by the task semantic setting instruction as the task semantic information of the target monitoring task, etc. The embodiments of the present invention do not limit this.

[0023] Optionally, task semantic information can also be referred to as Prompt information, etc.; optionally, the task semantic information of a task can be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, task semantic information can be a semantic embedding vector or a predefined description, and the predefined description may include, but is not limited to, the category of the target of interest, etc. Optionally, the edge-cloud collaborative target perception method proposed in this embodiment of the invention can also be referred to as the edge-cloud collaborative closed-loop method based on Prompt attention guidance. Based on this, this embodiment of the invention can enable the edge device to know what kind of interest area to find through task semantic information. This task semantic information guidance enables the edge device to only focus on the interest area that is meaningful to the current task, filtering out interference and irrelevant areas at the source; for example, in the traffic marking maintenance scenario (i.e., the traffic marking damage monitoring scenario), if the task semantic information indicates that attention should be paid to "road markings", the edge device will focus on extracting the road marking area and ignore vehicles, pedestrians, etc., without the need for further filtering in the cloud, greatly reducing the subsequent processing burden.

[0024] S102, based on the task semantic information, generate target attention heatmaps for each video frame to be processed in the video data to be processed; and determine the set of attention regions for each video frame to be processed based on the target attention heatmaps for each video frame to be processed.

[0025] In this embodiment of the invention, the edge device can analyze each video frame to be processed based on task semantic information to generate target attention heatmaps that can represent regions of interest. For example, a heatmap multimodal model can be used to analyze any video frame to be processed based on task semantic information to generate a target attention heatmap for any video frame to be processed, and so on. It should be noted that this embodiment of the invention does not limit the specific model structure of the heatmap multimodal model; for example, the heatmap multimodal model can be any lightweight neural network model, and so on. Based on this, this embodiment of the invention can achieve heatmap calculation through task semantic information and video frames to be processed, thereby making the target attention heatmap more accurate and more adaptable through prompts (here, task semantic information); that is, it can adapt to the corresponding task through task semantic information to identify the region of interest corresponding to the category of target interest (i.e., the monitoring object) to be monitored by the target monitoring task, and so on.

[0026] Optionally, when determining the set of attention regions for each video frame to be processed based on the target attention heatmap of each video frame to be processed, the edge device can traverse each video frame to be processed in the video data to be processed, and take the currently traversed video frame as the current video frame, and determine the attention region extraction index; wherein, the attention region extraction index can be updated according to the feedback control instructions issued by the cloud. Accordingly, the set of attention regions for the current video frame can be determined from the target attention heatmap of the current video frame according to the attention region extraction index.

[0027] Optionally, the region of interest extraction metrics may include, but are not limited to, at least one of the following: ROI area threshold (which can be used to limit the minimum ROI area), ROI selection threshold (i.e., heatmap pixel intensity threshold, also known as attention weight threshold, etc.), hysteresis threshold (which can be used to set upper and lower thresholds for attention heatmaps / scores to suppress jitter and frequent switching), region of interest expansion coefficient (which can be used to expand the region of interest indicated by the target attention heatmap according to the region of interest expansion coefficient so that the region of interest in the region of interest set is the expanded region of interest), and upper limit of the number of regions of interest extracted per frame, etc.; the embodiments of the present invention do not limit this. Optionally, the feedback control command may carry perception control feedback data, which may include, but is not limited to, at least one of the following: feedback ROI area threshold (used to control the ROI area threshold, i.e., to adjust the ROI area threshold to the feedback ROI area threshold, such as expanding or shrinking the ROI range), feedback ROI selection threshold (used to adjust the ROI selection threshold, i.e., to adjust the ROI selection threshold to the feedback ROI selection threshold, such as to increase the ROI selection threshold to converge the ROI area when bandwidth / latency is limited, or to decrease the ROI selection threshold to expand ROI coverage when recall is insufficient), feedback attention area expansion coefficient (used to adjust the attention area expansion coefficient), feedback task semantic information of the target monitoring task (used to update or switch the task semantic information of the target monitoring task), and feedback attention target category (used to change the attention target category, i.e., to change the attention target category in the task semantic information, i.e., to change the attention target category to the feedback attention target category). Other features include updating the category vocabulary or updating the semantic embedding corresponding to the category of interest, thereby enabling adjustments to the object of interest, such as switching from "road markings" to "vehicles / pedestrians / floating objects," etc.; feedback on the upper limit of the number of single-frame interest regions extracted (which can be used to adjust the upper limit of the number of single-frame interest regions extracted, i.e., to adjust the upper limit of the number of single-frame interest regions extracted to the feedback upper limit of the number of single-frame interest regions extracted); feedback on the keyframe sending frequency (which can be used to modify the keyframe sending frequency, i.e., to modify the keyframe sending frequency to the feedback keyframe sending frequency, such as to speed up or slow down the upload frame rate); and feedback on the keyframe sampling strategy (which can be used to modify the keyframe sampling strategy, i.e., to update the keyframe sampling strategy to the feedback keyframe sampling strategy, such as including but not limited to feedback on the keyframe sampling frequency and / or feedback on the specified acquisition time interval, etc., feedback on the keyframe sampling frequency can be used to modify the keyframe sampling frequency, and feedback on the specified acquisition time interval can be used to modify the specified acquisition time interval), etc.; the embodiments of the present invention do not limit these features.Correspondingly, the edge device can optimize the perception and control data through perception and control feedback data to update the perception and control data, thereby adjusting the processing of subsequent frames to form a closed-loop optimization of edge-cloud collaboration. Optionally, the perception and control data may include, but is not limited to, at least one of the following: indicators for extracting regions of interest and key frame sampling strategies, etc., which are not limited in this embodiment of the invention. Optionally, the perception and control data may be initialized according to experience or actual needs, which are not limited in this embodiment of the invention. Optionally, the model in this embodiment of the invention can support quantization, pruning, and non-stop parameter switching, which can ensure low latency and stability.

[0028] Optionally, the perception control feedback data and / or perception control data may be constrained by safety barriers (such as SLA barriers (Service Level Agreement), resource limits, etc.) and / or fallback mechanisms. That is, all feedback can be constrained by safety barriers and / or fallback mechanisms to ensure service stability. For example, when exceeding limits or with low confidence, it can automatically fall back to a conservative configuration, etc. Optionally, safety barriers, fallback mechanism constraints, conservative configurations, etc., can all be set according to experience or actual needs, and this embodiment of the invention does not limit them. Among them, the SLA barrier can deploy a set of constraints, such as maximum latency, minimum recall, bitrate limit, etc., which can be used for policy arbitration and fallback. It should be noted that this embodiment of the invention does not limit the specific content of safety barriers and fallback mechanisms, that is, they can be set according to experience or actual needs. For example, when a fallback trigger condition is detected, a fallback can be performed according to the fallback mechanism, such as restoring the default ROI selection threshold, reducing the keyframe sampling frequency, expanding the outer boundary of the region of interest, etc. Accordingly, this embodiment of the invention does not limit the fallback triggering conditions. Fallback triggering conditions may include, but are not limited to, at least one of the following: bitrate lower than a preset bitrate threshold, latency greater than a preset latency threshold, recall guardrail, etc. Optionally, the preset bitrate threshold and preset latency threshold can be set according to experience or actual needs, and this embodiment of the invention does not limit this. For example, the preset bitrate threshold can be jointly determined by the uplink available bandwidth budget and the video reporting strategy, such as converting the "maximum amount of data allowed to be reported per unit time" into the preset bitrate threshold (e.g., 80% or 90% of the maximum limit as the preset bitrate threshold); and / or, when network congestion or the concurrent use of multiple devices increases, the cloud can lower the preset bitrate threshold based on the link monitoring results, etc. As another example, the preset latency threshold can be determined by the "end-to-end processing timeliness requirements" on the service side, such as being broken down into "end-side processing latency budget + uplink transmission latency budget + cloud inference latency budget," and giving a total end-to-end latency upper limit as a guardrail, which serves as the preset latency threshold, etc.

[0029] Optionally, the cloud may include an adaptive feedback control module, which is responsible for adjusting the working parameters of the end-side device in real time based on the cloud analysis results and system status (such as indicators for the region of interest extraction, keyframe upload strategies, etc.) to achieve closed-loop control. Optionally, the adaptive feedback control module may receive at least one target region of interest perception result (also referred to as target region of interest inference result, such as the confidence level of a road marking damage) and network condition indicators (such as, but not limited to, bandwidth, latency, etc.), thereby determining perception control feedback data based on at least one target region of interest perception result and network condition indicators. For example, perception control feedback data can be determined according to a preset feedback control strategy based on at least one target region of interest perception result and network condition indicators; or, it can also be determined through reinforcement learning algorithms (such as Context-aware Multi-Agent RL (Context MARL)) based on at least one target region of interest perception result and network condition indicators, etc.; this embodiment of the invention does not limit this. Optionally, the preset feedback control strategy can be set according to experience or actual needs; this embodiment of the invention does not limit this. For example, the preset feedback control strategy may include, but is not limited to, at least one of the following: when the bitrate is observed to be continuously close to the upper limit, increase the ROI selection threshold, decrease the Top-K (the first K, where K is a positive integer), shrink the expansion coefficient of the region of interest, and extend the keyframe sampling interval (i.e., the specified acquisition time interval); when the recall is observed to be below the lower limit or the uncertainty increases, decrease the ROI selection threshold, increase the Top-K, expand the expansion coefficient of the region of interest, and shorten the keyframe sampling interval, etc. The context MARL can be a CTDE (Centralized Training, Decentralized Execution) paradigm, which can include edge-side agents and cloud-side scheduling agents. It can take states such as recall, false detection, uncertainty, latency, bitrate, energy consumption, temperature, bandwidth, packet loss, scene stability, queue load, etc., and take actions such as output threshold, Top-K, outer boundary, keyframe period or keyframe sampling frequency, QP-map (Quantization Parameter Map, a two-dimensional map that can set the quantization intensity for different regions, used for region weighted coding and bitrate allocation) or resolution, Prompt and negative prompts (such as prompting regions not to be extracted), cloud-side model depth, branch selection, etc.For example, the reward of the reinforcement learning algorithm can be: x1. Recall - x2. False Detection - x3. Latency - x4. Energy Consumption - x5. Jitter, where x1, x2, x3, x4 and x5 can be the weights of each item, and the jitter item constrains the instability caused by frequent parameter changes; optionally, each weight can be set according to SLA and / or resource budget, that is, it can be set according to experience or actual needs, or it can be adaptively tuned, and the embodiments of the present invention do not limit this. For example, when weights are set based on SLA and / or resource budget, the upper limit of bitrate, upper limit of latency, and lower limit of recall are guardrail constraints, which can have higher priority. The higher the priority, the greater the weight can be set. And / or, if resources are scarce, the weight of resource items is increased: for example, when bandwidth / computing power / energy consumption resources are scarce, the weight of bitrate / latency / energy consumption related indicators is increased, so that feedback control is more inclined to reduce transmission and computing overhead. And / or, if there is quality risk, the weight of quality items is increased: for example, when recall is insufficient or false positives are high, the weight of recall / false positives related indicators is increased, so that feedback control is more inclined to expand ROI coverage, increase sampling and inference power, and so on.

[0030] For example, when the cloud detection results are empty for a long time (e.g., no abnormal targets / no abnormalities, such as damaged areas), the system can increase the ROI selection threshold or reduce the keyframe sampling frequency to save resources. When potential abnormal targets or abnormal changes are detected (e.g., new abnormal targets or abnormal triggering conditions, such as a sudden increase in the number of abnormal targets, such as crowd gatherings or traffic congestion), the system can immediately notify the device to increase the keyframe sampling frequency or decrease the ROI selection threshold to ensure that subsequent detail capture is not missed, thus ensuring that the device can capture detailed information in a timely manner. When the cloud detects that network bandwidth is limited or latency is increasing, the feedback control command can require the device to perform image compression processing at a higher ratio (e.g., carry a larger feedback compression factor to update the compression factor, thereby achieving a higher compression ratio) or reduce keyframe uploads (e.g., reduce the keyframe sampling frequency) to alleviate network pressure, and so on. Based on this, this cloud-to-end feedback mechanism ensures the system's adaptability to scene changes: continuously filtering out spatiotemporal redundant information while not missing any important events; correspondingly, after the adaptive feedback control module generates feedback control commands, they can be sent to the end-side device via the network, where the corresponding module executes adjustments, enabling the entire system to form a self-optimizing closed loop. It is evident that this embodiment of the invention can adjust end-side strategies based on real-time conditions via the cloud, achieving dynamic balance between data flow and computing load between the end and cloud, ensuring efficient and robust system operation in different environments; based on this, it can coordinate computing power, battery life, and bandwidth even with limited single-device resources, and has the ability to smoothly scale to multi-device scenarios.

[0031] Optionally, when determining the set of attention regions for the current video frame from the target attention heatmap of the current video frame according to the attention region extraction index, if the attention region extraction index includes the ROI selection threshold, the target attention heatmap of the current video frame can be thresholded according to the ROI selection threshold to determine the initial set of attention regions for the current video frame (that is, extracting high response value regions, which can be regions with attention weights greater than the ROI selection threshold, as initial attention regions; i.e., one initial attention region can be one high response value region), and the set of attention regions for the current video frame can be determined based on the initial set of attention regions. For example, the attention weight (also called attention score) of each pixel in the target attention heatmap of the current video frame can be compared with the ROI selection threshold, and pixels with attention weights greater than the ROI selection threshold can be retained as attention regions, while pixels with attention weights less than or equal to the ROI selection threshold can be discarded. Then, connected component normalization can be performed on connected pixels that meet the ROI selection threshold (i.e., greater than the ROI selection threshold) to determine the initial set of attention regions for the current video frame. For example, when the region of interest extraction metrics also include an ROI area threshold and / or an ROI confidence threshold, connected pixels that meet the ROI selection threshold can be normalized according to the ROI area threshold and / or the ROI confidence threshold (i.e., filtering can be performed according to the ROI area threshold and / or the ROI confidence threshold). For example, regions with an area smaller than the ROI area threshold can be removed, and / or regions with a region confidence score less than the ROI confidence threshold can be removed (the region confidence score of a region can be the mean or median of the attention weights of each pixel in the corresponding region, etc.), etc. It should be noted that the specific implementation of connected component normalization in this embodiment of the invention is not limited, and may include, but is not limited to, at least one of the following: extracting connected components (grouping adjacent high-resolution pixels into an independent region), filling holes inside the region, merging adjacent discrete blocks that are too close (such as discrete blocks whose distance is less than a preset pixel distance threshold, which can be set according to experience or actual needs), contour smoothing, etc. This embodiment of the invention does not limit these aspects. Optionally, a region of interest can also be called a candidate ROI region. For example, when the target monitoring task is a lane line (i.e., road marking) damage monitoring scenario (also known as a traffic marking damage monitoring scenario), a region of interest can be a region containing the road marking, and so on.

[0032] Optionally, the set of attention regions for the current video frame can be determined based on the initial set of attention regions. Alternatively, when the attention region extraction metric includes an upper limit on the number of attention regions extracted per frame (e.g., indicating the extraction of the first K attention regions, where K is a positive integer), a set of undetermined attention regions can be selected from the initial set of attention regions. (When the number of initial attention regions in the initial set is greater than or equal to K, it may include the first K initial attention regions; when the number of initial attention regions in the initial set is less than K, it may include all initial attention regions in the initial set). Based on this undetermined set of attention regions, the set of attention regions for the current video frame can be determined, and so on. For example, the first K initial attention regions with the largest ROI area can be selected from the initial set of attention regions, or the first K initial attention regions with the largest region attention weight can be selected, etc. This embodiment of the invention does not limit this; wherein, the region attention weight of an attention region can be the average of the attention weights of each pixel in the corresponding attention region, or the median of the attention weights of each pixel in the corresponding attention region, etc.; this embodiment of the invention does not limit this.

[0033] Optionally, when determining the set of regions of interest for the current video frame based on the set of regions of interest to be determined, the set of regions of interest to be determined can be used as the set of regions of interest for the current video frame; or, when the region of interest extraction index also includes a region of interest expansion coefficient, each region of interest to be determined in the set of regions of interest to be determined can be expanded according to the region of interest expansion coefficient to obtain the set of regions of interest for the current video frame. In this case, the set of regions of interest for the current video frame may include the expanded regions of each region of interest to be determined, and so on. For example, assuming the region of interest expansion coefficient is 1.2, then the area of ​​the expanded region of a region of interest to be determined can be 1.2 times the area of ​​the corresponding region of interest to be determined, and so on.

[0034] Based on this, the set of regions of interest (ROIs) for the current video frame can be a set of ROIs filtered by one or more of the following: ROI area threshold (which can remove small noise areas), ROI confidence threshold, and upper limit on the number of ROIs extracted per frame; and / or, the ROIs in a set of ROIs can be ROIs expanded by an ROI expansion coefficient, and so on. Optionally, a set of ROIs can also be called a target set of ROIs, and the ROIs in a set of ROIs can also be called target ROIs, and so on. It should be understood that embodiments of the present invention can be guided by task semantic information, enabling the edge device to focus only on ROIs with specific semantics (i.e., the semantics indicated by the task semantic information), thereby achieving the filtering of irrelevant backgrounds.

[0035] Therefore, the embodiments of the present invention do not limit the specific implementation of determining the set of attention regions of the current video frame from the target attention heatmap of the current video frame according to the attention region extraction index.

[0036] Optionally, the edge device may include an image acquisition and ROI extraction module. In this case, the image acquisition and ROI extraction module can acquire the video data to be processed under the target monitoring task, as well as the task semantic information of the target monitoring task, and then determine the set of regions of interest and the masks of regions of interest for each video frame to be processed. For example, the image acquisition and ROI extraction module may include a video acquisition module, which can acquire real-time video data streams (such as from vehicle-mounted cameras, roadside cameras, or drones) to acquire the video data to be processed; and / or, it can also perform video frame preprocessing (such as resolution adjustment, distortion correction, etc.) on the original frames in the original video stream to acquire the video data to be processed, and so on. As another example, the image acquisition and ROI extraction module may also include an ROI extraction module to generate a target attention heatmap and extract the set of regions of interest, etc.; the ROI extraction module is equivalent to performing a semantically guided initial screening on the edge, filtering out backgrounds that are irrelevant to the task.

[0037] S103: Based on the set of interest regions of each video frame to be processed, determine the set of key frames from the video data to be processed; and based on the interest region indication data of each key frame in the set of key frames, determine the data packet to be processed in the cloud.

[0038] Among them, the region of interest indication data of a keyframe includes the region of interest indication data of each region of interest in the region of interest set of the corresponding keyframe, the region of interest indication data of a region of interest includes the region of interest image pixels and region metadata of the corresponding region of interest, and the region metadata of a region of interest includes the region location information of the corresponding region of interest.

[0039] In one implementation, when determining a set of keyframes from the video data to be processed based on the set of regions of interest (ROIs) for each video frame to be processed, the edge device can determine a region of interest (ROI) mask for any video frame to be processed based on the ROI set of that video frame. The ROI mask can be used to indicate the ROI set of that video frame (i.e., to indicate the ROI region in that video frame). For example, the pixel value of a pixel located within the ROI in the ROI mask can be a first pixel value, and the pixel value of a pixel located outside the ROI can be a second pixel value. Optionally, both the first and second pixel values ​​can be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, the first pixel value can be 1, the second pixel value can be 0, and so on.

[0040] Based on this, when a previous keyframe exists for any video frame to be processed, the degree of change of the region of interest of any video frame to be processed can be determined based on the region of interest mask of any video frame to be processed and the region of interest mask of the previous keyframe. The previous keyframe is the keyframe whose acquisition time is before the acquisition time of any video frame to be processed and is the closest to any video frame to be processed. Optionally, when determining the degree of change in the region of interest (ROI) of any video frame to be processed (also referred to as the difference index between the ROI of any video frame to be processed and the ROI of the previous keyframe) based on the ROI mask of any video frame to be processed and the ROI mask of the previous keyframe, the Intersection over Union (IoU) ratio of the ROI masks can be calculated based on the ROI masks of any video frame to be processed and the ROI masks of the previous keyframe, and the reciprocal of the IoU ratio can be used as the degree of change in the ROI of any video frame to be processed; or, the pixel change rate between the ROI mask of any video frame to be processed and the ROI mask of the previous keyframe can be determined, and the pixel change rate can be used as the degree of change in the ROI of any video frame to be processed; or, the structural similarity index Measure (SSIM) between the ROI mask of any video frame to be processed and the ROI mask of the previous keyframe can be determined, and the reciprocal of the structural similarity index can be used as the degree of change in the ROI of any video frame to be processed, etc.; the embodiments of the present invention do not limit this.

[0041] Furthermore, based on the degree of change in the area of ​​interest and the keyframe sampling strategy, it can be determined whether any video frame to be processed is considered a keyframe. If it is determined that any video frame to be processed is a keyframe, then that video frame to be processed is added to the keyframe set, thereby determining the keyframe set from the video data to be processed. Optionally, the keyframe sampling strategy can be set according to experience or actual needs, or it can be updated based on perception control feedback data (i.e., it can be updated based on feedback control commands issued from the cloud), etc.; this embodiment of the invention does not limit this. Optionally, the keyframe sampling strategy may include a specified acquisition time interval and / or a keyframe sampling frequency, and the keyframe sampling frequency can also be used to indicate a specified acquisition time interval; based on this, the keyframe sampling strategy can be used to indicate a specified acquisition time interval. For example, when determining whether any video frame to be processed is a key frame (i.e., determining whether any video frame to be processed is a key frame) based on the degree of change in the region of interest and the key frame sampling strategy, if the degree of change in the region of interest is greater than the degree of change threshold, and / or the time interval between any video frame to be processed and the previous key frame reaches the specified acquisition time interval, then any video frame to be processed can be determined to be a key frame (i.e., any video frame to be processed can be determined to be a key frame); if the region of interest is less than or equal to the degree of change threshold, and the time interval between any video frame to be processed and the previous key frame does not reach the specified acquisition time interval, then any video frame to be processed can be determined not to be a key frame (i.e., any video frame to be processed cannot be determined to be a key frame), in which case any video frame to be processed can be considered as redundant with the information of the previous key frame and skipped from uploading, etc.; the embodiments of the present invention do not limit this. Optionally, the change threshold can be set based on experience or actual needs (e.g., when uploading keyframes with significant differences, the change threshold can be set to a larger value, etc.), or it can be updated based on perception control feedback data (e.g., the perception control feedback data may include the feedback change threshold for updating the change threshold), etc.; this embodiment of the invention does not limit this. Based on this, this embodiment of the invention can introduce a dual-trigger strategy of "change degree driving + specified acquisition time interval constraint" in continuous video streams, which can effectively reduce repeated uploads.

[0042] Optionally, when the preceding keyframe of any video frame to be processed does not exist, any video frame to be processed can be used as a keyframe; or, when the area of ​​the largest region of interest in the set of regions of interest of any video frame to be processed is greater than a preset region of interest area threshold, any video frame to be processed can be used as a keyframe; or, when the number of regions of interest in the set of regions of interest of any video frame to be processed reaches a preset region of interest number threshold, any video frame to be processed can be used as a keyframe, and so on; the embodiments of the present invention do not limit this. Optionally, both the preset region of interest area threshold and the preset region of interest number threshold can be set according to experience or actual needs, and the embodiments of the present invention do not limit this.

[0043] In another implementation, for any video frame to be processed in the video data to be processed, the degree of change of the region of interest of any video frame to be processed can be determined directly according to the intersection-union ratio (IUGR) between the region of interest set of any video frame to be processed and the region of interest set of the previous keyframe (e.g., the reciprocal of the IUGR, etc.); the embodiments of the present invention do not limit this. Accordingly, based on the degree of change of the region of interest of any video frame to be processed and the keyframe sampling strategy, it can be determined whether any video frame to be processed is a keyframe; if it is determined that any video frame to be processed is a keyframe, then any video frame to be processed is added to the keyframe set to realize the determination of the keyframe set from the video data to be processed.

[0044] In another implementation, it is also possible to determine whether a video frame to be processed is a keyframe based solely on the degree of change in the region of interest of any video frame to be processed. If the degree of change in the region of interest of any video frame to be processed is greater than the degree of change threshold, then it can be determined that any video frame to be processed is a keyframe, and so on. Accordingly, if it is determined that any video frame to be processed is a keyframe, then the video frame to be processed is added to the keyframe set to achieve the determination of the keyframe set from the video data to be processed, and so on.

[0045] Optionally, the set of regions of interest for a video frame to be processed can be represented by the bounding rectangle of each region of interest in the set of regions of interest for the corresponding video frame to be processed, or by the binary mask of the corresponding video frame to be processed (in which case the mask of the region of interest for the corresponding video frame to be processed can be obtained directly), or by the outline point set of each region of interest, or by grid blocks, etc.; the embodiments of the present invention do not limit this.

[0046] Optionally, the number of data packets to be processed in the cloud can be one or more, and this embodiment of the invention does not limit this.

[0047] Optionally, when determining the cloud-based data packet to be processed based on the region of interest (ROI) indication data of each keyframe in the keyframe set, the ROI indication data of each keyframe can be added to the same remote data packet to be processed. In this case, after determining the ROI indication data of each keyframe, the ROI indication data of each keyframe can be uploaded to the cloud simultaneously. That is, one video data to be processed can correspond to one cloud-based data packet to be processed. For example, if a video data to be processed is a 15-second video, the ROI indication data of all keyframes within that 15-second video can be added to the same cloud-based data packet to be processed. Alternatively, the ROI indication data of each keyframe can be determined separately according to the keyframe transmission frequency. Domain indication data is added to the current cloud-based pending data packet. Then, each cloud-based pending data packet can be sent sequentially according to the keyframe sending frequency. For example, if the keyframe sending frequency indicates sending cloud-based pending data packets once per minute, the current cloud-based pending data packet can be uploaded to the cloud every minute. Alternatively, one keyframe can correspond to one cloud-based pending data packet. In this case, for each determined keyframe, the attention area indication data of the currently determined keyframe can be added to a cloud-based pending data packet, thereby reporting the cloud-based pending data packet corresponding to the current keyframe to the cloud. In this case, the keyframe sampling frequency can be the same as the keyframe sending frequency, and so on. This embodiment of the invention does not limit this. Optionally, the keyframe sending frequency can be set according to experience or actual needs, or it can be updated based on feedback keyframe sending frequency, etc. This embodiment of the invention does not limit this.

[0048] In this context, the image pixels of a region of interest can also be referred to as the image pixels of the corresponding region of interest. The image pixels of a region of interest can include the pixel values ​​of each pixel point within the corresponding region of interest. For example, the region of interest can be cropped for any keyframe to obtain the image pixels of the region of interest for any keyframe, and then only the region images of each region of interest in the set of regions of interest for any keyframe can be retained.

[0049] Optionally, the regional location information of a region of interest may be the regional location coordinates of the corresponding region of interest (such as the coordinates of the diagonal points of the region, or the coordinates of the upper left corner of the region, length and width, etc.), or the contour point set of the corresponding region of interest, or the polygonal representation of the corresponding region of interest (i.e., an ordered sequence of vertex coordinates), etc.; the embodiments of the present invention do not limit this.

[0050] Optionally, the region metadata (also known as alignment metadata) of a region of interest may include, but is not limited to, at least one of the following: frame alignment information of the corresponding region of interest (used to indicate the frame to which the region of interest belongs), original image size (i.e., the size of the frame to which the corresponding region of interest belongs, such as the target video frame size), intrinsic and extrinsic parameters of the video acquisition camera, region of interest expansion coefficient (also known as region of interest expansion coefficient, such as one that can be used to expand by 0.05-0.2 times), scaling factor (used to scale the pixels of the image in the region of interest, such as a compression factor), super-window index (used to indicate the super-window region to which it belongs), etc.; the embodiments of the present invention do not limit this. For example, the frame alignment information of a region of interest may be the frame identifier (such as frame ID or timestamp) of the frame to which the corresponding region of interest belongs, etc. Based on this, consistent decoding and reconstruction between the end and cloud can be determined through region metadata, thereby effectively adapting to bandwidth budget, etc.

[0051] Based on this, embodiments of the present invention can combine content changes with time limits through keyframe dual triggering, hysteresis, budget awareness, etc., to robustly control the reporting frequency and link resource status.

[0052] S104, upload the data packet to be processed in the cloud to the cloud so that target perception can be performed on the cloud side based on the data packet to be processed in the cloud, and obtain the target interest area perception result.

[0053] Based on this, the edge device can upload one or more cloud-based data packets corresponding to the video data to be processed to the cloud, thereby realizing the upload of cloud-based data packets to the cloud.

[0054] Optionally, target perception may include target detection and / or recognition reasoning. Accordingly, after receiving the data packet to be processed in the cloud, target detection and / or recognition reasoning can be performed through the data packet to be processed in the cloud. In this case, the cloud can perform sparse computation only on the ROI region to obtain the target interest region perception result of the ROI region, thereby realizing the transmission of only useful regions and the calculation of only regions of interest.

[0055] Optionally, the cloud can also output area-of-interest perception output data based on the target area-of-interest perception results. Optionally, the area-of-interest perception output data may include, but is not limited to, at least one of the following: the identification category of the monitored target (such as road markings, etc.) (e.g., damaged or intact), location information (indicating the location of the monitored target), and abnormal situation reports, etc., which are not limited in this embodiment. Optionally, the target area-of-interest perception results can be used as the area-of-interest perception output data; or, the area-of-interest perception output data can be output based on the target area-of-interest perception results according to preset output rules, etc., which are not limited in this embodiment. Optionally, the preset output rules can be set according to experience or actual needs, which are not limited in this embodiment. For example, the cloud can aggregate the target area-of-interest perception results to output area-of-interest perception output data (which may include conclusions from visual detection and / or identification, such as the location and extent of road marking damage), for use by the business layer.

[0056] Optionally, the edge device may also include a keyframe filtering and sending module. In this case, the edge device can use the keyframe filtering and sending module to adaptively select keyframes based on ROI changes and upload the corresponding region of interest indication data to the cloud. For example, the keyframe filtering and sending module may include a keyframe adaptive filtering module and a communication sending module; wherein, the keyframe adaptive filtering module can be used to filter keyframes (such as determining a set of keyframes), and the communication sending module can be used to determine and upload data packets to be processed to the cloud.

[0057] Optionally, the cloud-based component may include, but is not limited to, a ROI sparse inference module and a feedback control module. Optionally, the ROI sparse inference module can receive data packets to be processed from the cloud and utilize a pre-trained model to perform target perception (such as target detection and / or recognition inference) only on the ROI region, obtaining detection results (i.e., target interest region perception results). The feedback control module can generate feedback control instructions based on the detection results and / or preset feedback control strategies, and send them to the edge device via the network to adjust the edge device's interest region extraction metrics and / or keyframe sampling strategies, thereby achieving closed-loop control between the cloud and the edge. Based on this, the pre-trained model, knowing the region of interest (i.e., ROI) in advance, can employ sparse computation strategies, such as setting non-ROI regions in the input image to zero and skipping them, or performing convolution calculations only within the ROI bounding box, thereby reducing useless computation, etc.

[0058] As can be seen, this embodiment of the invention can effectively reduce redundant data transmission by uploading only frames containing new or significantly changed information through this adaptive frame selection strategy. Furthermore, the end device and the cloud can maintain a reference relationship for the same frame, enabling the cloud to map the ROI mask back to global image coordinates based on region metadata. This allows the cloud to restore the position and range of the ROI in the original image when needed. Specifically, it can backfill the position and size of the ROI on a canvas of the same size as the original video frame based on the region metadata, and / or backfill the image pixels of the region of interest (non-interested regions can be filled with default values, such as 0), to understand the received ROI content. This significantly reduces the amount of video data transmitted by transmitting only the region of interest indication data, thus significantly reducing bandwidth usage. Correspondingly, efficient and accurate target perception can be achieved through an end-to-cloud collaborative target perception method. Target perception can include target detection and / or anomaly recognition; that is, this embodiment of the invention can achieve efficient and accurate target detection and / or anomaly recognition.

[0059] Based on this, the embodiments of the present invention fully combine the semantic prior filtering of the edge device with the deep analysis capabilities of the cloud, realizing complementary collaboration between edge and cloud information: the edge device can be responsible for the macro-level task-related data filtering, while the cloud can be responsible for the precise perception of micro-details and continuously optimize the edge strategy through feedback; accordingly, the system can filter a large amount of redundant spatiotemporal information on the edge, extracting only the content meaningful to the task for uploading, while the cloud can efficiently process and guide the edge dynamic optimization, thereby maximizing the use of communication and computing resources while ensuring analysis accuracy. For example, transmitting only ROI region-related information can reduce the amount of data by more than 50%, while maintaining a high level of detection accuracy; the cloud's feedback adjustment enables the system to automatically adjust parameters according to scene changes, improving robustness to adverse factors such as changes in lighting and noise interference. Furthermore, through the combination of semantic guidance and closed-loop control, the embodiments of the present invention can improve the efficiency and reliability of video analysis while expanding the application depth of artificial intelligence technology in fields such as intelligent traffic inspection. Based on this, the embodiments of the present invention can significantly reduce bandwidth consumption and transmission volume through ROI adaptive encoding (i.e., determining the data indicating the region of interest).

[0060] This invention allows for the acquisition of video data to be processed under a target monitoring task, as well as the task semantic information of the target monitoring task, on the edge device side. Then, based on the task semantic information, a target attention heatmap for each video frame to be processed can be generated. Based on the target attention heatmaps of each video frame, a set of regions of interest for each video frame can be determined. Further, based on the set of regions of interest for each video frame, a set of keyframes can be determined from the video data to be processed. Based on the region of interest indication data of each keyframe in the keyframe set, a cloud-based data packet to be processed can be determined. The region of interest indication data for a keyframe includes the region of interest indication data for each region in the corresponding keyframe's region of interest set; the region of interest indication data for a region includes the region of interest image pixels and region metadata; and the region metadata for a region of interest includes the region location information. Based on this, the cloud-based data packet to be processed can be uploaded to the cloud for target perception on the cloud side, obtaining the target region of interest perception result. As can be seen, the embodiments of the present invention can accurately determine the set of attention regions of each video frame to be processed by the target attention heatmap of each video frame to be processed, and then accurately determine the set of key frames. This allows the cloud-based data packets with small data volume to be processed to be uploaded to the cloud, thereby effectively reducing redundant data transmission and invalid cloud computing. It can reduce the overall system latency and bandwidth usage while ensuring the accuracy of target perception. In other words, the embodiments of the present invention can effectively improve the accuracy of target perception while ensuring low latency.

[0061] Based on the above description, this embodiment of the invention also proposes a more specific edge-cloud collaborative target perception method. Accordingly, this edge-cloud collaborative target perception method can be executed by the aforementioned edge-cloud collaborative target perception system, which may include edge devices and the cloud, etc. Please refer to [link to relevant documentation]. Figure 2 The edge-cloud collaborative target perception method may include the following steps S201-S206: S201, acquire the video data to be processed under the target monitoring task on the end device side, and acquire the task semantic information of the target monitoring task.

[0062] S202, for any video frame to be processed in the video data to be processed, determine at least one superwindow size, and divide the video frame to be processed into superwindows according to each of the at least one superwindow size, so as to obtain the set of superwindow regions of any video frame to be processed under each superwindow size.

[0063] Here, "superwindow size" refers to the dimensions of the superwindow, which can be a large window (also called an expanded window), such as a window larger than the base patch (image patch, such as 16×16). Correspondingly, "superwindow partitioning" can refer to the partitioning of large windows, i.e., non-overlapping large window partitioning, also known as tiling superwindow segmentation or non-overlapping large window feature partitioning. Furthermore, a superwindow region can be a large window region, that is, the region obtained after superwindow partitioning.

[0064] Optionally, at least one hyperwindow size can be set based on experience or actual needs, or it can be updated according to feedback control instructions (i.e., the perception control feedback data may also include at least one feedback hyperwindow size for updating at least one hyperwindow size), etc.; the embodiments of the present invention do not limit this. For example, a hyperwindow size can be aligned with the basic image block step size (i.e., basic step size) / patch size of the image encoder, such as selecting an integer multiple of the basic step size; for example, at least one hyperwindow size can be selected from a candidate set (such as 2 times, 3 times, and 4 times the patch size) to simultaneously cover small targets and large semantic regions; and / or, the feedback control instructions issued by the cloud can select at least one enabled hyperwindow size from the above candidate set, etc. Among them, one hyperwindow size can correspond to one hyperwindow scale. In this embodiment of the invention, when there are multiple superwindow sizes within at least one superwindow size, the target attention heatmap can be generated through multi-scale tiling superwindows. A multi-scale tiling superwindow refers to a set of windows that divide the original frame at different scales in a non-sliding, full-frame coverage manner, used to generate a full-frame semantic heatmap with low redundancy. Specifically, dividing any video frame to be processed into superwindows according to each of the at least one superwindow size to obtain a set of superwindow regions for each video frame to be processed at each superwindow size means: for the same video frame to be processed (such as any video frame to be processed), dividing it into superwindows according to each superwindow size; then each superwindow size corresponds to a set of superwindow regions.

[0065] Optionally, for any of the at least one superwindow size, any video frame to be processed can be divided into superwindows according to any superwindow size to obtain an initial superwindow set of any video frame to be processed under any superwindow size, and based on the initial superwindow set of any video frame to be processed under any superwindow size, the superwindow region set of any video frame to be processed under any superwindow size can be determined. Optionally, for any initial superwindow in the initial superwindow set, any initial superwindow can be added to the superwindow region set of any video frame to be processed under any superwindow size, so that any initial superwindow is used as a superwindow region; or, according to the superwindow expansion coefficient (such as narrow redundancy width, used to set the narrow redundancy boundary of the superwindow (also called narrowband halo), etc., a narrow-width annular outer buffer around the edge of the superwindow is the expanded edge transition band, with very narrow bandwidth and not participating in the final feature output), any pair of initial superwindows can be expanded to obtain the expansion region corresponding to any initial superwindow, and then the expansion region corresponding to any initial superwindow can be added to the superwindow region set of any video frame to be processed under any superwindow size, etc.; the embodiments of the present invention do not limit this. Optionally, the superwindow expansion coefficient can be set according to experience or actual needs, and the embodiments of the present invention do not limit this. Optionally, in other embodiments, the backbone feature map of any video frame to be processed can be divided into windows according to any window size to achieve window division of any video frame to be processed, etc.; the present invention does not limit this. Based on this, a tiling window division can be performed starting from the coordinate starting point of any video frame to be processed (i.e., the coordinate starting point of the original frame) and according to the step size indicated by any window size.

[0066] S203, determine the image patch size corresponding to each hyperwindow size, and based on the task semantic information, the image patch size corresponding to each hyperwindow size, and the hyperwindow region set of any video frame to be processed under each hyperwindow size, determine the target attention heatmap of any video frame to be processed.

[0067] Optionally, the image patch size corresponding to a superwindow size can be set according to experience or actual needs, or it can be determined based on a preset superwindow size threshold (i.e., path selection), etc.; this embodiment of the invention does not limit this. Optionally, the image patch sizes corresponding to different superwindow sizes can be the same or different; this embodiment of the invention does not limit this. Optionally, the preset superwindow size threshold can be set according to experience or actual needs; this embodiment of the invention does not limit this.

[0068] For example, when determining the corresponding image patch size according to the path selection, the edge device can determine the image patch size corresponding to each hyperwindow size through an adaptive tokenization dual-path scheme. The adaptive tokenization dual-path can include fine-grained paths and coarse-to-fine paths. Fine-grained paths can align the token granularity with the backbone step size (such as being consistent with the scale of common patches) and adapt to fine-grained regions. For example, the image patch size in the fine-grained path can be the first image patch size (i.e., P×P, or P, where P is a positive integer, such as P=16, etc.). Then, the image patch size corresponding to the hyperwindow size that conforms to the fine-grained path can be the first image size. Correspondingly, the coarse-to-fine path can divide the hyperwindow area into blocks without sliding at a second image patch size step size (such as kP, where k can be any integer, such as k=2, 3, etc.). Then, the image patch size corresponding to the hyperwindow size that conforms to the coarse-to-fine path can be the second image patch size (i.e., kP). Here, k can represent the magnification factor, which can be set according to experience or actual needs. For example, for a coarse-to-fine path, ROIAlign (Region of Interest Alignment, a bilinear interpolation alignment operator that aligns irregular ROIs to a fixed-length grid) can be performed on each coarse block (i.e., a block divided according to the second image block size) without parameters (also known as ROIAlign alignment without learning parameters, such as bilinear interpolation), to obtain P×P resampled blocks, which are then projected into tokens through patch-embedding shared with the fine path; a sparse token grid with a step size of the second image block size (also represented as Pcoarse) is formed, and the internal resolution of each token is consistent with the backbone patch (i.e., P×P), enabling consistent token representation across scales. Optionally, both the first and second image block sizes can be set according to experience or actual needs, and this embodiment of the invention does not limit this. The alignment operator mentioned above can be a fixed operator rather than a learnable module.

[0069] Optionally, for any of the at least one superwindow size, when any superwindow size is greater than a preset superwindow size threshold, the second image patch size can be used as the image patch size corresponding to any superwindow size, i.e., a coarse-to-fine path can be selected; when any superwindow size is less than or equal to the preset superwindow size threshold, the first image patch size can be used as the image patch size corresponding to any superwindow size, i.e., a fine path can be selected, and so on. Optionally, when the number of superwindow sizes in the at least one superwindow size is one, and the at least one superwindow size supports being updated according to feedback control commands, the above two paths can be switched online according to the load of the end device and the feedback from the cloud. For example, when resources are limited, the superwindow size in the at least one superwindow size can be larger, and a coarse-to-fine path can be selected; when resources are sufficient, the superwindow size in the at least one superwindow size can be smaller, and a fine path can be selected, and so on.

[0070] For example, the edge device can also acquire the data corresponding to the superwindow size and the image block size, and determine the image block size corresponding to each superwindow size in at least one superwindow size from the superwindow size corresponding data, and so on. Optionally, the superwindow size corresponding data can be set according to experience or actual needs, and the embodiments of the present invention do not limit this; for example, the superwindow size corresponding data may include the image block size corresponding to each superwindow size in the superwindow size set, and so on.

[0071] Optionally, when determining the target attention heatmap of any video frame to be processed based on task semantic information, the image patch size corresponding to each superwindow size, and the set of superwindow regions of any video frame to be processed under each superwindow size, for any superwindow size in at least one superwindow size, and any superwindow region in the set of superwindow regions of any video frame to be processed under any superwindow size, the edge device can perform local attention feature extraction on any superwindow region according to the image patch size corresponding to any superwindow size and task semantic information to obtain the superwindow feature of any superwindow region; after obtaining the superwindow features of each superwindow region of any video frame to be processed under each superwindow size, the target attention heatmap of any video frame to be processed can be determined based on the superwindow features of each video frame to be processed under each superwindow size.

[0072] Optionally, the aforementioned heatmap multimodal model may include, but is not limited to, the initial attention modules corresponding to each hyperwindow size; optionally, the initial attention modules corresponding to different hyperwindow sizes may be the same module or different modules, and this embodiment of the invention does not limit this; or, the heatmap multimodal model may also include, but is not limited to, the initial attention modules corresponding to each image patch size in at least one image patch size, and the initial attention modules corresponding to different image patch sizes may be the same module or different modules, and this embodiment of the invention does not limit this, etc. For example, the heatmap multimodal model may also include any shared modules other than the initial attention modules (such as fusion modules, etc.), etc.; optionally, different initial attention modules may also share a semantic encoding module (such as a text feature extraction module, etc.), etc. Based on this, this embodiment of the invention does not limit the specific model structure of the heatmap multimodal model. Optionally, the edge device can invoke a heatmap multimodal model to determine the target attention heatmap of any video frame to be processed based on task semantic information, the image patch size corresponding to each hyperwindow size, and the hyperwindow region set of any video frame to be processed under each hyperwindow size. For example, the task semantic information, the image patch size corresponding to each hyperwindow size, and the hyperwindow region set of any video frame to be processed under each hyperwindow size can be input into the heatmap multimodal model to output the target attention heatmap of any video frame to be processed. Alternatively, the heatmap multimodal model can be invoked to generate the target attention heatmap of any video frame to be processed based on task semantic information. For example, the task semantic information and any video frame to be processed can be input into the heatmap multimodal model to output the target attention heatmap of any video frame to be processed. In this case, at least one hyperwindow size and the hyperwindow region set of any video frame to be processed under each hyperwindow size can all be determined by the heatmap multimodal model, etc. The embodiments of the present invention do not limit this.

[0073] Optionally, an attention module may include, but is not limited to, a cross-attention module (such as a lightweight cross-attention module, also known as a prompt-adaptor), which can use the image token sequence (i.e., the token sequence of image patches) within the window (i.e., the superwindow) as Q (Query, query vector), and a vector composed of one or more of the following: task semantic information embedding, category prototype, negative cues (such as content that does not need to be detected, which may include a set of cues that suppress irrelevant semantics) as K (Key, key vector) and V (Value, value vector); for example, a lightweight cross-attention module can be used to illustrate, which can perform single-layer, few-head multi-head cross-attention (Multi Cross Attention (MCA, e.g., ≤4 attention heads). In this embodiment of the invention, attention interactions do not occur between different hyperwindows, and they can converge during the fusion stage. Optionally, to control computational power, low-rank projection / bottleneck mapping can be used on the key and value vector sides, combined with FiLM (Feature-wise Linear Modulation) channel gating, residual and layer normalization (LayerNorm, LN). Optionally, to alleviate boundary truncation, a narrowband halo can be set for each hyperwindow (i.e., hyperwindow region) and only the core region can be output (i.e., computation / attention is only performed in the non-interfering core region at the center of the hyperwindow, discarding the outer narrowband halo and not sending it to subsequent alignment, fusion, and classification branches); feathered weighted fusion (which may include normalization) is performed in the overlapping band of adjacent core regions according to the distance from their respective boundaries to eliminate seams and improve cross-window continuity. For example, an attention module may also include, but is not limited to, at least one of the following: a small Transformer (a deep neural network structure based on a self-attention mechanism), a CNN (Convolutional Neural Network), etc. Based on this, embodiments of the present invention can introduce cross-attention between the Prompt vector (i.e., the feature vector of the task semantic information) and the image features during the feature extraction stage, so that the network assigns higher weights to the image regions related to the Prompt. This mechanism can efficiently generate semantically related attention heatmaps on the edge, thereby accurately delineating the regions of interest. This allows even small models to achieve high recall and low false positives for specific targets, and can achieve the effect of customized detection with a small computational cost.

[0074] Optionally, the super-window feature of any super-window region can be determined by the target attention module corresponding to that super-window region. Correspondingly, the edge device can also determine the initial attention module corresponding to any super-window region. When a feature enhancement requirement for any super-window region is detected, a feature enhancement module is inserted into the initial attention module corresponding to that super-window region to obtain the target attention module corresponding to that super-window region. When no feature enhancement requirement for any super-window region is detected, the initial attention module corresponding to that super-window region can be used as the target attention module corresponding to that super-window region. Alternatively, the initial attention module corresponding to any super-window region can be directly used as the target attention module corresponding to that super-window region, and so on. This embodiment of the invention does not limit this. Based on this, when extracting local attention features for any super-window region, the target attention module corresponding to any super-window region can be determined through a heatmap multimodal model, thereby further extracting local attention features for any super-window region through the target attention module corresponding to that super-window region, and so on.

[0075] Optionally, when the heatmap multimodal model includes initial attention modules corresponding to each hyperwindow size, the initial attention module corresponding to the hyperwindow size (i.e., the hyperwindow size used to divide any hyperwindow region) of any hyperwindow region can be used as the initial attention module corresponding to any hyperwindow region; or, when the heatmap multimodal model includes initial attention modules corresponding to each image patch size, the initial attention module corresponding to the image patch size (i.e., the image patch size corresponding to the corresponding hyperwindow size) of any hyperwindow region can be used as the initial attention module corresponding to any hyperwindow region; or, when the heatmap multimodal model includes only one initial attention module (i.e., the initial attention modules corresponding to each hyperwindow size or image patch size are all the same module), this initial attention module can be used as the initial attention module corresponding to any hyperwindow region, etc.; this embodiment of the invention does not limit this. Optionally, when the computing power budget allows, the feature enhancement requirement for any detected hyperwindow region can be determined; or, when the uncertainty of any detected hyperwindow region is greater than a preset uncertainty threshold, the feature enhancement requirement for any detected hyperwindow region can be determined, etc.; this embodiment of the invention does not limit this. Optionally, the preset uncertainty threshold can be set according to experience or actual needs, and the embodiments of the present invention do not limit this.

[0076] Optionally, the uncertainty of any super-window region can be determined by the boundary overlap information entropy method (in which case the larger the entropy value, the higher the uncertainty of the super-window boundary), or by the coordinate residual quantization method (in which case the larger the residual, the higher the spatial positioning uncertainty), or by the characteristic variance fluctuation method (in which case the larger the variance, the more obvious the abrupt change in edge features, and the higher the uncertainty), etc.; the embodiments of the present invention do not limit this.

[0077] Optionally, the feature enhancement module can be a lightweight context mixing layer to enhance local consistency and suppress noise. Optionally, the feature enhancement module can be a single-layer window MHSA (Multi-Head Self-Attention, such as heads ≤ 4), or depthwise convolution, mean pooling, or linear attention, etc.; this embodiment of the invention does not limit this. Optionally, the feature enhancement module can be inserted before the cross-attention module (for example, the feature enhancement module can be used to mix the tokens within the window before performing cross-attention, etc.), or it can be inserted after the cross-attention module; this embodiment of the invention does not limit this.

[0078] Optionally, when extracting local attention features from any hyperwindow region according to the image patch size corresponding to any hyperwindow size and the task semantic information to obtain the hyperwindow features of any hyperwindow region, a semantic feature vector can be determined based on the task semantic information (which can be used as a key vector and value vector, or even based on negative cues, etc.). Furthermore, the image patch of any hyperwindow region can be divided into multiple hyperwindow partitioned image patches according to the image patch size corresponding to any hyperwindow size, resulting in multiple hyperwindow partitioned image patches for any hyperwindow region. Based on these multiple hyperwindow partitioned image patches and the semantic feature vector, the hyperwindow features of any hyperwindow region are determined. Thus, local attention features can be extracted from any hyperwindow region based on these multiple hyperwindow partitioned image patches and the semantic feature vector to obtain the hyperwindow features of any hyperwindow region. For example, in low-light, glare, rain, and fog scenarios, negative cues, ground / sky occlusion, etc., can be enabled so that the task semantic information can include negative cues, thereby effectively enhancing stability in harsh scenes.

[0079] In one implementation, when the image patch size corresponding to any hyperwindow size is the target image patch size (e.g., the size indicated by the target backbone step size of the heatmap multimodal model), the target attention module corresponding to any hyperwindow region can be directly invoked. Based on the task semantic information, local attention features are extracted from any hyperwindow region to obtain the hyperwindow features of any hyperwindow region. For example, the task semantic information and any hyperwindow region (i.e., the hyperwindow region image) can be input into the target attention module corresponding to any hyperwindow size, so that the hyperwindow features of any hyperwindow region can be output through the target attention module corresponding to any hyperwindow size. In this case, the above-mentioned image patch division, etc., can all be performed through the target attention module corresponding to any hyperwindow size. Optionally, the target image patch size and the target backbone step size can both be set according to experience or actual needs. For example, the target image patch size can be the first image patch size mentioned above, etc.; the embodiments of the present invention do not limit this. For example, at this time, within any superwindow region, a set of P×P image blocks can be obtained by performing a non-slipping, whole-grid segmentation according to the target image block size (e.g., P×P) and stride P (e.g., P can be the patch scale of the backbone network, preferably P∈{14,16}). Each block is mapped to a token of the same dimension through shared patch-embedding (e.g., convolution, linear projection) to form a dense token grid, so as to obtain the superwindow feature of any superwindow region.

[0080] Based on this, when determining the superwindow features of any superwindow region based on multiple superwindow-divided image blocks and semantic feature vectors, feature extraction (such as convolution, linear projection, etc.) can be performed on each superwindow-divided image block in the multiple superwindow-divided image blocks of any superwindow region to obtain the image block features of each superwindow-divided image block of any superwindow region (which can also be represented as image block tokens). Thus, the superwindow features of any superwindow region can be determined by using the image block features and semantic feature vectors of each superwindow-divided image block of any superwindow region. For example, the image block feature sequence of each superwindow-divided image block of any superwindow region can be used as a query vector, and the semantic feature vector can be used as a key vector and a value vector. Cross-attention can be performed using the query vector, key vector, and value vector to achieve local attention feature extraction of any superwindow region, and so on. For example, local attention feature extraction can be performed using the image patch features and semantic feature vectors of each superwindow-divided image block within any superwindow region. This yields the attention feature extraction results for each superwindow-divided image block (which can also be represented as tokens after local attention feature extraction). The superwindow features of any superwindow region can then be determined using these attention feature extraction results. For example, for any point (i.e., a pixel) within any superwindow region, the attention feature extraction result of the superwindow-divided image block containing that point can be used as the feature of that point, thus obtaining the superwindow features of any superwindow region. In other words, the attention feature extraction result of any superwindow-divided image block can be used as the feature of each point within that superwindow-divided image block. Alternatively, the token-level relevance score of any superwindow region can be determined using the attention feature extraction results of each superwindow-divided image block (which may include the token-level relevance scores of each superwindow-divided image block) and then dimensionality-reduced to that of any superwindow region. The thermal icon quantity of the domain (which may include the thermal icon quantity of each hyperwindow-divided image block) is used to backfill or upsample the thermal icon quantity of any hyperwindow region (e.g., backfill or upsample according to the hyperwindow-original image coordinate relationship) to obtain the window thermal map of any hyperwindow region (that is, the thermal icon quantity of each hyperwindow-divided image block is tiled and backfilled to obtain the window thermal map of any hyperwindow region). The window thermal map of any hyperwindow region is used as the hyperwindow feature of any hyperwindow region, etc. The specific implementation method of determining the hyperwindow feature of any hyperwindow region using the attention feature extraction results of each hyperwindow-divided image block is not limited in the embodiments of the present invention.

[0081] In another implementation, when the image patch size corresponding to any hyperwindow size is not the target image patch size, if the image patch partitioning size within the target attention module corresponding to any hyperwindow region is the same as the image patch size corresponding to any hyperwindow size, the target attention module corresponding to any hyperwindow region can be directly invoked. Based on the task semantic information, local attention features are extracted from any hyperwindow region to obtain the hyperwindow features of any hyperwindow region. In this case, local attention features can be extracted from any hyperwindow region according to the image patch size corresponding to any hyperwindow size. In this case, the initial attention module and the target attention module corresponding to different image patch sizes can be different. For example, the image patch partitioning size within the initial attention module corresponding to an image patch size can be the corresponding image patch size.

[0082] Based on this, when determining the superwindow features of any superwindow region based on multiple superwindow-divided image blocks and semantic feature vectors of any superwindow region, feature extraction can also be performed on each superwindow-divided image block in the multiple superwindow-divided image blocks of any superwindow region to obtain the image block features of each superwindow-divided image block of any superwindow region. Thus, the superwindow features of any superwindow region can be determined by using the image block features and semantic feature vectors of each superwindow-divided image block of any superwindow region.

[0083] In another implementation, if the image patch partitioning size within the target attention module corresponding to any hyperwindow region differs from the image patch size corresponding to any hyperwindow size (in this case, the image patch partitioning size within the target attention module corresponding to any hyperwindow region can be the target image patch size), then parameterless alignment can be performed on each hyperwindow partitioning image patch in the multiple hyperwindow partitioning image patches of any hyperwindow region to obtain the parameterless alignment result of each hyperwindow partitioning image patch of any hyperwindow region. Then, the target attention module corresponding to any hyperwindow region is called, and based on the semantic feature vector and the parameterless alignment result of each hyperwindow partitioning image patch of any hyperwindow region, local attention feature extraction is performed on any hyperwindow region to obtain the hyperwindow feature of any hyperwindow region, and so on. Previously, local attention feature extraction could be performed according to the target image patch size.

[0084] Based on this, when determining the superwindow features of any superwindow region based on multiple superwindow-divided image blocks and semantic feature vectors, parameterless alignment can be performed on each superwindow-divided image block in the multiple superwindow-divided image blocks of any superwindow region to obtain the parameterless alignment results of each superwindow-divided image block of any superwindow region. Local attention features can be extracted according to the semantic feature vector and the parameterless alignment results of each superwindow-divided image block of any superwindow region (also called parameterless aligned image blocks) to obtain the alignment attention feature extraction results of each superwindow-divided image block of any superwindow region. Thus, the superwindow features of any superwindow region can be determined using the alignment attention feature extraction results of each superwindow-divided image block of any superwindow region. For example, the alignment attention feature extraction results of each hyperwindowed image block in any hyperwindow region can be used to determine the features of each position point in the parameterless alignment result of each hyperwindowed image block in any hyperwindow region (e.g., the alignment attention feature extraction results of any hyperwindowed image block can be used as the features of each position point in the parameterless alignment result of any hyperwindowed image block). Furthermore, the features of each position point in the parameterless alignment result of each hyperwindowed image block in any hyperwindow region can be used to determine the features of each position point in each hyperwindowed image block in any hyperwindow region (e.g., the features of any position point in the parameterless alignment result of any hyperwindowed image block can be used as the features of each position point in a corresponding grid region within any hyperwindowed image block, thereby mapping each position in the small-scale parameterless alignment result to fill the original large-scale hyperwindowed image block). This can be achieved by determining the superwindow features of any superwindow region; alternatively, the alignment attention feature extraction results of each superwindow segmented image block in any superwindow region can be used to determine the token-level relevance score of each superwindow segmented image block in any superwindow region (also known as the token-level relevance score of the parameterless alignment result of the superwindow segmented image block), and the dimensionality can be reduced to the heat map quantity of each superwindow segmented image block in any superwindow region (the heat map quantity of a superwindow segmented image block can also be known as the heat map quantity of the parameterless alignment result of the corresponding superwindow segmented image block), thereby backfilling the heat map quantity of each superwindow segmented image block to obtain the window heat map of any superwindow region (the backfilling process can be the same as the above alignment attention feature extraction result), so as to use the window heat map of any superwindow region as the superwindow feature of any superwindow region, etc.; the embodiments of the present invention do not limit this.

[0085] Optionally, when determining the target attention heatmap of any video frame to be processed based on the superwindow features of each superwindow size, if a superwindow feature is a window heatmap, the attention heatmap of any video frame to be processed can be determined based on the superwindow features of each superwindow size (i.e., the attention heatmap of any video frame to be processed can be determined by using the superwindow features of any video frame to be processed at any superwindow size, such as by stitching and fusing the superwindow features of any video frame to be processed at any superwindow size, such as by weighted summation of overlapping areas, etc.), and the attention of any video frame to be processed at each superwindow size can be determined. Heatmaps are fused to obtain an initial attention heatmap for any video frame to be processed. The target attention heatmap for any video frame to be processed is then determined using the initial attention heatmap. Alternatively, for any pixel in any video frame to be processed, at least one heatmap value can be determined from the various superwindow features of the video frame at various superwindow sizes. The at least one heatmap value of any pixel is then fused (e.g., by weighted summation) to obtain the heatmap value of any pixel, thereby obtaining the initial attention heatmap for any video frame to be processed. The target attention heatmap for any video frame to be processed is then determined using the initial attention heatmap. Based on this, it is possible to perform intra-scale normalization on heatmaps at various scales and then weightedly fuse them to the original resolution to obtain an initial attention heatmap for any video frame to be processed (which can be from the same source and scale as the original image). For example, intra-scale normalization and scale compensation (such as normalization by the number of window tokens or effective area) can be performed on the heatmaps at each scale first, and then weighted summation can be performed across multiple scales using fixed or learnable weights. The heatmaps can then be aligned to the original resolution using distance weighting or feathering of the halo overlap area to obtain the entire frame's Prompt-Attention semantic heatmap (i.e., the attention heatmap). Optionally, when there are multiple types of Prompts or prototypes, multi-channel class-specific heatmaps can be obtained first, and then the attention heatmap of the target channel can be obtained through maxing, weighted summation, or gating to obtain the final attention heatmap, and so on.

[0086] For example, the initial attention heatmap of any video frame to be processed can be used as the target attention heatmap of any video frame to be processed; or, the initial attention heatmap of any video frame to be processed can be preprocessed to obtain the target attention heatmap of any video frame to be processed, etc.; the embodiments of the present invention do not limit this. Optionally, the heatmap preprocessing process may include, but is not limited to, at least one of the following: denoising, connected component processing, morphological smoothing, and EMA (Exponential Moving Average) temporal smoothing (which can be used to perform exponential moving average on heatmaps / masks to stabilize the output), etc., the embodiments of the present invention do not limit this.

[0087] Optionally, when a superwindow feature is not a window heatmap (i.e., a location point is a feature), the superwindow features of any video frame to be processed under various superwindow sizes can be stitched and fused (i.e., the superwindow features of any video frame to be processed under any superwindow size can be stitched and fused) to obtain the intermediate image features of any video frame to be processed under various superwindow sizes. Then, feature extraction (such as convolution, pooling, etc.) is performed on the intermediate image features of any video frame to be processed under various superwindow sizes to obtain the attention heatmap of any video frame to be processed under various superwindow sizes. Then, the attention heatmaps of any video frame to be processed under various superwindow sizes are fused to obtain the initial attention heatmap of any video frame to be processed. Based on the initial attention heatmap of any video frame to be processed, the target attention heatmap of any video frame to be processed is determined, and so on.

[0088] Based on this, the heatmap multimodal model may include a fusion module. This fusion module can determine the target attention heatmap of any video frame to be processed based on the features of each superwindow at various superwindow sizes. For example, the features of each video frame to be processed at various superwindow sizes can be input into the fusion module to output the target attention heatmap of any video frame to be processed. In this embodiment of the invention, the specific structure of the fusion module is not limited, that is, the specific implementation of the fusion module is not limited.

[0089] Optionally, in other embodiments, any video frame to be processed and task semantic information can be directly input into the heatmap multimodal model to output a target attention heatmap of any video frame to be processed. In this case, the heatmap multimodal model may not include a fusion module, etc., so that no window segmentation is performed on any video frame to be processed, but feature extraction is performed directly on any video frame to be processed and task semantic information, and then the target attention heatmap of any video frame to be processed is output, etc. The present invention does not limit the generation method of the target attention heatmap.

[0090] S204, based on the target attention heatmap of each video frame to be processed, determine the set of attention regions for each video frame to be processed.

[0091] S205: Based on the set of interest regions of each video frame to be processed, determine the set of key frames from the video data to be processed; and based on the interest region indication data of each key frame in the set of key frames, determine the data packet to be processed in the cloud.

[0092] In this embodiment of the invention, the region of interest indication data of a key frame can be added to the same cloud-based data packet to be processed, thereby achieving batch processing of the same key frame and reducing scheduling overhead.

[0093] S206, upload the data packet to be processed in the cloud to the cloud so that target perception can be performed on the cloud side based on the data packet to be processed in the cloud, and obtain the target interest area perception result.

[0094] Among them, the process executed on the cloud side can be the process executed through the cloud, that is, the process executed in the cloud.

[0095] Optionally, the edge-cloud collaborative target perception system can determine the regional image data of each key frame based on the data packets to be processed in the cloud. The regional image data of a key frame may include the image pixels of the region of interest of each region in the set of regions of interest of the corresponding key frame, and the image size of the regional image data of a key frame is the size of the target video frame. For any regional image data in all the determined regional image data (which may include the regional image data of each key frame), the first perception model can be called to determine the instance-level region of interest perception result of any regional image data. The instance-level region of interest perception result may include the instance-level region of interest image indication data of any regional image data and the region type of each instance-level region of interest in any regional image data (also known as region type indication information, such as region type name, region type number or region type representation vector, etc.). The instance-level region of interest image indication data can be used to indicate each instance-level region of interest. Then, based on the region type of each instance-level region of interest, the category prototype of each instance-level region of interest can be determined. The second perception model can then be invoked to infer the meaning of each instance-level region of interest based on its category prototype, obtaining the instance-level inference result for each region of interest. This instance-level inference result is then added to the target region of interest perception result to achieve target perception and obtain the target region of interest perception result. Based on this, embodiments of the present invention can implement a two-stage target perception method based on Prompt semantic transfer on the cloud side, such as a two-stage traffic marking damage detection method based on Prompt semantic transfer. Optionally, the first and second perception models can be set based on experience or actual needs, or they can be obtained through model training; embodiments of the present invention do not limit this. Optionally, the first and second perception models can be the same or different; embodiments of the present invention do not limit this. Optionally, the first and second perception models may or may not include a shared module; embodiments of the present invention do not limit this. Optionally, a perception model can be a deep learning model, such as any convolutional neural network, a multimodal large model, etc.; the embodiments of the present invention do not limit this.

[0096] Optionally, when determining the regional image data of each keyframe based on the cloud-based data packet to be processed, for any keyframe in the keyframe set (i.e., for any keyframe in the cloud-based data packet to be processed), the region of interest indication data of any keyframe can be determined from the cloud-based data packet to be processed. Then, according to the region metadata of each region of interest in the region of interest set of any keyframe, the region of interest image pixels of each region of interest in the region of interest set of any keyframe are superimposed on the target video frame size. Thus, according to the target video frame size and the region metadata of each region of interest in the region of interest set of any keyframe, the regional image data of any keyframe can be determined, so that the position of the region of interest image pixels of any region of interest in any keyframe is the position indicated by the region metadata of any region of interest (e.g., the position indicated by the region location information of any region of interest). Based on this, the edge-side extraction and cloud-side extraction use the same region of interest (ROI) at the same location within the same frame (e.g., a homologous ROI mask, which corresponds one-to-one with the pixel coordinates of the original frame and can be mapped back to the original image from the region metadata). This homologous design ensures that the cloud understands the contextual correspondence of the content sent by the edge, and the cloud can overlay the ROI back onto the global background when needed. Furthermore, since the edge and cloud share the same ROI reference (i.e., homologous), the cloud can perform multi-source data fusion or comparison (e.g., comparing historical ROIs at the same location), thereby enhancing the reliability of the analysis. The target video frame size can be the frame size in the video data to be processed, which is also the frame size of a keyframe, or the image size. Based on this, the region of interest image pixels of each region of interest in the set of regions of interest of any keyframe can be backfilled onto the canvas of the target video frame size (i.e., the same size as the original video frame), thus determining the region image data of any keyframe. Correspondingly, the non-interested region of any keyframe can also be used as the non-interested region in the region image data of any keyframe, and the non-interested region in a region image data can be zeroed out, blacked out, or skipped in calculations.

[0097] Optionally, the goal of the first perception model can be high recall. Optionally, the instance-level interest region image indication data of any region image data may include the instance-level interest region bounding box of each instance-level interest region in the instance-level interest region of the ... Optionally, when calling the first perception model to determine the instance-level region of interest perception result of any region image data, the image data of any region can be input into the first perception model so that the instance-level region of interest perception result of the image data of any region can be output through the first perception model; alternatively, the region of interest mask of any region image data (which can be used to indicate the region of interest in any region image data) can also be input into the first perception model so that convolution operations at non-region of interest locations can be skipped in the intermediate feature layer through the region of interest mask of any region image data, thereby enhancing the model's ability to focus on the region of interest.

[0098] For example, taking traffic marking damage detection as an example of target monitoring task, the first perception model can be a marking detection model (also known as a general marking detection model, which can be used to detect all marking areas as instance-level regions of interest). Furthermore, the system can provide a category prototype for each type of marking (i.e., each area type) under the traffic marking damage detection task (this can be a normal marking prototype; for example, the category prototype corresponding to an area type can be an ideal shape, appearance cues, feature representation vectors, or geometric templates, such as zebra crossings, guide lines, straight lines, turning arrows, etc., each type of marking can correspond to a category prototype). The determination of the area image data for each keyframe and the determination of the category prototype for each instance-level region of interest can be considered the first perception stage; correspondingly, the invocation of the second perception model to determine the target region of interest perception result can be considered the second perception stage, and so on. Optionally, the marking detection model can be a one-stage detector or segmenter trained with markings (intact + damaged) as a single class or a small number of subclasses (zebra crossings, stop lines, arrows, etc.), ensuring recall and localization accuracy.

[0099] Optionally, when calling the second perception model to infer each instance-level region of interest based on the category prototype of each instance-level region of interest, and obtaining the instance-level inference result of each instance-level region of interest, the instance-level region of interest image of any region image data can be determined based on the instance-level region of interest image indicator data of any region image data (e.g., it can be determined from any region image data). The instance-level region of interest image of any region image data may include the image pixels of each instance-level region of interest in any region image data. The instance-level region of interest image of any region image data and the category prototype of each instance-level region of interest in any region image data can be input into the second perception model to output the instance-level inference result of each instance-level region of interest through the second perception model. Alternatively, the instance-level region image and the category prototype of each instance-level region of interest in any region image data can be input into the second perception model, and the instance-level region of interest mask of any region image data can be used as guiding information input into the model to skip the convolution operation of non-instance-level region of interest positions in the intermediate feature layer through the instance-level region of interest mask of any region image data, etc. The embodiments of the present invention do not limit this. Based on this, the second perception stage can use the instance-level attention area and the category prototype of each instance-level attention area in the first perception stage as prompts, and perform reasoning only within the instance-level attention area, such as determining whether a marking is complete or damaged only within the instance-level attention area, etc. Optionally, the instance-level reasoning result of an instance-level attention area may include, but is not limited to, at least one of the following: the reasoning perception category of the corresponding instance-level attention area (such as complete or damaged) and at least one monitoring quantification indicator, etc., which are not limited in this embodiment of the present invention; Optionally, a monitoring task may correspond to at least one monitoring quantification indicator, and the at least one monitoring quantification indicator corresponding to a monitoring task may be set according to experience or actual needs, which are not limited in this embodiment of the present invention. For example, at least one monitoring quantification indicator may include, but is not limited to, at least one of the following: a residual area mask (which can be used to indicate the area where the remaining monitoring object (such as a marking) is located in the instance-level attention area), a damage mask (which can be used to indicate the area where the abnormal target (such as a damaged marking) is located in the instance-level attention area), a damage area, a damage length, etc.

[0100] Optionally, in the traffic marking damage detection task, the second perception model can be a fine-grained detection model for traffic marking damage; optionally, the fine-grained detection model for traffic marking damage may include a cross-attention module (such as a lightweight cross-attention module) to map the category prototypes determined in the first stage into key vectors and value vectors, and to map the instance-level region of interest image into a query vector, thereby achieving cross-attention.

[0101] For example, the instance-level region of interest mask of any region image data can also be used as an explicit input branch or attention guidance information to the second perceptual model, thereby enhancing the model's ability to focus on instance-level regions of interest. For example, at the input layer of the second perceptual model, pixels in non-instance-level regions of interest can be masked or filled with default values ​​(e.g., all pixels in any region image data except for instance-level regions of interest can be 0), and at the intermediate feature layer, the instance-level region of interest mask of any region image data can be used to skip the convolution operation at non-instance-level regions of interest locations, thereby reducing the computational overhead of irrelevant regions, and so on. Optionally, the cloud can also adaptively adjust the model inference process based on the size and complexity of the region of interest and / or instance-level regions of interest: for example, when the region of interest is small (e.g., less than the region size threshold), a smaller sub-model or a reduction in network layers can be automatically selected for inference; when the region of interest is large and the scene is complex (e.g., greater than or equal to the region size threshold, one task scene can correspond to one scene complexity, etc.), the full model is enabled for computation, and so on. Through this mechanism of dynamically switching models based on the complexity of the region of interest, excessive computation can be further avoided, and inference efficiency can be improved. Furthermore, through the above-mentioned sparse computation strategy (e.g., setting non-interested regions in the input image to zero and skipping them, or performing convolution computation only within the bounding box of the region of interest, thereby reducing useless computation), the cloud can only parse the region of interest, effectively reducing computational latency and resource consumption, while ensuring detection accuracy. Optionally, the region size threshold can be set according to experience or actual needs, and this embodiment of the invention does not limit this. As can be seen, the embodiments of the present invention can achieve sparse convolution (also known as masked convolution, where the convolution kernel slides only in the region of interest, and all features in non-interested regions are set to zero and do not participate in the calculation), reduce invalid FLOPs (Floating Point Operations per Second), and suppress boundary artifacts by channel-based and block-based gating; furthermore, it can achieve sparse attention, that is, in the Transformer-based structure, the attention calculation range can be limited by the region of interest mask, retaining only the query-key interaction within the region of interest or the region of interest-neighborhood, supporting a combination of local block sparsity and cross-block sparsity to balance details and necessary context. For example, the image of the input model can first black out, zero out, or fill the non-interested regions and / or non-instance-level interest regions with preset non-interest indicator values ​​(which can be set according to experience or actual needs, i.e., fill with default values), and then skip the calculation of the non-interested regions and / or non-instance-level interest regions based on the corresponding mask during the model inference stage.

[0102] Optionally, when performing inference using a perceptual model (such as a first perceptual model and / or a second perceptual model), the weights of non-interested regions and / or non-instance-level interest regions can be reduced. For example, the computational weights and / or feature fusion weights of spatial pixels in non-interested regions such as zero-based regions and / or non-instance-level interest regions can be reduced, weakening their influence in convolution aggregation, context interaction, and model parameter iteration, retaining only the dominant feature calculations of the effective foreground region. It should be noted that the specific implementation of the weight reduction processing in this embodiment of the invention is not limited, and can be set according to experience or actual needs. Optionally, the inference method that skips non-interested regions and / or non-instance-level interest regions in the calculation, and the inference method that uses weight reduction processing, can be used separately or in combination, and this embodiment of the invention does not limit this.

[0103] Based on this, the embodiments of the present invention implement a two-stage detection scheme for semantic transfer of prompts (also known as a two-stage perception scheme). Further, taking traffic marking damage detection as an example, in the first perception stage (also referred to as the first stage): the system detects and segments all traffic markings, providing them as instance-level regions of interest to the second stage (i.e., the second perception stage); the second stage focuses on the interior of these instance-level regions of interest, detecting the damage details of the markings (i.e., inferring the perception category, such as paint leakage, mottled patches, cracks, etc.). To address this, semantic transfer between the two stages can be achieved through a Prompt mechanism: the prompt information received by the second-stage model comes from the image indication data of the first stage (category prototypes and instance-level regions of interest, such as semantics like the shape / position of road markings). This essentially tells the second stage "what a normal road marking should look like," allowing the model to find deviations from the normal and identify damage. This approach of using prior semantics for anomaly detection improves the sensitivity to detect minor defects and avoids the problem of overlooking details or being interfered with by noise in simple end-to-end detection. By using instance-level category prototypes as prompts, it can achieve greater sensitivity to deviation detection, such as low contrast, small cracks, and edge gaps, thus balancing generalization and interpretability and effectively improving perception accuracy. Furthermore, through two-stage collaboration, this embodiment of the invention can achieve high-precision detection of the location and extent of road marking damage, improving the intelligent level of traffic infrastructure operation and maintenance. It is worth mentioning that there were previously no mature automated solutions for traffic marking damage detection; this solution fills this gap, demonstrating its engineering feasibility and industry value.

[0104] Optionally, the edge-cloud collaborative target perception system can also acquire a set of training image samples on the cloud side, and determine the background region in each training image sample based on the training interest region in each training image sample in the training image sample set. Optionally, the cloud can acquire the set of training image samples from its own storage space, or download the set of training image samples using a training image sample set download link, etc.; this embodiment of the invention does not limit this. Optionally, the interest region of a training image sample can be determined by the interest region mask of the corresponding training image sample, or by the region coordinates of the interest region in the corresponding training image sample, etc.; this embodiment of the invention does not limit this. Optionally, the number of interest regions in a training image sample can be one or more; this embodiment of the invention does not limit this.

[0105] Based on this, the cloud can randomly mask the background regions in each training image sample to obtain background masking indication information for each training image sample. The background masking indication information for a training image sample can be either the background masked image sample (i.e., the image after random background masking) or a background masking region of interest mask (i.e., the mask of the background masked image sample, such as 0 for the masked area and the original pixels for the unmasked area). In this regard, embodiments of the present invention can improve the robustness of the model under partial input by randomly masking non-interested areas of the training image samples to simulate the situation where only the interest area is visible. It should be understood that embodiments of the present invention, by randomly occluding to simulate the application process, can make the training more thorough and accurate. The background masking region of interest mask can be used as an additional input branch of the model or introduced into the attention mechanism to guide the model to focus on the interest area during inference.

[0106] Furthermore, the perception model to be trained can be invoked to determine the training perception result of each training image sample based on the background occlusion indication information of each training image sample. The perception model to be trained may include the initial perception model corresponding to the first perception model and / or the initial perception model corresponding to the second perception model. Optionally, the initial perception model corresponding to the first perception model and the initial perception model corresponding to the second perception model can both be set according to experience or actual needs, or they can be randomly initialized. This embodiment of the invention does not limit this.

[0107] In one implementation, when the background occlusion indication information is a background occlusion image sample, the background occlusion indication information of each training image sample can be input into the perception model to be trained to determine the training perception result of each training image sample. Optionally, when the perception model to be trained is the initial perception model corresponding to the first perception model, the background occlusion indication information of any training image sample can be input into the perception model to be trained, so that the instance-level region of interest perception result of any training image sample can be output through the perception model to be trained. In this case, the training perception result of a training image sample may include the instance-level region of interest perception result of the corresponding training image sample. When the perception model to be trained includes the initial perception model corresponding to the first perception model and the initial perception model corresponding to the second perception model, the background occlusion indication information of any training image sample can be input into the perception model to be trained. Background occlusion indication information of a training image sample is input into the initial perception model corresponding to the first perception model. The initial perception model outputs the instance-level region of interest perception result for any training image sample. This allows the initial perception model corresponding to the second perception model to be invoked. Based on the category prototype of each instance-level region of interest in any training image sample, inference is performed on each instance-level region of interest in any training image sample to obtain the instance-level inference result for each instance-level region of interest in any training image sample (e.g., the instance-level region of interest image and its corresponding category prototype can be input into the initial perception model corresponding to the second perception model). This yields the training perception result for each training image sample, and so on. This embodiment of the invention does not limit this. Therefore, the training perception result of a training image sample may include, but is not limited to, at least one of the following: the instance-level region of interest perception result of the corresponding training image sample, the instance-level inference result of each instance-level region of interest in the corresponding training image sample, etc. This embodiment of the invention does not limit this.

[0108] In another implementation, when the background occlusion indication information is a background occlusion region of interest mask, for any training image sample in the training image sample set, the input of the initial perception model corresponding to the first perception model can be any training image sample, and the background occlusion indication information can be input so that the model automatically skips the feature map calculation of non-intention regions during inference; the input of the initial perception model corresponding to the second perception model can be any training image sample and the corresponding category prototype (at this time, the instance-level region of interest mask of any training image sample can also be input to guide the model to skip the calculation of non-instance-level regions of interest), or it can be the instance-level region of interest image of any training image sample and the corresponding category prototype, etc.; the embodiments of the present invention do not limit this.

[0109] Optionally, when the perception model to be trained is the initial perception model corresponding to the second perception model, the first perception model can be called. Based on the background occlusion indication information of any training image sample, the instance-level attention region perception result of any training image sample can be determined. Then, the initial perception model corresponding to the second perception model can be called. Based on the category prototype of each instance-level attention region in any training image sample, inference can be performed on each instance-level attention region in any training image sample to obtain the instance-level inference result of each instance-level attention region in any training image sample. In this case, only the initial perception model corresponding to the second perception model can be trained to obtain the second perception model. In this case, the first perception model can be set according to experience or actual needs, etc.

[0110] Based on this, embodiments of the present invention can add branch input or gating mechanisms for mask information to the model architecture. This allows for enhanced training through region-of-interest masks, enabling the model to automatically skip feature map calculations for non-interested regions during inference. It also guides the model to focus on the regions of interest indicated by the mask during feature extraction. Furthermore, by randomly masking non-interested regions of training image samples, the model learns to pay more attention to local key information while ignoring background interference. Consequently, the model trained in this way can perform on-demand computation during the inference phase, calculating features only in regions of interest and maintaining sparseness otherwise, thus reducing wasted computational power. Simultaneously, since the model has adapted to incomplete input, it avoids accuracy degradation and ensures reliable analysis results. In other words, embodiments of the present invention, through sparse inference, both guarantee detection accuracy and reduce computational costs.

[0111] Accordingly, the model loss value of the trained perception model can be determined based on the training perception results and the annotation results of each training image sample; alternatively, the loss value under each model loss index can be determined based on the training perception results and the annotation results of each training image sample, and the loss values ​​under each model loss index can be weighted and summed to obtain the model loss value. Optionally, at least one model loss metric may include, but is not limited to, at least one of the following: detection loss (which can guarantee recall, such as being calculated by the difference between the region bounding box of the instance-level region of interest indicated by the training perception results and the annotation box of the instance-level region of interest indicated by the training image sample annotation results, or by the difference between the instance-level region of interest mask and the corresponding annotation mask, etc.), prototype alignment loss (which can make complete samples close to the corresponding category prototype and / or damaged samples far away from the corresponding category prototype, such as being calculated by the difference between the region type representation vector of the complete instance-level region of interest indicated by the training perception results and the representation vector of the corresponding category prototype, etc.), damage classification loss (such as being calculated by using the damage classification probability indicated by the training perception results and the corresponding damage classification label indicated by the training image sample annotation results), damage segmentation loss (such as being calculated by using the segmentation result indicated by the training perception results (such as whether a pixel is damaged or complete, etc.) and the corresponding segmentation label indicated by the training image sample annotation results, etc.), and consistency loss (which can be used to constrain the prediction of only the region of interest input to be consistent with the prediction of the whole image input, and alleviate the drift caused by missing information), etc.; the embodiments of the present invention do not limit this. It should be noted that the embodiments of the present invention do not limit the specific content of the training perception results and the training image sample annotation results.

[0112] Furthermore, the model parameters in the training perception model can be optimized in the direction of reducing the model loss value to obtain an optimized training perception model. Based on the optimized training perception model, a target perception model is determined, which may include a first perception model and / or a second perception model. Specifically, when the training perception model includes the initial perception model corresponding to the first perception model, the target perception model may include the first perception model; when the training perception model includes the initial perception model corresponding to the second perception model, the target perception model may include the second perception model. Optionally, when determining the target perception model based on the optimized training perception model, the optimized training perception model can be further optimized using a set of training image samples until the model convergence condition is met (e.g., the number of iterations reaches a preset iteration threshold, the model loss value is less than a preset model loss threshold, etc.), thereby using the training perception model that meets the convergence condition as the target perception model. Optionally, both the preset iteration threshold and the preset model loss threshold can be set according to experience or actual needs, and this embodiment of the invention does not limit this.

[0113] Based on this, the embodiments of the present invention can achieve the effects of local injection within a stage and weak coupling between stages: for example, it can be indicated that injection is limited to the instance-level area of ​​interest to avoid noise in the whole image; the parameters in the first stage and the second stage can be optimized independently, which is beneficial to porting to other fine-grained anomaly scenarios.

[0114] In summary, the embodiments of the present invention can be organically combined in the overall system to form a closed-loop innovation system of "collection-filtering-transmission-analysis-feedback" under the end-cloud collaboration, which has significant advantages in both algorithm principle and engineering implementation.

[0115] This invention allows for the acquisition of video data to be processed under a target monitoring task at the edge device side, along with the task semantic information of the target monitoring task. For any video frame to be processed within the video data, at least one hyperwindow size can be determined. Then, each video frame is divided into hyperwindows according to each of the at least one hyperwindow size, resulting in a hyperwindow region set for each video frame under each hyperwindow size. This allows the determination of the image block size corresponding to each hyperwindow size. Based on the task semantic information, the image block size corresponding to each hyperwindow size, and the hyperwindow region set for each video frame under each hyperwindow size, a target attention heatmap for each video frame can be determined. Furthermore, based on the target attention heatmaps of each video frame, a set of attention regions for each video frame can be determined; and based on the set of attention regions for each video frame, a set of keyframes can be determined from the video data to be processed; and based on the attention region indication data of each keyframe in the keyframe set, a cloud-based data packet to be processed can be determined. Correspondingly, the data packets to be processed in the cloud can be uploaded to the cloud for target perception on the cloud side, obtaining the target attention region perception result. It can be seen that the embodiments of the present invention can inject features through task semantic information to remove large areas of irrelevant background from the source, and can achieve multi-scale tiling (i.e., non-sliding whole-grid sampling) and dual-path unified representation (i.e., unified to a fixed grid) through multi-scale superwindows. This can balance wide-area retrieval and fine-grained localization, thereby further improving the accuracy of the target attention heatmap, improving the accuracy of the attention region set and keyframe set, and thus effectively improving the accuracy of the target attention region perception result while ensuring efficiency, thereby improving output accuracy.

[0116] Based on the description of the relevant embodiments of the edge-cloud collaborative target perception method above, this invention also proposes an edge-cloud collaborative target perception system; such as Figure 3 As shown, the edge-cloud collaborative target perception system includes an edge device 301 and a cloud device 302. This edge-cloud collaborative target perception system can perform... Figure 1 or Figure 2The edge-cloud collaborative target perception method shown, i.e., the edge-cloud collaborative target perception system, can run the above-mentioned units: The end-side device 301 is used to acquire video data to be processed under the target monitoring task, and to acquire the task semantic information of the target monitoring task; The edge device 301 is further configured to generate target attention heatmaps for each video frame to be processed in the video data to be processed based on the task semantic information; and to determine the set of attention regions for each video frame to be processed based on the target attention heatmaps for each video frame to be processed. The edge device 301 is further configured to determine a set of keyframes from the video data to be processed based on the set of interest regions of each video frame to be processed; and to determine a data packet to be processed in the cloud based on the interest region indication data of each keyframe in the set of keyframes; wherein, the interest region indication data of a keyframe includes the interest region indication data of each interest region in the interest region set of the corresponding keyframe, the interest region indication data of an interest region includes the interest region image pixels and region metadata of the corresponding interest region, and the region metadata of an interest region includes the region location information of the corresponding interest region; The end-side device 301 is also used to upload the data packet to be processed in the cloud to the cloud; The cloud 302 is used to perform target perception based on the cloud-based data packet to be processed, and obtain the target area of ​​interest perception result.

[0117] In one embodiment, when the end-side device 301 determines the set of attention regions for each video frame to be processed based on the target attention heatmap of each video frame to be processed, it can be specifically used for: The process iterates through each video frame in the video data to be processed, and takes the currently traversed video frame as the current video frame, and determines the region of interest extraction index; wherein, the region of interest extraction index can be updated according to the feedback control instructions issued by the cloud. Based on the aforementioned region of interest extraction index, the set of regions of interest for the current video frame is determined from the target attention heatmap of the current video frame.

[0118] In another embodiment, when the edge device 301 generates target attention heatmaps for each video frame to be processed in the video data to be processed based on the task semantic information, it can be specifically used for: For any video frame to be processed in the video data to be processed, at least one superwindow size is determined, and the video frame to be processed is divided into superwindows according to each of the at least one superwindow size, so as to obtain the set of superwindow regions of the video frame to be processed under each superwindow size. The image patch size corresponding to each hyperwindow size is determined, and based on the task semantic information, the image patch size corresponding to each hyperwindow size, and the hyperwindow region set of any video frame to be processed under each hyperwindow size, the target attention heatmap of any video frame to be processed is determined.

[0119] In another embodiment, when the edge device 301 determines the target attention heatmap of any video frame to be processed based on the task semantic information, the image patch size corresponding to each hyperwindow size, and the hyperwindow region set of any video frame to be processed under each hyperwindow size, it can be specifically used for: For any of the at least one superwindow size, and any superwindow region in the set of superwindow regions of any video frame to be processed under any superwindow size, local attention features are extracted for any superwindow region according to the image block size corresponding to any superwindow size and the task semantic information, so as to obtain the superwindow features of any superwindow region. After obtaining the superwindow features of each superwindow region of any video frame to be processed under each superwindow size, the target attention heatmap of any video frame to be processed is determined based on the superwindow features of any video frame to be processed under each superwindow size.

[0120] In another embodiment, the super-window feature of any super-window region is determined by the target attention module corresponding to the super-window region. An attention module includes a cross-attention module. The end-side device 301 can also be used for: Determine the initial attention module corresponding to any of the hyperwindow regions; When a feature enhancement requirement is detected for any of the super-window regions, a feature enhancement module is inserted into the initial attention module corresponding to any super-window region to obtain the target attention module corresponding to any super-window region. If no feature enhancement requirement is detected for any of the super-window regions, the initial attention module corresponding to any super-window region is used as the target attention module corresponding to any super-window region.

[0121] In another embodiment, when the end-side device 301 determines the set of keyframes from the video data to be processed based on the set of regions of interest of each video frame to be processed, it can specifically be used to: For any video frame to be processed in the video data to be processed, a region of interest mask for the video frame to be processed is determined based on the region of interest set of the video frame to be processed. When a previous keyframe of any video frame to be processed exists, the degree of change of the region of interest of any video frame to be processed is determined based on the region of interest mask of any video frame to be processed and the region of interest mask of the previous keyframe. The previous keyframe is the keyframe whose acquisition time is before the acquisition time of any video frame to be processed and is the closest to any video frame to be processed. Based on the degree of change in the region of interest and the keyframe sampling strategy, it is determined whether any video frame to be processed is considered a keyframe; if it is determined that any video frame to be processed is considered a keyframe, then the video frame to be processed is added to the keyframe set, so as to determine the keyframe set from the video data to be processed.

[0122] In another implementation, when the cloud-based 302 performs target perception based on the cloud-based data packet to be processed and obtains the target area of ​​interest perception result, it can be specifically used for: Based on the cloud-based data packets to be processed, the regional image data of each key frame is determined. The regional image data of a key frame includes the image pixels of the region of interest of each region in the set of regions of interest of the corresponding key frame, and the image size of the regional image data of a key frame is the target video frame size. For any region image data among all the identified region image data, the first perception model is invoked to determine the instance-level attention region perception result of the any region image data. The instance-level attention region perception result includes instance-level attention region image indication data of the any region image data and the region type of each instance-level attention region in the any region image data. The instance-level attention region image indication data is used to indicate each instance-level attention region. Based on the region type of each instance-level region of interest, determine the category prototype of each instance-level region of interest; The second perception model is invoked to infer each instance-level region of interest based on the category prototype of each instance-level region of interest, thereby obtaining the instance-level inference result of each instance-level region of interest, and adding the instance-level inference result of each instance-level region of interest to the perception result of the target region of interest.

[0123] In another implementation, the cloud 302 can also be used for: Obtain a set of training image samples, and determine the background region in each training image sample based on the training interest region in each training image sample in the set of training image samples; The background regions in each training image sample are randomly masked to obtain background masking indication information for each training image sample. The background masking indication information for a training image sample is either the background masking image sample or the background masking region of interest mask for the corresponding training image sample. The training perception model is invoked, and the training perception result of each training image sample is determined based on the background occlusion indication information of each training image sample. The training perception model includes the initial perception model corresponding to the first perception model and / or the initial perception model corresponding to the second perception model. Based on the training perception results and training image sample annotation results of each training image sample, the model loss value of the perception model to be trained is determined. The model parameters in the perception model to be trained are optimized in the direction of reducing the model loss value to obtain the optimized perception model to be trained; and based on the optimized perception model to be trained, a target perception model is determined, the target perception model including the first perception model and / or the second perception model.

[0124] According to one embodiment of the present invention, Figure 3 Each device in the edge-cloud collaborative target perception system shown can be individually or entirely merged into one or more other units, or one or more of the devices can be further divided into multiple functionally smaller devices. This achieves the same operation without affecting the technical effects of the embodiments of the present invention. In practical applications, the function of one device can also be implemented by multiple devices, or the function of multiple devices can be implemented by one device. In other embodiments of the present invention, any edge-cloud collaborative target perception system may also include other devices. In practical applications, these functions can also be implemented with the assistance of other devices, and can be implemented collaboratively by multiple devices.

[0125] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.

[0126] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.

[0127] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. An end-to-cloud coordination target perception method, characterized in that, The edge-cloud collaborative target perception method is applied to an edge-cloud collaborative target perception system, which includes edge devices and a cloud. The method includes: The device acquires the video data to be processed under the target monitoring task, and also acquires the task semantic information of the target monitoring task. Based on the task semantic information, target attention heatmaps are generated for each video frame in the video data to be processed, including: for any video frame in the video data to be processed, determining at least one hyperwindow size, and dividing the video frame into hyperwindows according to each of the at least one hyperwindow size to obtain a hyperwindow region set for the video frame under each hyperwindow size; determining the image block size corresponding to each hyperwindow size, and determining the target attention heatmap for the video frame based on the task semantic information, the image block size corresponding to each hyperwindow size, and the hyperwindow region set for the video frame under each hyperwindow size; and determining the attention region set for each video frame based on the target attention heatmaps for each video frame. Based on the set of regions of interest for each video frame to be processed, a set of keyframes is determined from the video data to be processed; and based on the region of interest indication data of each keyframe in the set of keyframes, a data packet to be processed in the cloud is determined; wherein, the region of interest indication data of a keyframe includes the region of interest indication data of each region of interest in the set of regions of interest of the corresponding keyframe, the region of interest indication data of a region of interest includes the region of interest image pixels and region metadata of the corresponding region of interest, and the region metadata of a region of interest includes the region location information of the corresponding region of interest; The data packet to be processed in the cloud is uploaded to the cloud so that target perception is performed on the cloud side based on the data packet to be processed in the cloud, and the target interest area perception result is obtained. The step of determining the attention region set of each video frame to be processed based on the target attention heatmap of each video frame to be processed includes: traversing each video frame to be processed in the video data to be processed, and taking the currently traversed video frame as the current video frame, and determining the attention region extraction index; wherein the attention region extraction index can be updated according to the feedback control command issued by the cloud; and determining the attention region set of the current video frame from the target attention heatmap of the current video frame according to the attention region extraction index. The step of determining the target attention heatmap of any video frame to be processed based on the task semantic information, the image patch size corresponding to each hyperwindow size, and the hyperwindow region set of any video frame to be processed under each hyperwindow size includes: for any hyperwindow size in the at least one hyperwindow size, and any hyperwindow region in the hyperwindow region set of any video frame to be processed under any hyperwindow size, extracting local attention features of any hyperwindow region according to the image patch size corresponding to any hyperwindow size and the task semantic information to obtain the hyperwindow features of any hyperwindow region; after obtaining the hyperwindow features of each hyperwindow region of any video frame to be processed under each hyperwindow size, determining the target attention heatmap of any video frame to be processed based on the hyperwindow features of any video frame to be processed under each hyperwindow size.

2. The method of claim 1, wherein, The superwindow feature of any superwindow region is determined by the target attention module corresponding to the superwindow region. An attention module includes a cross-attention module. The method further includes: Determine the initial attention module corresponding to any of the hyperwindow regions; When a feature enhancement requirement is detected for any of the super-window regions, a feature enhancement module is inserted into the initial attention module corresponding to any super-window region to obtain the target attention module corresponding to any super-window region. If no feature enhancement requirement is detected for any of the super-window regions, the initial attention module corresponding to any super-window region is used as the target attention module corresponding to any super-window region.

3. The method of claim 1, wherein, The process of determining a set of keyframes from the video data to be processed based on the set of regions of interest for each of the video frames to be processed includes: For any video frame to be processed in the video data to be processed, a region of interest mask for the video frame to be processed is determined based on the region of interest set of the video frame to be processed. When a previous keyframe of any video frame to be processed exists, the degree of change of the region of interest of any video frame to be processed is determined based on the region of interest mask of any video frame to be processed and the region of interest mask of the previous keyframe. The previous keyframe is the keyframe whose acquisition time is before the acquisition time of any video frame to be processed and is the closest to any video frame to be processed. Based on the degree of change in the region of interest and the keyframe sampling strategy, it is determined whether any video frame to be processed is considered a keyframe; if it is determined that any video frame to be processed is considered a keyframe, then the video frame to be processed is added to the keyframe set, so as to determine the keyframe set from the video data to be processed.

4. The method according to claim 1, characterized in that, The step of performing target perception based on the cloud-to-be-processed data packets on the cloud side to obtain target interest area perception results includes: On the cloud side, based on the cloud-to-process data packet, the regional image data of each key frame is determined. The regional image data of a key frame includes the region image pixels of each region of interest in the region of interest set of the corresponding key frame, and the image size of the regional image data of a key frame is the target video frame size. For any region image data among all the identified region image data, the first perception model is invoked to determine the instance-level attention region perception result of the any region image data. The instance-level attention region perception result includes instance-level attention region image indication data of the any region image data and the region type of each instance-level attention region in the any region image data. The instance-level attention region image indication data is used to indicate each instance-level attention region. Based on the region type of each instance-level region of interest, determine the category prototype of each instance-level region of interest; The second perception model is invoked to infer each instance-level region of interest based on the category prototype of each instance-level region of interest, thereby obtaining the instance-level inference result of each instance-level region of interest, and adding the instance-level inference result of each instance-level region of interest to the perception result of the target region of interest.

5. The method according to claim 4, characterized in that, The method further includes: A set of training image samples is obtained on the cloud side, and the background region in each training image sample is determined based on the training interest region in each training image sample in the set of training image samples. The background regions in each training image sample are randomly masked to obtain background masking indication information for each training image sample. The background masking indication information for a training image sample is either the background masking image sample or the background masking region of interest mask for the corresponding training image sample. The training perception model is invoked, and the training perception result of each training image sample is determined based on the background occlusion indication information of each training image sample. The training perception model includes the initial perception model corresponding to the first perception model and / or the initial perception model corresponding to the second perception model. Based on the training perception results and training image sample annotation results of each training image sample, the model loss value of the perception model to be trained is determined. The model parameters in the perception model to be trained are optimized in the direction of reducing the model loss value to obtain the optimized perception model to be trained; and based on the optimized perception model to be trained, a target perception model is determined, the target perception model including the first perception model and / or the second perception model.

6. An edge-cloud collaborative target perception system, characterized in that, The edge-cloud collaborative target perception system includes edge devices and a cloud platform; wherein... The end-side device is used to acquire video data to be processed under the target monitoring task, and to acquire the task semantic information of the target monitoring task; The edge device is further configured to generate target attention heatmaps for each video frame to be processed in the video data to be processed based on the task semantic information, including: for any video frame to be processed in the video data to be processed, determining at least one hyperwindow size, and dividing the video frame to be processed into hyperwindows according to each of the at least one hyperwindow size, to obtain a hyperwindow region set for the video frame to be processed under each hyperwindow size; determining the image block size corresponding to each hyperwindow size, and determining the target attention heatmap for the video frame to be processed based on the task semantic information, the image block size corresponding to each hyperwindow size, and the hyperwindow region set for the video frame to be processed under each hyperwindow size; and determining the attention region set for each video frame to be processed based on the target attention heatmaps for each video frame to be processed. The edge device is further configured to determine a set of keyframes from the video data to be processed based on the set of regions of interest of each video frame to be processed; and to determine a data packet to be processed in the cloud based on the region of interest indication data of each keyframe in the set of keyframes; wherein, the region of interest indication data of a keyframe includes the region of interest indication data of each region of interest in the set of regions of interest of the corresponding keyframe, the region of interest indication data of a region of interest includes the region of interest image pixels and region metadata of the corresponding region of interest, and the region metadata of a region of interest includes the region location information of the corresponding region of interest; The end-side device is also used to upload the data packets to be processed in the cloud to the cloud; The cloud is used to perform target perception based on the data packets to be processed in the cloud, and obtain the target area of ​​interest perception result. Specifically, when the edge device determines the set of attention regions for each video frame to be processed based on the target attention heatmap of each video frame to be processed, it is used to: traverse each video frame to be processed in the video data to be processed, and take the currently traversed video frame as the current video frame, and determine the attention region extraction index; wherein, the attention region extraction index can be updated according to the feedback control command issued by the cloud; and determine the set of attention regions for the current video frame from the target attention heatmap of the current video frame according to the attention region extraction index; When the edge device determines the target attention heatmap of any video frame to be processed based on the task semantic information, the image patch size corresponding to each superwindow size, and the superwindow region set of any video frame to be processed under each superwindow size, it specifically performs the following steps: for any superwindow size in the at least one superwindow size, and any superwindow region in the superwindow region set of any video frame to be processed under any superwindow size, it performs local attention feature extraction on any superwindow region according to the image patch size corresponding to any superwindow size and the task semantic information to obtain the superwindow feature of any superwindow region; after obtaining the superwindow features of each superwindow region of any video frame to be processed under each superwindow size, it determines the target attention heatmap of any video frame to be processed based on the superwindow features of any video frame to be processed under each superwindow size.

7. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.