Long video semantic analysis method and device based on cloud edge collaboration, equipment and medium

CN122473717BActive Publication Date: 2026-09-25SHENZHEN RES INST OF BIG DATA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610955637.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

[0004]然而,相关技术这种方式在上传原始长视频或大量的采样视频帧时,由于视频分辨率高、视频时长久、网络信号薄弱等因素影响,这导致上传至云服务器时不仅需要较高的通信带宽开销,且传输时延长,降低了对长视频的识别效率

Benefits of technology

[0017]本申请实施例可应用于长视频语义分析系统,长视频语义分析系统至少包括云服务器和边缘端,首先,云服务器接收终端发送的原始问题,将原始问题拆解为多个子问题,并确定每个子问题对应的重要性评分,将多个子问题以及每个子问题对应的重要性评分下发至边缘端,如此,云服务器将接收到的原始问题拆分得到多个子问题,并标定每个子问题的重要性评分作为把控视频帧选取数量的依据,为云边协同处理的奠定基础;然后,边缘端接收到多个子问题以及每个子问题对应的重要性评分,并接收终端上传的长视频文件,并按照预设帧率对长视频文件进行视频帧抽取,得到候选视频帧集合,如此,通过边缘端抽帧预处理,初步缩减待筛选数据体量,规避原始长视频全量传输带来的带宽损耗;接着,边缘端结合每个子问题对应的重要性评分,确定每个子问题对应的上传帧数,如此,在有限的通信带宽传输资源下优先保障核心子问题的帧数量,从而在源头控制整体上传数据量;进而,边缘端针对每个子问题,结合子问题与候选视频帧集合中每个候选视频帧之间的特征相似度,从候选视频帧集合中选取每个子问题对应的上传帧数的目标视频帧,如此,通过子问题定向筛选关联的视频帧,仅保留各子问题所需有效画面,剔除无关冗余帧,在有限通信带宽下减少上传视频帧的数量,提高后续的识别效率;进一步的,边缘端结合每个子问题对应的上传帧数的目标视频帧生成目标视频帧集合,并将目标视频帧集合上传至云服务器,如此,仅上传少量筛选后的关键帧,大幅降低传输带宽占用,缩短传输时延;最后,云服务器通过预设的多模态推理模型基于目标视频帧集合、原始问题以及多个子问题进行推理得到目标推理结果,并将目标推理结果返回至终端,如此,结合云端算法优势,对精简的目标视频帧集合实现高效推理,在保证识别准确率的同时提升长视频整体识别效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473717B_ABST
    Figure CN122473717B_ABST
Patent Text Reader

Abstract

The application discloses a long video semantic analysis method and device based on cloud edge cooperation, equipment and medium. The cloud server splits the original question sent by the terminal into multiple sub-questions, determines the importance score of each sub-question, and issues it to the edge end. The edge end frames the long video uploaded by the terminal to obtain a candidate video frame set. Based on each sub-question and the corresponding importance score, a specific number of target video frames that meet each sub-question are selected from the candidate video frame set to form a target video frame set containing key frames for multiple sub-questions. The edge end only needs to upload the target video frame set containing only key frames to the cloud server for identification for the multiple sub-questions issued by the cloud server, without uploading the complete long video. In this way, the communication bandwidth is reduced under the condition of limited communication bandwidth, the communication bandwidth overhead is reduced, the transmission efficiency of video resources is improved, and the identification efficiency of the long video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device and medium for long video semantic analysis based on cloud-edge collaboration. Background Technology

[0002] With the development of large language models and multimodal large models, long video understanding has become a key capability in scenarios such as smart cameras, embodied intelligence, smart homes, robot assistants, industrial inspection, and in-vehicle intelligent agents. Long video understanding tasks typically require the system to locate visual evidence related to the user's natural language questions in videos containing a large number of redundant frames and complex temporal relationships, and further perform event recognition, causal judgment, temporal reasoning, action relationship analysis, or video question answering.

[0003] In related technologies, cloud servers receive raw long videos or a large number of sampled video frames uploaded by terminals, as well as natural language questions uploaded by the receiving terminals. Through large models, semantic understanding and complex reasoning are performed on the received raw long videos or a large number of sampled video frames to obtain answers to the natural language questions.

[0004] However, when uploading original long videos or a large number of sampled video frames, this method of technology suffers from high communication bandwidth overhead and long transmission time when uploading to the cloud server due to factors such as high video resolution, long video duration, and weak network signal, which reduces the recognition efficiency of long videos. Summary of the Invention

[0005] This application provides a cloud-edge collaborative long video semantic analysis method, apparatus, device, and medium, which can reduce communication bandwidth overhead, improve the transmission efficiency of video resources, and thus improve the recognition efficiency of long videos.

[0006] Firstly, this application provides a cloud-edge collaborative long video semantic analysis method, applied to a long video semantic analysis system, wherein the long video semantic analysis system includes at least a cloud server and an edge terminal, comprising: The cloud server receives the original question sent by the terminal, breaks down the original question into multiple sub-questions, determines the importance score corresponding to each sub-question, and sends the multiple sub-questions and the importance score corresponding to each sub-question to the edge terminal. The edge device receives the multiple sub-questions and the importance score corresponding to each sub-question, and receives the long video file uploaded by the terminal, and extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. The edge terminal combines the importance score corresponding to each sub-problem to determine the number of upload frames corresponding to each sub-problem; For each sub-problem, the edge terminal combines the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set to select the target video frame corresponding to the number of uploaded frames for each sub-problem from the candidate video frame set; The edge device generates a target video frame set by combining the target video frames corresponding to the number of uploaded frames for each sub-problem, and uploads the target video frame set to the cloud server; The cloud server uses a preset multimodal reasoning model to reason based on the target video frame set, the original question, and the multiple sub-questions to obtain the target reasoning result, and then returns the target reasoning result to the terminal.

[0007] Secondly, this application provides a cloud-edge collaborative long video semantic analysis device, applied at the edge of a long video semantic analysis system, wherein the long video semantic analysis system further includes a cloud server, and the device includes: The receiving unit is used to receive multiple sub-problems and the importance score corresponding to each sub-problem, as well as a long video file uploaded by the receiving terminal, and to extract video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. The multiple sub-problems are obtained by the cloud server from the original problem sent by the terminal, and each importance score is generated by the cloud server for the corresponding sub-problem. The determination unit is used to determine the number of upload frames for each sub-problem by combining the importance score corresponding to each sub-problem; The selection unit is used to select the target video frame corresponding to the number of uploaded frames for each sub-problem from the candidate video frame set, based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set. The upload unit is used to generate a target video frame set by combining the target video frames corresponding to the upload frame number of each sub-problem, and upload the target video frame set to the cloud server. The cloud server then uses a preset multimodal inference model to infer the target inference result based on the target video frame set, the original problem, and the multiple sub-problems, and returns the target inference result to the terminal.

[0008] In some embodiments, the determining unit is further configured to: Get the preset total number of uploaded frames and the preset minimum number of frames for each sub-question; Determine the number of sub-problems corresponding to multiple sub-problems; By combining the preset total number of upload frames, the preset minimum number of frames, the number of questions, and the importance score corresponding to each sub-question, the number of upload frames corresponding to each sub-question is determined.

[0009] In some embodiments, the determining unit is further configured to: The initial total number of frames occupied for the multiple sub-questions is determined by combining the number of questions and the preset minimum number of frames, and the remaining number of allocated frames is determined based on the difference between the preset total number of uploaded frames and the initial total number of frames occupied. When the remaining number of allocated frames is greater than or equal to the preset frame number threshold, the allocation weight corresponding to each sub-problem is determined by combining the importance score corresponding to each sub-problem. The remaining number of allocated frames is then allocated to the multiple sub-problems to obtain the redistributed frame number for each sub-problem, and the uploaded frame number corresponding to each sub-problem is obtained by combining the redistributed frame number for each sub-problem and the preset minimum frame number. When the remaining number of allocated frames is less than the preset frame number threshold, the score ranking relationship between the multiple sub-questions is determined by combining the importance score corresponding to each sub-question, and at least one sub-question ranked first in the score ranking relationship is selected from the preset minimum frame number for frame allocation to obtain the number of upload frames corresponding to each sub-question.

[0010] In some embodiments, the determining unit is further configured to: Based on the allocation weight corresponding to each sub-problem and the remaining allocation frames, the initial redistribution value of each sub-problem is determined, and the initial redistribution value of each sub-problem is rounded down to obtain the candidate redistribution frame number of each sub-problem. Calculate the total number of candidate reassignment frames for the multiple sub-problems by combining the number of candidate reassignment frames for each sub-problem, and determine the remaining amount of reassignment between the remaining number of reassignment frames and the total number of candidate reassignment frames; By combining the importance score corresponding to each sub-problem, the score ranking relationship among the multiple sub-problems is determined, and the remaining redistribution amount is sequentially allocated to at least one sub-problem ranked first in the score ranking relationship to obtain the number of secondary redistribution frames for each sub-problem. For each sub-problem, the number of reassignment frames for each sub-problem is determined by combining the corresponding candidate reassignment frame number and the secondary reassignment frame number.

[0011] In some embodiments, the selection unit is further configured to: Each candidate video frame in the candidate video frame set is encoded to obtain the corresponding frame feature matrix, and each sub-problem is encoded to obtain the corresponding sub-problem feature matrix. For each sub-problem, the feature matrix of the corresponding sub-problem is compared with the feature matrix of each frame to obtain the feature similarity between the current sub-problem and each candidate video frame; For each sub-problem, candidate video frames with the corresponding number of uploaded frames are selected from the candidate video frame set in descending order of feature similarity to obtain the target video frame for each sub-problem.

[0012] In some embodiments, the selection unit is further configured to: For each sub-problem, candidate video frames with the corresponding number of uploaded frames are selected from the candidate video frame set in descending order of the feature similarity to obtain the initial video frame set for each sub-problem. By combining the initial video frame set of each sub-problem, deduplication is performed to obtain the candidate video frame set corresponding to each sub-problem; Determine the number of video frames in the candidate video frame set corresponding to each sub-problem, and determine the number of supplementary frames for each sub-problem based on the number of video frames corresponding to each sub-problem and the corresponding number of uploaded frames; Based on the number of supplementary frames for each sub-problem, supplementary video frames corresponding to each sub-problem are selected from the candidate video frame set; The supplementary video frames corresponding to each sub-problem are added to the corresponding set of candidate video frames to obtain the target video frames corresponding to each sub-problem.

[0013] Thirdly, this application provides a cloud-edge collaborative long video semantic analysis device, applied to a cloud server in a long video semantic analysis system, wherein the long video semantic analysis system further includes an edge terminal, and the device includes: The receiving unit is used to receive the original question sent by the terminal, decompose the original question into multiple sub-questions, determine the importance score corresponding to each sub-question, and send the multiple sub-questions and the importance score corresponding to each sub-question to the edge terminal. The edge device receives the multiple sub-questions and the importance score corresponding to each sub-question, and receives the long video file uploaded by the terminal. It extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. Combining the importance score corresponding to each sub-question, it determines the number of upload frames corresponding to each sub-question. For each sub-question, it selects the target video frames corresponding to the number of upload frames corresponding to each sub-question from the candidate video frame set based on the feature similarity between the sub-question and each candidate video frame in the candidate video frame set. It generates a target video frame set by combining the target video frames corresponding to the number of upload frames corresponding to each sub-question, and uploads the target video frame set to the cloud server. The inference unit is used to perform inference based on the target video frame set, the original question, and the multiple sub-questions using a preset multimodal inference model to obtain a target inference result, and then return the target inference result to the terminal.

[0014] In some embodiments, the receiving unit is further configured to: Determine the logical order of events among the multiple sub-problems, and construct a directed acyclic graph (DAG) for the multiple sub-problems according to the logical order of events, wherein the DAG represents the temporal relationship between the multiple sub-problems; Based on the directed acyclic graph, an importance score for each subproblem relative to the original problem is generated.

[0015] Furthermore, this application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the above-mentioned long video semantic analysis method based on cloud-edge collaboration.

[0016] Furthermore, embodiments of this application also provide a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute the aforementioned cloud-edge collaborative long video semantic analysis method.

[0017] This application embodiment can be applied to a long video semantic analysis system, which includes at least a cloud server and an edge device. First, the cloud server receives the original question sent by the terminal, breaks it down into multiple sub-questions, and determines the importance score for each sub-question. The cloud server then sends the multiple sub-questions and their corresponding importance scores to the edge device. In this way, the cloud server decomposes the received original question into multiple sub-questions and assigns an importance score to each sub-question as a basis for controlling the number of video frames selected, laying the foundation for cloud-edge collaborative processing. Next, the edge device receives the multiple sub-questions and their corresponding importance scores, and receives the long video file uploaded by the terminal. It then extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. Thus, through frame extraction preprocessing at the edge device, the amount of data to be filtered is initially reduced, avoiding bandwidth loss caused by the full transmission of the original long video. Finally, the edge device, based on the importance score of each sub-question, determines the number of frames to be uploaded for each sub-question. In this way, core sub-questions are prioritized within limited communication bandwidth resources. The number of frames is controlled at the source to manage the overall amount of uploaded data. Then, for each sub-question, the edge device, based on the feature similarity between the sub-question and each candidate video frame in the candidate video frame set, selects the target video frames corresponding to the upload frame count for each sub-question. This targeted filtering of related video frames by sub-question retains only the effective images required for each sub-question, eliminating irrelevant and redundant frames, reducing the number of uploaded video frames under limited communication bandwidth, and improving subsequent recognition efficiency. Furthermore, the edge device generates a target video frame set based on the target video frames corresponding to the upload frame count for each sub-question and uploads the target video frame set to the cloud server. This uploads only a small number of filtered key frames, significantly reducing transmission bandwidth usage and shortening transmission latency. Finally, the cloud server uses a preset multimodal inference model to infer the target video frame set, the original question, and multiple sub-questions to obtain the target inference result, and returns the target inference result to the terminal. Thus, combining the advantages of cloud-based algorithms, efficient inference is achieved on the streamlined target video frame set, improving the overall recognition efficiency of long videos while ensuring recognition accuracy.

[0018] As can be seen from the above, compared with related technologies that, when uploading original long videos or a large number of sampled video frames, suffer from high communication bandwidth overhead and long transmission time due to factors such as high video resolution, long video duration, and weak network signal, thus reducing the recognition efficiency of long videos, the cloud server in this application breaks down the original question sent by the terminal into multiple sub-questions and determines the importance score of each sub-question. By distributing multiple sub-questions and their corresponding importance scores to the edge, the edge extracts frames from the long video uploaded by the terminal to obtain a set of candidate video frames. Based on each sub-question and its corresponding importance score, it selects a specific number of target video frames that meet each sub-question from the set of candidate video frames, thereby forming a set of target video frames containing key frames for multiple sub-questions. The edge only needs to upload the set of target video frames containing only key frames to the cloud server for recognition, without uploading the complete long video. In this way, under limited communication bandwidth conditions, the occupation of communication bandwidth is reduced, the communication bandwidth overhead is lowered, the transmission efficiency of video resources is improved, and thus the recognition efficiency of long videos is enhanced. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A schematic diagram of a long video semantic analysis system based on cloud-edge collaboration provided in an embodiment of this application; Figure 2 A flowchart illustrating the steps of the cloud-edge collaborative long video semantic analysis method provided in this application embodiment; Figure 3 Example diagram of a cloud-edge collaborative long video semantic analysis system architecture provided in the embodiments of this application; Figure 4 A flowchart of long video semantic analysis based on cloud-edge collaboration is provided for embodiments of this application; Figure 5 A schematic diagram of the structure of a cloud-edge collaborative long video semantic analysis device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the terminal structure provided in the embodiments of this application; Figure 7 This is a partial structural block diagram of the server provided in an embodiment of this application. Detailed Implementation

[0021] To enable those skilled in the art to better understand the solutions of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, other embodiments obtained by those skilled in the art without creative effort are all within the scope of protection of this application.

[0022] It is understood that in the specific implementation of this application, data related to the original problem, sub-problems, importance scores, long video files, candidate video frame sets, number of uploaded frames, target video frames, and target inference results are involved. When the above embodiments of this application are applied to specific products or technologies, permission or consent from the target is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0023] Furthermore, when this application embodiment needs to obtain relevant data, it will obtain separate permission or separate consent for the original question, sub-questions, importance scores, long video files, candidate video frame sets, uploaded frame counts, target video frames, target inference results, and other related data through pop-up windows or redirection to a confirmation page. Only after explicitly obtaining separate permission or separate consent for the original question, sub-questions, importance scores, long video files, candidate video frame sets, uploaded frame counts, target video frames, target inference results, and other related data will it obtain the necessary data for enabling the application embodiment to operate normally.

[0024] It should be noted that while some processes described in the specification, claims, and accompanying drawings include multiple steps appearing in a specific order, it should be clearly understood that these steps may not be performed in the order they appear herein, or may be performed in parallel. The step numbers are merely used to distinguish different steps and do not represent any execution order. Furthermore, descriptions such as "first," "second," or "objective" in this document are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, other embodiments obtained by those skilled in the art without creative effort are all within the scope of protection of this application.

[0026] This application provides a method, apparatus, device, and medium for long-video semantic analysis based on cloud-edge collaboration. Specifically, the long-video semantic analysis method based on cloud-edge collaboration in this application can be implemented in a computer device, which can be a server or a terminal device. The server can be an independent physical server, a server cluster or service cluster composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The user terminal device can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart home appliance, vehicle terminal, smart voice interaction device, aircraft, drone, etc., but is not limited to these.

[0027] To facilitate understanding, this application will describe the implementation process of the cloud-edge collaborative long video semantic analysis method through several embodiments, as follows: This application provides a cloud-edge collaborative long video semantic analysis method, applicable to long video semantic analysis systems. The long video semantic analysis system includes at least a cloud server and an edge terminal. First, the cloud server receives the original question sent by the terminal, breaks it down into multiple sub-questions, and determines the importance score for each sub-question. The cloud server then sends the multiple sub-questions and their corresponding importance scores to the edge terminal. Thus, the cloud server decomposes the received original question into multiple sub-questions and assigns an importance score to each sub-question as a basis for controlling the number of video frames selected, laying the foundation for cloud-edge collaborative processing. Next, the edge terminal receives the multiple sub-questions and their corresponding importance scores, and receives the long video file uploaded by the terminal. It then extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. This edge-end frame extraction preprocessing initially reduces the amount of data to be filtered, avoiding bandwidth loss caused by transmitting the entire original long video. Finally, the edge terminal, based on the importance score of each sub-question, determines the number of frames to be uploaded for each sub-question. This allows for efficient transmission of limited communication bandwidth resources. Prioritizing the number of frames for core sub-problems controls the overall amount of uploaded data at the source. Then, at the edge, for each sub-problem, the target video frames corresponding to the uploaded frame count for each sub-problem are selected from the candidate video frame set based on the feature similarity between the sub-problem and each candidate video frame. This targeted filtering of related video frames by sub-problem retains only the effective images required for each sub-problem, eliminating irrelevant and redundant frames, reducing the number of uploaded video frames under limited communication bandwidth, and improving subsequent recognition efficiency. Furthermore, the edge generates a target video frame set based on the target video frames corresponding to the uploaded frame count for each sub-problem and uploads the target video frame set to the cloud server. This uploads only a small number of filtered key frames, significantly reducing transmission bandwidth consumption and shortening transmission latency. Finally, the cloud server uses a preset multimodal inference model to infer the target video frame set, the original problem, and multiple sub-problems to obtain the target inference result, and returns the target inference result to the terminal. Thus, combining the advantages of cloud-based algorithms, efficient inference is achieved on the streamlined target video frame set, improving the overall recognition efficiency of long videos while ensuring recognition accuracy. Please refer to the following specific embodiments for details.

[0028] For example, see Figure 1 This is a schematic diagram of a long video semantic analysis system based on cloud-edge collaboration provided in this application embodiment. The system includes a terminal 110, an edge terminal 120, and a cloud server 130.

[0029] The terminal 110 can have a target application installed on it, and the corresponding application business can be run through the target application. The target application can be called a client. Taking the terminal 110 as a user's mobile phone or computer device as an example, the terminal 110 can send the user's original input question to the cloud server 130 and send the long video file to the edge terminal 120.

[0030] The cloud server 130 can be a single service node, a distributed system composed of multiple service nodes, or a service node within a distributed system. The cloud server 130 can receive the original question sent by the terminal, break down the original question into multiple sub-questions, determine the importance score corresponding to each sub-question, and distribute the multiple sub-questions and their corresponding importance scores to the edge terminal 120.

[0031] The edge terminal 120 can be a single service node or a service node in a distributed system. The edge terminal 120 can receive multiple sub-problems and their corresponding importance scores, and receive long video files uploaded by terminals. It extracts video frames from the long video files according to a preset frame rate to obtain a candidate video frame set. Combining the importance scores for each sub-problem, it determines the number of upload frames for each sub-problem. For each sub-problem, based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set, it selects the target video frames corresponding to the upload frame count for each sub-problem from the candidate video frame set. Finally, it generates a target video frame set by combining the target video frames corresponding to the upload frame count for each sub-problem, and uploads the target video frame set to the cloud server 130.

[0032] The cloud server 130 uses a preset multimodal reasoning model to reason based on the target video frame set, the original question, and multiple sub-questions to obtain the target reasoning result, and then returns the target reasoning result to the terminal 110.

[0033] Therefore, the cloud server of this application breaks down the original question sent by the terminal into multiple sub-questions and determines the importance score of each sub-question. By sending multiple sub-questions and their corresponding importance scores to the edge, the edge extracts frames from the long video uploaded by the terminal to obtain a candidate video frame set. Based on each sub-question and its corresponding importance score, a specific number of target video frames that meet each sub-question are selected from the candidate video frame set, thus forming a target video frame set containing key frames for multiple sub-questions. The edge only needs to upload the target video frame set containing only key frames to the cloud server for recognition, without uploading the complete long video. In this way, the bandwidth usage is reduced under limited communication bandwidth conditions, the communication bandwidth overhead is reduced, the transmission efficiency of video resources is improved, and the recognition efficiency of long videos is improved.

[0034] To facilitate understanding, the steps of the cloud-edge collaborative long video semantic analysis method will be described in detail below. It should be noted that the order of the following embodiments is not intended to limit the preferred order of the embodiments.

[0035] See Figure 2 , Figure 2 This is a flowchart illustrating the steps of a cloud-edge collaborative long video semantic analysis method provided in an embodiment of this application. In this embodiment, the cloud-edge collaborative long video semantic analysis method can be executed by at least two computer devices, with at least two computer devices serving as... Figure 1 Taking edge device 120 and cloud server 130 as examples, the long video semantic analysis method based on cloud-edge collaboration is executed. The specific process is as follows: 101. The cloud server receives the original question sent by the terminal, breaks down the original question into multiple sub-questions, determines the importance score corresponding to each sub-question, and sends the multiple sub-questions and their corresponding importance scores to the edge.

[0036] With the development of large language models and multimodal large models, long video understanding has become a key capability in scenarios such as smart cameras, embodied intelligence, smart homes, robot assistants, industrial inspection, and in-vehicle intelligent agents. Long video understanding tasks typically require the system to locate visual evidence related to the user's natural language questions in videos containing a large number of redundant frames and complex temporal relationships, and further perform event recognition, causal judgment, temporal reasoning, action relationship analysis, or video question answering.

[0037] However, in related technologies, cloud servers receive raw long videos or a large number of sampled video frames uploaded by terminals, along with natural language questions uploaded by the terminals. A large model is then used to perform semantic understanding and complex reasoning on the received raw long videos or large numbers of sampled video frames to obtain answers to the natural language questions. However, this method suffers from drawbacks when uploading raw long videos or large numbers of sampled video frames. Factors such as high video resolution, long video duration, and weak network signals result in high communication bandwidth overhead and extended transmission time when uploading to the cloud server, reducing the efficiency of long video recognition.

[0038] To address the above issues, the embodiments of this application can be applied to a long video semantic analysis system. This system includes at least a cloud server and an edge device. First, the cloud server receives the original question sent by the terminal, breaks it down into multiple sub-questions, and determines the importance score for each sub-question. The cloud server then distributes the sub-questions and their corresponding importance scores to the edge device. In this way, the cloud server decomposes the received original question into multiple sub-questions and assigns an importance score to each sub-question as a basis for controlling the number of video frames selected, laying the foundation for cloud-edge collaborative processing. Next, the edge device receives the multiple sub-questions and their corresponding importance scores, and receives the long video file uploaded by the terminal. It then extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. This edge device preprocessing reduces the amount of data to be filtered, avoiding bandwidth loss caused by transmitting the entire original long video. Finally, the edge device, based on the importance score of each sub-question, determines the number of frames to be uploaded for each sub-question. Thus, with limited communication bandwidth resources, core data is prioritized. The number of frames for each sub-problem is controlled at the source to manage the overall amount of uploaded data. Then, for each sub-problem, the edge device selects the target video frames corresponding to the number of uploaded frames for each sub-problem based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set. This targeted filtering of related video frames by sub-problem retains only the effective images required for each sub-problem, eliminating irrelevant and redundant frames, reducing the number of uploaded video frames under limited communication bandwidth, and improving subsequent recognition efficiency. Furthermore, the edge device generates a target video frame set based on the target video frames corresponding to the number of uploaded frames for each sub-problem and uploads the target video frame set to the cloud server. This uploads only a small number of filtered key frames, significantly reducing transmission bandwidth usage and shortening transmission latency. Finally, the cloud server uses a preset multimodal inference model to infer the target video frame set, the original problem, and multiple sub-problems to obtain the target inference result, and returns the target inference result to the terminal. Thus, by combining the advantages of cloud-based algorithms, efficient inference is achieved on the streamlined target video frame set, improving the overall recognition efficiency of long videos while ensuring recognition accuracy.

[0039] In this embodiment, to avoid receiving long video files uploaded by the terminal under limited communication bandwidth resources, which would affect transmission time, the terminal can send the original question input by the user to the cloud server. The cloud server can receive the original question sent by the terminal, break it down into multiple sub-questions, determine the importance score corresponding to each sub-question, and send the multiple sub-questions and their corresponding importance scores to the edge device. Thus, by breaking down the original question and setting corresponding importance scores, the edge device is instructed to subsequently collect relevant target video frames for each sub-question based on the importance scores, and upload the target video frames of each sub-question to the cloud server for recognition processing. In this way, through cloud-edge collaboration, only target video frames related to multiple sub-questions are uploaded under limited communication bandwidth resources, reducing the amount of data transmitted in the video file, shortening transmission time, and improving the overall recognition efficiency of long videos. This changes the existing technology's passive reception of the entire video in the cloud, relying on question decomposition to clarify the evidence requirements of each part and using importance scores to achieve resource quantification, laying the foundation for precise frame control at the edge and avoiding the uploading of invalid frames.

[0040] The original question is a complete natural language query input by the user on the terminal side. It is a long video query instruction containing complete semantic logic and covering multiple events, temporal sequences, or causal relationships. It serves as the original basis for subsequent sub-question decomposition, keyframe filtering, and cloud-based inference. For example, the original question could be "Did the person turn off the lights after cooking?"

[0041] Among them, multiple sub-problems can be several atomic queries obtained by the cloud server after decomposing the original problem. Each sub-problem corresponds to only a single fact, action or state that can be verified through video footage. The composite logic embedded in the original problem is stripped away to guide the edge end to retrieve corresponding video evidence in different dimensions.

[0042] Each sub-problem refers to any independent atomic query unit after splitting, with independent image verification requirements. It is the smallest execution unit for allocating frame quotas and performing frame similarity matching sequentially at the edge. Each sub-problem can describe a single object, action, state, moment, or event that can be verified through video images.

[0043] The importance score is a quantitative parameter assigned by the cloud server based on the contribution percentage of each sub-question in the original question's logical chain. It characterizes the criticality of the video evidence for each sub-question to the generation of the final answer, serving as a quantitative criterion for differentiated allocation of upload frames at the edge. Specifically, the cloud server generates a quantitative score based on the event logic sequence between sub-questions and the constructed directed acyclic graph, combined with the contribution of each sub-question to the original question's logical chain. This score characterizes the criticality of the video evidence corresponding to a single sub-question to the generation of the original question's answer, and is the quantitative criterion for differentiated allocation of upload frames for each sub-question at the edge. In some implementations, the importance score of each sub-problem relative to the original problem is determined based on the logical order of events among multiple sub-problems. Step 101, "determining the importance score corresponding to each sub-problem," may include: determining the logical order of events among multiple sub-problems, and constructing a directed acyclic graph (DAG) for the multiple sub-problems according to the logical order of events, wherein the DAG represents the temporal relationship between the multiple sub-problems; and generating an importance score for each sub-problem relative to the original problem based on the DAG.

[0044] The logical sequence of events can be the objective temporal order, causal pre- and post-dependent relationships between the sub-problems obtained by the cloud server after breaking down the original problem. This reflects the sequential logic of the events to be verified in the videos and serves as the factual basis for constructing the sub-problem association structure.

[0045] The directed acyclic graph can be a topological structure graph constructed with each sub-problem as an independent node and the logical order of events obtained from sorting out as oriented edges. It is used to intuitively and quantitatively represent the temporal and causal dependency constraints between each sub-problem, and to provide structured data support for the quantitative calculation of the importance score of the sub-problem.

[0046] Specifically, after decomposing the original problem into all subproblems, the process first involves identifying the implicit sequence of events and the causal dependencies within each subproblem. Then, using each subproblem as a node and the temporal relationships between subproblems as directional edges, a directed acyclic graph (DAG) is constructed to represent the temporal dependencies between the subproblems. Furthermore, after constructing the topological structure of the DAG, the importance score of each subproblem relative to the original problem is quantitatively calculated based on its topological position, its degree of dependency on the entire problem's logical chain, and its contribution weight to the derivation of the complete answer. Therefore, by sorting out the logical sequence of events and building a directed acyclic graph (DAG) to represent the temporal relationship, the objective causal and temporal constraints between the sub-problems after the original problem is decomposed can be fully solidified, avoiding the isolated evaluation of the value of individual sub-problems and providing an objective logical basis for importance scoring. Based on the topological relationship of the DAG, the importance score can be generated, which can accurately distinguish the criticality of different sub-problems in the answer formation process, so that the importance score fits the internal logic of the problem. Subsequently, the edge can allocate the number of uploaded frames differently according to the score, ensuring that key sub-problems on the logical chain get sufficient frame resources, preventing the loss of key evidence due to insufficient frames, and improving the key information omission problem caused by traditional unified frame allocation and global similarity screening from the logical root, thereby improving the screening effectiveness of key frames under limited transmission bandwidth.

[0047] By breaking down the original problem and assigning corresponding importance scores, the cloud server instructs the edge device to collect relevant target video frames for each sub-problem based on the importance scores. The target video frames for each sub-problem are then uploaded to the cloud server for recognition and processing. In this way, through cloud-edge collaboration, only target video frames related to multiple sub-problems are uploaded under limited communication bandwidth resources, which reduces the amount of data transmitted in the video file, shortens the transmission time, and improves the recognition efficiency of the entire long video.

[0048] 102. The edge device receives multiple sub-problems and the importance score corresponding to each sub-problem, and receives long video files uploaded by the terminal. It then extracts video frames from the long video files according to a preset frame rate to obtain a set of candidate video frames.

[0049] In this embodiment, after the cloud server sends multiple sub-problems and their corresponding importance scores to the edge, the edge receives these sub-problems and their corresponding importance scores, and receives a long video file uploaded by the terminal. It then extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. Thus, the edge simultaneously receives the sub-problems, importance scores, and the terminal's long video file from the cloud, and forms a candidate video frame set by extracting frames based on the preset frame rate. On the one hand, unlike related technologies that directly upload the entire long video to the cloud, this method performs video frame extraction preprocessing locally at the edge, pre-screening a large number of redundant frames in the video, reducing the amount of data to be transmitted from the source, and effectively reducing subsequent data transmission bandwidth consumption and latency. On the other hand, the edge pre-stores the sub-problem and importance score data sent from the cloud, and can subsequently directly use this quantitative indicator to carry out differentiated frame allocation and targeted frame screening, achieving precise linkage between frame selection logic and front-end problem decomposition results, providing preliminary data support for finely screening target video frames and avoiding the upload of invalid frames.

[0050] The long video file can be raw video data captured by the terminal, containing continuous temporal frames and a large number of redundant video frames. It serves as the data source for edge-end frame extraction processing and subsequent selection of target video frames. The preset frame rate can be a fixed sampling parameter for extracting video frames per unit time. The edge end regularly extracts video frames from the long video file based on this parameter to generate a candidate video frame set.

[0051] Specifically, after acquiring the complete long video file at the edge, the video temporal data stream is traversed according to pre-configured frame rate parameters. Images are extracted frame-by-frame from the corresponding continuous video frames at fixed time sampling intervals. All extracted video frames are then aggregated and organized into a candidate video frame set. In this way, by systematically extracting frames according to the preset frame rate, redundant continuous frames with repetitive content in the long video are eliminated. This reduces the volume of the original video data while preserving the video's temporal information, preventing the original complete video from directly participating in subsequent similarity calculations and reducing the computational overhead of edge feature encoding. Simultaneously, a standardized candidate frame set is formed, providing a standardized data foundation for subsequent feature matching and targeted selection of target frames in conjunction with sub-problems. This reduces the probability of invalid frames being selected for upload during preprocessing, and helps control cloud transmission bandwidth and latency.

[0052] Through the above methods, the edge device synchronously receives sub-questions, importance scores, and long video files from the cloud. Based on a preset frame rate, it extracts frames to form a candidate video frame set. In this way, video frame extraction preprocessing is completed locally at the edge, filtering out a large number of redundant images in the video in advance, reducing the amount of data to be transmitted from the source, and avoiding the direct full upload of the complete long video to the cloud, effectively reducing the bandwidth consumption and transmission latency of subsequent data transmission. On the other hand, the edge device pre-stores the sub-questions and importance score data from the cloud, and can directly carry out differentiated frame allocation and targeted frame screening based on these quantitative indicators. This achieves precise linkage between the frame selection logic and the front-end problem decomposition results, providing pre-data support for fine-tuning the selection of target video frames and avoiding the upload of invalid frames.

[0053] 103. At the edge, the number of upload frames corresponding to each sub-problem is determined by combining the importance score of each sub-problem.

[0054] In this embodiment, after the edge device extracts a set of candidate video frames from the long video file, it determines the number of upload frames for each sub-problem based on the importance score corresponding to each sub-problem. Thus, by using the importance score, which represents the degree of importance of each sub-problem, the number of upload frames for each sub-problem is allocated. This ensures that sub-problems with higher importance are allocated more upload frames, while sub-problems with lower importance are allocated relatively fewer upload frames. Under the constraint of a fixed total upload resource, this prioritizes ensuring that key sub-problems have sufficient candidate frame selection space, effectively preventing the loss of key evidence required for core logic due to insufficient frames, while reducing bandwidth waste caused by secondary sub-problems occupying too much transmission resources. This optimizes the total amount of data uploaded from the source of frame resource allocation.

[0055] The number of uploaded frames can be the number of capture frames allocated for the corresponding sub-problem, i.e., how many video frames need to be captured for this sub-problem, and how many video material frames corresponding to this sub-problem need to be uploaded to the cloud server under limited communication bandwidth resources. Specifically, the number of uploaded frames can be understood as the maximum number of target video frames that can be filtered and uploaded to the cloud server for a single sub-problem, allocated based on the importance score of the corresponding sub-problem, and is the resource upper limit indicator for filtering key frames at the edge.

[0056] In some implementations, a preset total number of upload frames and a preset minimum number of frames for each sub-problem are obtained. Based on the preset total number of upload frames, the preset minimum number of frames, and the importance score corresponding to each sub-problem, the number of upload frames corresponding to each sub-problem is determined. For example, step 103 may include: (103.1) Obtain the preset total number of uploaded frames and the preset minimum number of frames for each sub-problem; (103.2) Determine the number of problems corresponding to multiple subproblems; (103.3) Combine the preset total number of uploaded frames, the preset minimum number of frames, the number of questions, and the importance score corresponding to each sub-question to determine the number of uploaded frames corresponding to each sub-question.

[0057] The preset total upload frame count can be a pre-defined upper limit on the total number of video frames allowed to be uploaded to the cloud after summing all sub-problems at the edge. This limit is used to constrain the overall upload data scale and manage transmission bandwidth consumption. The preset total upload frame count is determined based on the limited communication bandwidth. Specifically, it can be calculated by combining the actual available communication bandwidth resources between the edge and the cloud server, calculating the transmission cost per frame based on the size of the data per unit frame, and considering link stability and real-time transmission latency constraints. This allows for the calculation of the maximum number of video frames that the edge can stably upload at a time, and then fixing this value as the preset total upload frame count. In this way, by limiting the total frame limit based on actual bandwidth resources, the data uploaded by the edge is prevented from exceeding the link's carrying capacity. This top-level constraint on the overall transmission volume prevents network congestion and transmission timeouts caused by excessive instantaneous data, accurately matches the hardware transmission conditions of the edge-cloud link, and continuously reduces bandwidth consumption and transmission latency.

[0058] The preset minimum number of frames can be the minimum number of frames that a single sub-problem can obtain, ensuring that any sub-problem can obtain the basic number of search frames and avoiding the situation where there are no available filter frames for the sub-problem.

[0059] The number of questions can be the total number of all sub-questions obtained after splitting the original question in the cloud, providing a basis for calculating the total occupancy of the basic minimum frame for all sub-questions.

[0060] Specifically, firstly, the edge retrieves the pre-configured total number of upload frames and the unified minimum number of frames for each sub-problem; then, it counts the total number of all sub-problems currently pending; finally, it calculates the total number of frames required based on the total number of sub-problems and the minimum guaranteed number of frames for each sub-problem, and subtracts the minimum guaranteed number of frames from the total number of upload frames to obtain the remaining number of frames that can be flexibly allocated. Based on the importance score of each sub-problem, a weight is calculated, and the remaining number of frames is allocated differentially. After adding the minimum guaranteed number of frames for each sub-problem, the dedicated number of upload frames for each sub-problem is finally determined. Thus, by first configuring a basic frame quota for all sub-problems by setting a minimum number of frames, we can ensure that each sub-problem has a minimum number of search samples and prevent minor sub-problems from completely lacking corresponding evidence due to an allocation of zero frames. The total number of uploaded frames constrains the overall upper limit of uploaded data, thereby controlling the scale of transmitted data globally and managing bandwidth consumption and transmission latency. Then, based on the importance score, we complete the weighted allocation of the remaining frames, so that high-scoring key sub-problems receive more frames and low-scoring non-key sub-problems receive fewer frames. This breaks the drawback of evenly distributing frames and prioritizes the evidence required for core reasoning under limited transmission resources, while taking into account both the integrity of evidence and the lightweight nature of uploaded data. This avoids resource mismatch and omission of key evidence caused by fixed frame selection for each sub-problem.

[0061] In some implementations, the initial total number of frames occupied for multiple sub-problems and the remaining allocated frames can be determined based on a preset minimum frame number. When the remaining allocated frames are greater than or equal to a preset frame number threshold, the remaining allocated frames are redistributed to determine the number of upload frames for each sub-problem. When the remaining allocated frames are less than the preset frame number threshold, the preset total number of upload frames is preferentially allocated to sub-problems with higher importance scores according to the preset minimum frame number, thus obtaining the number of upload frames for each sub-problem. For example, step (103.3) may include: (103.3.1) Determine the initial total number of frames occupied for multiple sub-problems by combining the number of problems and the preset minimum number of frames, and determine the remaining number of allocated frames based on the difference between the preset total number of uploaded frames and the initial total number of frames occupied; (103.3.2) When the remaining number of allocated frames is greater than or equal to the preset frame number threshold, the allocation weight corresponding to each sub-problem is determined by combining the importance score corresponding to each sub-problem. The remaining number of allocated frames is allocated to multiple sub-problems by combining the allocation weight corresponding to each sub-problem to obtain the redistributed frame number of each sub-problem. The number of uploaded frames corresponding to each sub-problem is obtained by combining the redistributed frame number of each sub-problem and the preset minimum frame number. (103.3.3) When the remaining number of allocated frames is less than the preset frame number threshold, the score ranking relationship between multiple sub-problems is determined by combining the importance score corresponding to each sub-problem, and at least one sub-problem ranked first in the score ranking relationship is selected for frame allocation according to the preset minimum frame number, so as to obtain the number of upload frames corresponding to each sub-problem.

[0062] The initial total number of frames occupied can be the total number of frames obtained by multiplying the number of sub-problems by the preset minimum number of frames, which is the number of frames occupied by the minimum quota reserved for all sub-problems.

[0063] The remaining allocated frames can be the number of frames remaining after deducting the initial total occupied frames from the preset total uploaded frames, and can be flexibly allocated according to the importance of the sub-problems.

[0064] The preset frame rate threshold can be a pre-defined critical frame rate value used to determine whether the remaining allocated frames meet the conditions for weighted fine-grained allocation, serving as the basis for switching between two different frame rate allocation strategies. For example, the preset frame rate threshold can be set to 0.

[0065] The allocation weight can be a quantified coefficient obtained by taking the ratio of the importance score of each sub-problem to the total score of all sub-problems, representing the resource priority of a single sub-problem in the allocation of the remaining frames.

[0066] The number of redistributed frames can be the additional frames obtained by splitting the remaining allocated frames based on the allocation weight, which are then added to the preset minimum number of frames to form the final number of uploaded frames for the sub-problem.

[0067] The scoring ranking relationship can be a sequential order formed by ranking the importance scores of each sub-problem from high to low, which can be used as a basis for determining the optimal allocation of frames when frame resources are scarce.

[0068] Specifically, firstly, the number of problems is multiplied by the preset minimum frame count for each sub-problem to calculate the initial total number of frames occupied by all sub-problems. Then, the preset total number of upload frames is subtracted from the initial total number of occupied frames to calculate the remaining number of frames available for flexible allocation. Next, the remaining number of allocated frames is compared with a preset frame count threshold to determine the relationship between the two thresholds and obtain the comparison result. In this way, by first calculating the basic occupied frames and the remaining adjustable frames, the overall upload volume is controlled, ensuring that all frame resources do not exceed the preset total frame count limit, continuously constraining the amount of transmitted data, and controlling bandwidth and latency.

[0069] On one hand, if the remaining allocated frames are greater than or equal to a preset frame threshold (e.g., greater than 0), the allocation weight is calculated based on the importance score of each sub-question. For example, the importance scores of each sub-question are summed to obtain the total score for multiple sub-questions, and the allocation weight for each sub-question is determined based on the ratio between the importance score of each sub-question and the total score. Further, the remaining frames are split according to the allocation weight of each sub-question to form the redistributed frame count for each sub-question. Finally, the redistributed frame counts of each sub-question are added to a preset minimum frame count to obtain the final uploaded frame count for each sub-question. In this way, resources are allocated based on importance weights when there are sufficient remaining frames, allowing high-scoring key sub-questions to receive additional frame quotas, thereby increasing the retrieval sample size for key evidence.

[0070] On the other hand, if the remaining allocated frames are less than the preset frame threshold, it means that the preset total upload frames cannot meet the frame resource allocation for multiple sub-problems according to the preset minimum frame count. Therefore, all sub-problems can be sorted in descending order of importance score to obtain the score ranking relationship. Then, according to the preset minimum frame count, the preset total upload frames are preferentially allocated to the higher-ranked sub-problems to determine the upload frame count for each sub-problem. In this way, when the remaining frames are scarce, priority allocation is given to the best, discarding the minimum allocation for some low-scoring sub-problems, and maximizing the evidence collection needs of high-importance sub-problems in resource-scarce scenarios. This overcomes the drawbacks of resource waste or insufficient frames for key sub-problems caused by fixed average frame allocation, and optimizes frame resource utilization under limited bandwidth conditions.

[0071] In some implementations, since redistributing the remaining allocated frames according to the allocation weight may involve non-integer issues, which does not conform to the video frame acquisition standard, when redistributing the remaining allocated frames according to the allocation weight, each redistributed frame is rounded down, and the remaining amount is redistributed to the sub-problems with high importance scores in order to make full use of communication bandwidth resources in the future. For example, step (103.3.2), "allocating the remaining allocation frames to multiple sub-problems based on the allocation weights corresponding to each sub-problem to obtain the redistribution frame number for each sub-problem," may include: determining the initial redistribution value for each sub-problem based on the allocation weights corresponding to each sub-problem and the remaining allocation frames, and rounding down the initial redistribution value for each sub-problem to obtain the candidate redistribution frame number for each sub-problem; calculating the total number of candidate redistribution frames for multiple sub-problems based on the candidate redistribution frame number for each sub-problem, and determining the redistribution surplus between the remaining allocation frames and the total number of candidate redistribution frames; determining the scoring ranking relationship between multiple sub-problems based on the importance scores corresponding to each sub-problem, and sequentially allocating the redistribution surplus to at least one sub-problem ranked first in the scoring ranking relationship to obtain the secondary redistribution frame number for each sub-problem; and determining the redistribution frame number for each sub-problem based on the corresponding candidate redistribution frame number and the secondary redistribution frame number.

[0072] The initial redistribution value can be a floating-point value obtained by multiplying the allocation weight of the corresponding individual subproblem by the remaining allocation frames. It represents the theoretical number of frames that can be allocated to the corresponding subproblem when the frame number is not constrained to be an integer. The candidate redistribution frame number can be an integer number of frames obtained by rounding down the initial redistribution value, or the basic additional frame number of the subproblem determined based on the weight allocation.

[0073] The total number of candidate redistribution frames can be the sum of the candidate redistribution frames of all subproblems, used to calculate the remaining frame resources consumed in the integer allocation stage.

[0074] The remaining amount of redistribution can be the integer number of frames remaining after deducting the total number of candidate redistribution frames from the remaining redistribution frames. It is the scattered frame resources left over after the rounding operation that need to be redistributed in a second round of selection.

[0075] The number of frames to be redistributed in the second round can be the number of additional frames selected from the remaining amount of redistribution and sent to the corresponding subproblems in order of importance. It is the frame quota obtained by redistributing the remainder after rounding.

[0076] It should be noted that in the process of quantifying and allocating the remaining allocation frames based on the allocation weights of each sub-problem, since the allocation weights are small quantified coefficients obtained by multiplying the weights by the remaining allocation frames, the resulting initial redistribution value is usually a non-integer value. However, video frames are indivisible, independent, and visual data units. The business logic of edge video extraction, filtering, and cloud transmission all use a complete single frame as the smallest execution unit. Non-integer frame numbers do not conform to the actual execution standards of video frame acquisition, filtering, and uploading, and cannot be directly applied to actual frame allocation. Therefore, in the implementation method of this application, during the weight allocation of the remaining frame numbers, the non-integer initial redistribution values ​​corresponding to each sub-problem are first uniformly rounded down, converting the theoretically calculated floating-point frame numbers into integer candidate redistribution frame numbers that conform to actual business specifications, thereby ensuring that the frame allocation results can be implemented. At the same time, the rounding down operation will generate some unallocated, scattered remaining frames. These frames are usable and effective transmission resources. If they are directly discarded, it will cause the limited edge-cloud communication bandwidth resources to be idle and wasted, and the transmission value of the limited bandwidth cannot be maximized. Therefore, a second optimal allocation is performed on the remaining redistribution amount after rounding. Based on the importance of each sub-problem to the reasoning result of the original problem, the scattered remaining frames are preferentially allocated to the core sub-problems with higher importance scores and greater impact on the reasoning result. This not only solves the technical problem of the mismatch between the non-integer weight allocation and the integer acquisition and transmission rules of video frames, but also fully utilizes all available frame resources, achieving the ultimate utilization of limited communication bandwidth resources. At the same time, the effective frame quota of the core sub-problems is specifically improved, taking into account the rationality, standardization and resource utilization of frame allocation, and effectively avoiding the resource waste and lack of core evidence caused by the traditional equal allocation and direct discarding of the remaining frames.

[0077] Specifically, firstly, the allocation weight corresponding to each sub-problem is multiplied by the remaining number of allocated frames to calculate the initial redistribution value for each sub-problem. This value is generally non-integer and cannot be directly used as the actual number of video frames available for allocation. Therefore, all initial redistribution values ​​are uniformly rounded down to generate candidate redistribution frame numbers for each sub-problem. Then, the remaining frame number is split by multiplying the weights and rounding down, ensuring that the base allocation frame number is a valid integer while adhering to the weight allocation ratio, thus adapting to the business rules for actual frame selection and uploading. Next, the candidate redistribution frame numbers for all sub-problems are summed to obtain the total number of candidate redistribution frames. The remaining allocation frame number is then subtracted from the total number of candidate redistribution frames to calculate the remaining redistribution amount after allocating integer frames. In this way, by statistically calculating the total number of candidate redistribution frames, the scattered remaining unallocated frames are locked in, avoiding the problem of idle and wasted frame resources. Subsequently, a ranking relationship is established based on the importance scores of sub-problems from high to low. The remaining redistribution frames are then sequentially added to the higher-ranking sub-problems, thus generating a secondary redistribution frame count for each sub-problem. This process of allocating remaining frames according to importance scores prioritizes sub-problems more critical to answer generation, further aligning with the actual evidence requirements of each sub-problem. Finally, the candidate redistribution frame count and the secondary redistribution frame count for the same sub-problem are added together, and the combined result is the final redistribution frame count for that sub-problem. This method of combining candidate redistribution frame counts with secondary redistribution frame counts ensures that the overall allocation result conforms to the weighting logic while maximizing the utilization of all remaining adjustable frame resources. Under the constraints of limited total uploaded frames and limited communication bandwidth, this further optimizes the frame quota for key sub-problems, reduces the probability of missing key evidence, and continuously achieves the invention's effects of reducing uploaded bandwidth and optimizing evidence screening quality.

[0078] By using the above method, the edge can allocate the number of upload frames for each sub-problem based on the importance score representing the importance of each sub-problem. This allows sub-problems with higher importance to be allocated more upload frames, while sub-problems with lower importance are allocated relatively fewer upload frames. In this way, under the constraint of a fixed total upload resource, key sub-problems are given priority to ensure that they have sufficient candidate frame filtering space, effectively avoiding the loss of key evidence required for core logic due to insufficient frames. At the same time, it reduces the bandwidth waste caused by secondary sub-problems occupying too much transmission resources, thus optimizing the total amount of data uploaded from the source of frame resource allocation.

[0079] 104. At the edge, for each sub-problem, the target video frame corresponding to the number of uploaded frames for each sub-problem is selected from the candidate video frame set based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set.

[0080] In this embodiment, after allocating the corresponding number of upload frames for each sub-problem, the edge device can select the target video frame corresponding to the number of upload frames for each sub-problem from the candidate video frame set, based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set. In this way, by using the semantic similarity between the sub-problem and the video frame, matching video frames are selected for each sub-problem, realizing the targeted selection of key video frames for each separated sub-problem. Only the effective images required for each sub-problem are selected, and irrelevant and redundant frames are further eliminated. This reduces the number of video frame data uploaded under limited communication bandwidth resources, and accurately obtains the target video frame matching each sub-problem, improving the accuracy of the subsequent answer based on the semantic analysis results of long videos.

[0081] The feature similarity can be a quantitative matching value obtained by cross-modal feature comparison between the sub-question feature matrix and the frame feature matrix. It is used to characterize the degree of fit between the picture content of a single candidate video frame and the corresponding sub-question query intent, and is the core quantitative evaluation criterion for screening valid evidence frames.

[0082] The target video frame can be a corresponding sub-problem. The candidate video frames with high matching degree, selected based on feature similarity and whose number matches the number of uploaded frames for the corresponding sub-problem, are the final effective video evidence frames that are aggregated and uploaded to the cloud server for multimodal reasoning.

[0083] In some implementations, each candidate video frame in the candidate video frame set can be encoded to obtain a corresponding frame feature matrix, and each sub-problem can be encoded to obtain a corresponding sub-problem feature matrix. For each sub-problem, based on the feature similarity between the corresponding sub-problem feature matrix and each frame feature matrix, candidate video frames corresponding to the number of uploaded frames are selected from the candidate video frame set to obtain the target video frame corresponding to each sub-problem. For example, step 104 may include: (104.1) Encode each candidate video frame in the candidate video frame set to obtain the corresponding frame feature matrix, and encode each sub-problem to obtain the corresponding sub-problem feature matrix; (104.2) For each sub-problem, the feature matrix of the corresponding sub-problem is compared with the feature matrix of each frame to obtain the feature similarity between the current sub-problem and each candidate video frame; (104.3) For each sub-problem, select the candidate video frames corresponding to the number of uploaded frames from the candidate video frame set in descending order of feature similarity to obtain the target video frame for each sub-problem.

[0084] The frame feature matrix can be a structured multidimensional numerical matrix obtained by encoding the image features of a single candidate video frame at the edge. For example, it can be obtained by encoding each candidate video frame separately using a lightweight visual encoder. Each frame feature matrix is ​​used to quantify and characterize the image content, scene information, object features, and visual semantics of the video frame. It is an image feature data carrier that realizes the digital and computable expression of video frames.

[0085] The sub-question feature matrix can be a structured multidimensional numerical matrix output by semantically encoding the corresponding individual sub-question text at the edge, for example, by encoding each sub-question separately using a text encoder or a visual language model. This sub-question feature matrix is ​​used to quantify the textual semantics, query intent, and verification requirements of the sub-questions, transforming natural language query instructions into textual feature data carriers that can be compared with image features.

[0086] It should be noted that after obtaining the candidate video frame set and the specific number of uploaded frames corresponding to each sub-question, a unified structured feature transformation is performed on the text semantics and image images to achieve accurate matching between text sub-questions and video frames. First, the edge processing unit performs feature encoding on each candidate video frame in the candidate video frame set, transforming each visualized video image into a frame feature matrix with uniform dimensions that can be used for numerical calculation. Simultaneously, semantic encoding is performed on the text content of each decomposed sub-question, transforming the natural language form of the sub-question semantic information into a structured sub-question feature matrix. This completes the same-dimensional feature digitization conversion between images and text, eliminating the barrier of incomparability between cross-modal data. In this way, by performing feature encoding on video frames and sub-questions separately, unstructured video images and natural language questions are transformed into quantifiable and computable feature matrices, achieving accurate cross-modal matching between video images and text questions. This avoids the shortcomings of traditional global random frame sampling and fixed-interval frame sampling, which result in images being irrelevant to the query question and having poor matching.

[0087] Then, for each sub-question, the sub-question feature matrix corresponding to that sub-question is compared with the frame feature matrices of all video frames in the candidate video frame set using cross-modal feature similarity calculation. This quantifies the degree of matching between the current sub-question and each candidate video frame, generating a one-to-one corresponding feature similarity value, thus achieving quantitative matching between the video content and the query question. In this way, by calculating feature similarity frame by frame, the evidentiary value of each candidate video frame for the corresponding sub-question can be objectively and accurately determined, abandoning subjective and coarse-grained screening methods and significantly improving the matching accuracy between video frames and sub-question query requirements.

[0088] Finally, for each sub-problem, all candidate video frames are prioritized based on the calculated similarity of all features, following a descending order of similarity values. Then, according to the pre-allocated number of upload frames for that sub-problem, a corresponding number of top-ranked candidate video frames are extracted, ultimately yielding the target video frames suitable for each sub-problem's query requirements. In this way, by prioritizing similarity and strictly adhering to the allocated number of upload frames, each sub-problem can be matched with the optimal and most relevant video evidence. This approach strictly adheres to the edge-end frame upload total limit, without increasing the bandwidth pressure and latency of the edge cloud transmission, while providing a highly matched and effective set of target video frames for the cloud-based multimodal inference model. This frame selection ensures the accuracy and reliability of long-video semantic analysis and inference, effectively solving the problems of large upload redundancy, insufficient effective evidence frame selection, and low recognition accuracy in long videos.

[0089] In some implementations, candidate video frames corresponding to the number of uploaded frames are selected from the candidate video frame set for each sub-question in descending order of feature similarity. These frames are then deduplicated to obtain a set of candidate video frames for each sub-question. The number of supplementary frames for each sub-question is determined based on the number of deduplicated frames. Based on this supplementary number, supplementary video frames corresponding to each sub-question are selected from the candidate video frame set and added to the corresponding set of candidate video frames to obtain the target video frame for each sub-question. This approach fully utilizes limited communication bandwidth resources to upload a sufficient number of video frame materials, thereby improving the accuracy of subsequent answers based on long-video semantic analysis results. For example, step (104.3) may include: (104.3.1) For each sub-problem, select the candidate video frames corresponding to the number of uploaded frames from the candidate video frame set in descending order of feature similarity to obtain the initial video frame set for each sub-problem; (104.3.2) Combine the initial video frame set of each sub-problem to perform deduplication processing to obtain the candidate video frame set corresponding to each sub-problem; (104.3.3) Determine the number of video frames in the candidate video frame set corresponding to each sub-problem, and determine the number of supplementary frames for each sub-problem based on the number of video frames corresponding to each sub-problem and the corresponding number of uploaded frames. (104.3.4) Based on the number of supplementary frames for each sub-problem, select the supplementary video frames corresponding to each sub-problem from the candidate video frame set; (104.3.5) Add the supplementary video frames corresponding to each sub-problem to the corresponding candidate video frame set to obtain the target video frames corresponding to each sub-problem.

[0090] The initial set of video frames can correspond to a sub-problem. Specifically, it is a set of candidate video frames selected according to the descending order of feature similarity between the candidate video frames and the sub-problem, and the corresponding number of uploaded frames. This set is the original set of frames obtained from the initial similarity-based selection, and includes video frames that may have duplicate or redundant frames. The candidate video frame set can be a purified frame set obtained by performing deduplication processing on the initial video frame set corresponding to the respective sub-problem and removing redundant video frames with highly overlapping content, or it can be obtained by deduplicating multiple initial video frame sets corresponding to multiple sub-problems to obtain the candidate video frame set for each sub-problem. Each video frame in this candidate video frame set is a highly matched candidate frame with independent and valid image information and no duplication or redundancy, which serves as the basic frame set for subsequent frame completion.

[0091] The number of video frames can be the total number of valid, non-duplicate video frames actually contained in the candidate video frame set corresponding to a single sub-problem. It represents the actual amount of valid frame resources currently possessed by the sub-problem after deduplication and is the core basis for calculating the number of supplementary frames. The number of supplementary frames can be the difference between the preset number of upload frames for the corresponding single sub-problem and the number of video frames in the candidate video frame set after deduplication. It represents the number of quota frames missing due to frame deduplication in the sub-problem and is the quantitative basis for subsequent secondary selection and supplementation of frames, used to complete the standard upload frame quota of the sub-problem.

[0092] It should be noted that, based on the selection of video frames based on feature similarity, a frame deduplication and frame number completion mechanism is added to maximize the rational use of frame resources. In this way, through multi-level optimization of the selection logic, the problems of image redundancy and frame resource waste that exist in the traditional fixed number of frame selection are effectively solved.

[0093] Specifically, firstly, for each sub-question, based on the pre-calculated feature similarity of each candidate video frame, and in descending order of similarity value, a number of candidate video frames equal to the number of uploaded frames corresponding to that sub-question are selected from the candidate video frame set to form the initial video frame set for each sub-question, completing the initial screening of high-matching video frames. In this way, the initial selection of frames based on similarity ensures that the initially screened frames are all highly matching images that closely meet the query requirements of the sub-question, guaranteeing the validity of the reasoning evidence.

[0094] Then, deduplication processing is performed on the initial video frame sets corresponding to each sub-problem. This deduplication processing can refer to deduplication among multiple initial video frame sets corresponding to multiple sub-problems, or it can be deduplication within the initial video frame set itself. This process removes redundant video frames with highly overlapping content and repeated frames, eliminating the problem of similar duplicate frames occupying frame quotas, and filtering to obtain each candidate video frame set after removing duplicate frames. In this way, deduplication processing of the initial video frame sets can eliminate invalid duplicate frames with homogeneous content, preventing the limited upload frame quota from being occupied by duplicate video frames with no incremental information, and significantly improving the information density and effective information ratio of a single uploaded frame.

[0095] Next, the actual number of video frames in the candidate video frame set corresponding to each sub-problem is counted. The difference between this actual number of frames and the preset number of upload frames allocated for that sub-problem is calculated to accurately determine the number of frames missing due to the deduplication operation. This number serves as the supplementary frame count for each sub-problem. Furthermore, based on the supplementary frame count for each sub-problem, the sorting rule from highest to lowest feature similarity is continued. From the remaining unselected candidate video frames, a corresponding number of high-similarity video frames are selected as supplementary video frames. In this way, by calculating the difference in the number of deduplicated frames, the number of supplementary frames is determined, accurately locating the missing frame quota for each sub-problem and ensuring the accuracy and targeting of frame resource completion. Moreover, by supplementing and selecting high-similarity video frames and merging them into the candidate video frame set, the standard upload frame count for each sub-problem can be completed without exceeding the preset upload frame quota and without consuming additional edge-cloud communication bandwidth. This fully utilizes limited communication bandwidth resources and maximizes the quantity and diversity of effective video materials.

[0096] Finally, the selected supplementary video frames are added one by one to the corresponding candidate video frame set to complete the preset upload frame quota for each sub-question, ultimately forming the target video frame for each sub-question. This avoids the waste of bandwidth resources caused by uploading duplicate frames while ensuring that each sub-question receives a sufficient number of differentiated and highly effective video evidence frames. This enriches the dimensions and quantity of materials for multimodal reasoning in the cloud, effectively improving the completeness and accuracy of subsequent video reasoning results, and balancing the lightweight requirements of edge-cloud transmission with the high-precision needs of long-video question-answering reasoning.

[0097] In some implementations, for each sub-problem, a candidate video frame is selected from the candidate video frame set in descending order of feature similarity. At each selected time step, the currently selected candidate video frame is deduplicated from the previously selected video frames until the number of selected candidate video frames reaches the upload frame count for that sub-problem, thus obtaining the initial video frame set corresponding to that sub-problem. For example, step (104.3.1) may include: For each sub-problem, a separate set of frames to be compared is created, and the initial number of frames in the set of frames to be compared is determined as the number of selected frames, with the initial number of frames being zero. When the number of selected frames is less than the number of uploaded frames for the corresponding sub-problem, a candidate video frame is selected from the candidate video frame set in descending order of feature similarity. If the similarity between the currently selected candidate video frame and each selected video frame in the frame set to be compared is greater than the preset frame similarity threshold, the currently selected candidate video frame is added to the frame set to be compared, and the number of selected frames in the frame set to be compared is updated. When the number of selected frames is detected to be greater than or equal to the number of uploaded frames for the corresponding sub-problem, the set of frames to be compared corresponding to the current number of selected frames is determined as the initial video frame set for the corresponding sub-problem.

[0098] The set of frames to be compared can be a temporary frame storage set independently created by the edge for each individual sub-problem. That is, each sub-problem has a separate set of frames to be compared, which is used to store valid video frames that have been filtered for correlation and deduplication in real time during the iterative frame selection process. This enables the independent collection, iterative update and redundancy filtering of candidate frames for a single sub-problem, and serves as a temporary data carrier for generating the initial set of video frames.

[0099] The selected frame count can be the statistical number of valid video frames currently included in the frame set to be compared. It is initialized to zero and is used to represent the amount of frame resources that have been filtered, verified and stored in the database during the single sub-problem iterative frame selection process in real time. It serves as the quantitative basis for determining whether to terminate the iterative frame selection.

[0100] The preset frame similarity threshold can be a pre-configured critical value for content similarity between video frames. It is used to judge whether there is high overlap of images or information homogeneity and redundancy between any two video frames. If the frame similarity between the current candidate frame and the selected frames in the set exceeds the threshold, the current candidate frame is determined to be a redundant and duplicate frame and will not be included in the set of frames to be compared. It is the core discrimination standard for realizing iterative real-time deduplication and ensuring frame diversity.

[0101] Specifically, firstly, for each sub-problem to be processed, an independent set of frames to be compared is created at the edge, and the number of selected frames in the set is initialized to zero, so that the frame filtering process of each sub-problem is independent and does not interfere with each other, providing a blank statistical and storage medium for frame-by-frame iterative filtering.

[0102] Furthermore, during the frame filtering loop, the number of selected frames corresponding to the current sub-problem is continuously monitored in real time to determine whether the number of selected frames is less than the pre-allocated number of upload frames for that sub-problem. If the number of selected frames does not reach the upload frame quota, the candidate video frame with the highest similarity is selected from the candidate video frame set in a single iteration, based on the descending order of feature similarity between the candidate video frames and the current sub-problem. Further, redundancy verification is performed on the currently selected candidate video frame, calculating the inter-frame similarity between the candidate video frame and all existing selected video frames in the comparison frame set. When the inter-frame similarity between the candidate video frame and all existing selected video frames in the set is greater than a preset similarity threshold, the current candidate video frame is determined to be a valid differentiated frame without content redundancy, and it is added to the comparison frame set. Simultaneously, the number of selected frames in the comparison frame set is updated and increased, completing a single iteration filtering operation. This iterative process of frame-by-frame selection, redundancy verification, frame storage, and quantity update is repeated to continuously replenish valid candidate video frames. On the other hand, when it is detected that the number of selected frames in the set of frames to be compared is greater than or equal to the number of uploaded frames corresponding to the current sub-problem, it is determined that the frame filtering quota is full, the iterative filtering process is terminated, and the set of frames to be compared that stores valid video frames at this time is determined as the initial video frame set corresponding to the sub-problem.

[0103] Therefore, a successive filtering mechanism of single-frame iterative selection and real-time deduplication verification is adopted. Each sub-problem is configured with its own dedicated frame filtering set, achieving logical isolation between multiple sub-problem frame filtering steps. This avoids confusion and cross-interference of frame resources from different sub-problems, ensuring the independence and accuracy of frame filtering for each sub-problem. Compared to a one-time batch frame selection method, this real-time method synchronously completes frame correlation filtering and inter-frame redundancy filtering in each frame selection time step. While prioritizing the retention of video frames highly relevant to the sub-problem, it removes redundant frames with highly overlapping content and homogeneous information in real time. From the initial filtering stage, it avoids the problem of a large number of duplicate frames entering the frame set, significantly improving the effective information density and image diversity of the frame content within the initial video frame set. Through iterative loop control logic and quota termination, the final number of selected video frames is precisely controlled to match the preset upload frame quota, strictly adhering to the bandwidth resource constraints of edge cloud transmission and avoiding bandwidth pressure from excessive frame uploads. The overall filtering logic takes into account both the relevance of frame content and the diversity between frames. Within the limited quota of uploaded frames, it maximizes the filtering of video frames with differentiated and effective information. This ensures that the initial video frames meet the requirements of sub-question verification while reducing the consumption of communication resources by invalid and redundant frames. It provides high-quality basic frame data for subsequent frame completion and target frame determination, effectively improving the accuracy and completeness of cloud-based video reasoning and answer generation.

[0104] For example, a pre-configured weighting coefficient for balancing frame content relevance and inter-frame diversity establishes a two-dimensional evaluation standard for frame selection. The feature similarity between the sub-problem and candidate video frames is used as the relevance evaluation criterion, while the content similarity between the candidate video frame and the selected video frames is used as the redundancy evaluation criterion. A comprehensive evaluation is performed on all unselected candidate video frames within the candidate video frame set. In each round of iterative selection, the optimal candidate video frame that is highly semantically relevant to the current sub-problem and has the lowest content overlap with all selected video frames is selected first and included in the initial candidate frame group corresponding to the current sub-problem. The optimal selection operation is continuously iterated, and after each frame selection, the selected video frame set is updated as the benchmark for the next round of diversity evaluation, until the total number of selected candidate video frames reaches the preset number of uploaded frames corresponding to the sub-problem. All candidate video frames obtained after iterative selection are summarized and integrated to form an initial video frame set that balances problem matching and frame content diversity. In this way, based on similarity-based frame selection, the relevance and diversity of frames are simultaneously constrained. This ensures that the selected video frames fit the sub-question query requirements while avoiding the problems of high image repetition and information redundancy caused by single similarity-based frame selection. This allows the video frames in the initial video frame set to have both high matching degree and content differentiation, maximizing the enrichment of effective video information within the limited upload frame quota, and providing high-quality basic frame data for subsequent deduplication and frame supplementation to improve the accuracy of cloud inference.

[0105] For example, for each sub-problem, the maximum marginal correlation is used. The Relevance (MMR) strategy iteratively selects the best frames from the candidate video frame set. In each iteration, a preset weighting coefficient lambda is used to balance the relevance of frame content and the diversity between frames. The MMR score for each unselected candidate frame is calculated. The formula for calculating the MMR score is: lambda*M(t,i)-(1-lambda)*maxsim(f_t,f_k); where M(t,i) is the feature similarity between the current candidate frame and the corresponding sub-problem, sim(f_t,f_k) is the image similarity between the current candidate frame and each selected video frame in the selected frame set, lambda is a preset weighting coefficient used to balance the semantic relevance of frames and the diversity of frame content, with a value between 0 and 1, and K_i is the selected frame set corresponding to the current sub-problem. In each iteration, the candidate video frame with the maximum MMR score is selected and added to the selected frame set. The selection operation is continuously performed iteratively until the number of selected candidate video frames reaches the number of uploaded frames corresponding to the sub-problem, at which point the iteration terminates, and the initial video frame set corresponding to the current sub-problem is finally obtained. Among them, the MMR strategy ensures that the selected video frames are highly matched with the query intent of the sub-question through the relevance term, and suppresses the redundancy of the picture content between the selected video frames through the diversity penalty term. In the initial frame selection stage, it simultaneously takes into account the matching accuracy of a single frame and the overall difference of the frame set, effectively reducing the computational overhead of subsequent frame deduplication processing. Within the limited upload frame quota, it maximizes the effective information density of the initial video frame set, adapting to the lightweight and high-precision processing requirements of long video edge-cloud collaborative inference.

[0106] By using the above methods, the semantic similarity between sub-questions and video frames can be used to select matching video frames for each sub-question. This enables targeted selection of key video frames for each separated sub-question, selecting only the effective images required for each sub-question, and further eliminating irrelevant and redundant frames. This reduces the number of video frame data uploaded under limited communication bandwidth resources, while accurately obtaining the target video frames that match each sub-question, thus improving the accuracy of subsequent answers based on the semantic analysis results of long videos.

[0107] 105. At the edge, the target video frame set is generated by combining the target video frames corresponding to the number of uploaded frames for each sub-problem, and the target video frame set is uploaded to the cloud server.

[0108] In this embodiment, after obtaining the target video frames for the number of uploaded frames corresponding to each sub-problem at the edge, the edge device can generate a target video frame set by combining the target video frames for the number of uploaded frames corresponding to each sub-problem. Specifically, the target video frames for the number of uploaded frames corresponding to each sub-problem are merged to obtain target video frame sets for multiple problems. Then, the target video frame sets are uploaded to the cloud server, allowing the cloud server to perform inference based on the target video frame sets, the original problem, and the multiple sub-problems using a preset multimodal inference model to obtain the target inference result. It is understood that since the target video frame set contains key video frames collected from the long video file for each sub-problem, the number of video frames it contains is far less than the number of frames in the original long video file. Thus, only a small number of filtered key video frames are ultimately uploaded. Compared to uploading the entire original long video / mass-sampled frames to the cloud, this significantly reduces transmission bandwidth usage, shortens transmission latency, and improves the transmission efficiency of video material data.

[0109] In some implementations, the target video frames corresponding to the number of uploaded frames for each sub-problem can be merged and sorted to obtain a set of target video frames. For example, step 105, "generating a set of target video frames by combining the target video frames corresponding to the number of uploaded frames for each sub-problem," may include: merging the target video frames corresponding to the number of uploaded frames for each sub-problem to obtain a merged video frame set; determining the timestamp corresponding to each target video frame in the merged video frame set in the long video file; and sorting the multiple target video frames in the merged video frame set based on the timestamp corresponding to each target video frame to obtain the target video frame set.

[0110] Each timestamp corresponds to a target video frame and serves as a unique temporal position identifier for that frame within the original long video file. It accurately records the capture time and temporal arrangement of the corresponding video frame, providing the sole quantitative basis for temporal regularization and sequential sorting of the merged multi-sub-problem target video frames. The timestamps reconstruct the true playback sequence of each video frame within the long video, breaking the disordered and mixed state of frames selected from different sub-problems. This ensures that the final set of target video frames strictly adheres to the original video temporal logic, providing accurate temporal prior conditions for cloud-based temporal reasoning and video content evolution analysis.

[0111] First, the edge processing unit merges and summarizes all target video frames, corresponding to their respective uploaded frame numbers, obtained from the screening of each sub-problem. Duplicate and redundant frames are removed during the aggregation process, resulting in a merged video frame set containing valid evidence frames from all sub-problems. This achieves unified aggregation of the screening results from multiple sub-problems. By merging target frames from multiple sub-problems, it can completely aggregate all highly matched and diverse valid evidence frames selected from all sub-problems, integrating multi-dimensional video semantic information and avoiding the fragmentation of evidence caused by the independent dispersion of frames selected from a single sub-problem. This achieves unified aggregation and management of global valid frame resources. Second, for each target video frame in the merged video frame set, its corresponding temporal position information in the original long video file is matched and read one by one. This accurately determines the unique video timestamp corresponding to each target video frame, providing a temporal quantification basis for subsequent ordered sorting. Thus, by matching the original timestamps of each video frame, the original temporal position of each frame is accurately anchored, ensuring that the sorting criteria completely conform to the actual playback logic of the long video, avoiding temporal disorder caused by manual sorting or feature-based sorting. Finally, using the timestamp size of each target video frame as the sorting criterion, and following the playback sequence logic of the original long video, the scattered target video frames in the merged video frame set are sorted in ascending order. The disordered discrete target frames are then rearranged according to the original temporal relationship of the video, ultimately generating a standardized set of target video frames that are temporally continuous and orderly arranged. In this way, global temporal rearrangement is completed based on timestamps, transforming the disordered target frames collected across sub-questions into a temporally ordered set of frames. This ensures that the finally uploaded video frames retain complete video temporal evolution characteristics and scene change logic. Through this process, not only can temporally coherent, informationally complete, and dimensionally rich video evidence be provided to the cloud-based inference model, significantly improving the accuracy of cloud-based understanding of long video content, scene deduction, and question answering, but the standardized frame set structure is also more compatible with edge-cloud data transmission protocols and model input specifications, reducing cloud data preprocessing overhead and further optimizing the overall edge-cloud collaborative inference efficiency.

[0112] Using the above methods, a target video frame set can be generated by combining the target video frames corresponding to the number of uploaded frames for each sub-question, and then the target video frame set can be uploaded to the cloud server. In this way, only a small number of filtered key video frames are uploaded, which significantly reduces the transmission bandwidth usage, shortens the transmission latency, and improves the transmission efficiency of video material data compared to uploading the original long video / large batch of sampled frames to the cloud. This improves the overall efficiency of the long video semantic analysis process and question-and-answer dialogue.

[0113] 106. The cloud server uses a preset multimodal reasoning model to reason based on the target video frame set, the original question, and multiple sub-questions to obtain the target reasoning result, and then returns the target reasoning result to the terminal.

[0114] In this embodiment, after receiving the target video frame set sent by the edge terminal, the cloud server can use a preset multimodal inference model to infer the target inference result based on the target video frame set, the original question, and multiple sub-questions. Furthermore, the cloud server can return the target inference result to the terminal. This achieves question-and-answer based on long video semantic analysis results. Because the cloud server decomposes the complete original question and limits the importance score of each sub-question obtained from the decomposition, the edge terminal allocates a corresponding number of video upload frames for each sub-question according to the importance score, and selectively selects a small number of matching target video frames from the long video file for each sub-question. This reduces the amount of video frame material transmitted under limited communication bandwidth resources, thereby improving the transmission efficiency of video material. Furthermore, it ensures that the precisely selected target video frame set can support the inference of the corresponding original question, improving the accuracy of the answer based on long video semantic analysis results, reducing the overall question-and-answer cycle based on long video, and improving the overall efficiency of long video semantic analysis and question-and-answer dialogue.

[0115] To better understand the embodiments of this application, Figure 3 This is an example diagram of the cloud-edge collaborative long video semantic analysis system architecture provided in the embodiments of this application. Figure 4 The flowchart of long video semantic analysis based on cloud-edge collaboration provided in the embodiments of this application, combined with Figure 3 and Figure 4 As shown, an example is provided to illustrate the process of semantic analysis of long videos based on cloud-edge collaboration, as detailed below: like Figure 3 As shown, the cloud-edge collaborative long video semantic analysis system consists of a device layer (terminal), an edge layer (edge ​​terminal), and a cloud layer (cloud server). The device layer includes a user query interface and a long video acquisition or caching module, used to receive the user's original question q and obtain the long video file V to be analyzed.

[0116] The edge layer includes a video frame extraction and encoding module, a similarity calculation module, and a keyframe selection module. The cloud layer includes a semantic query planner and a large-scale multimodal reasoning module.

[0117] Among them, the cloud-based semantic query planner is used to process the original questions in complex natural language input by the user. Decomposed into several visually verifiable atomic subproblems And construct directed acyclic graphs representing temporal, causal, or logical dependencies. .in, Represents the set of sub-problem nodes. This represents the dependency edges between subproblems. The cloud also assigns a semantic importance score (i.e., importance rating) to each subproblem. This is used to indicate the importance of the visual evidence corresponding to the sub-question in answering the original question.

[0118] The edge layer receives planning instructions from the cloud. Based on the total upload budget and minimum frame rate guarantee Determine the frame budget corresponding to each subproblem. And retrieve the set of key frames that meet the requirements of each sub-problem from the candidate video frames. The cloud-based large model receives the merged keyframe set uploaded from the edge device. In addition to the original question context, the final answer is generated.

[0119] Combination Figure 4 As shown, the process of semantic analysis of long videos based on cloud-edge collaboration is introduced as follows: (S1) Obtain user questions and long videos. The device layer receives natural language questions input by the user. And obtain the long video to be analyzed. .video It can come from a real-time camera, mobile terminal, camera equipment, robot vision sensor, vehicle camera, or local cache file.

[0120] (S2) Semantic planning is performed in the cloud. The device layer will handle user questions. The query is sent to the cloud-based semantic query planner. The cloud-based semantic query planner utilizes a large language model or a multimodal large model to plan the query based on preset prompts. Decomposed into The problem of individual atoms Each sub-problem should describe a single object, action, state, moment, or event that can be verified via video footage.

[0121] (S3) Construct a logical dependency graph and estimate its importance. The cloud-based semantic query planner further generates a directed acyclic graph. , among which the side Representing subproblems Subproblems The temporal premise, causal premise, or contextual premise. Meanwhile, the cloud provides each sub-problem with... Generate importance score , Discrete values ​​from 1 to 10 can be used, or normalized continuous values ​​can be employed. The higher the importance score, the more crucial the evidence corresponding to that sub-question is to answering the original question.

[0122] (S4) The edge end performs video frame extraction and feature encoding. The edge end extracts video frames from the long video file according to a preset frame rate or an adaptive frame rate. Obtain candidate frame set The frame feature matrix is ​​obtained using a lightweight visual encoder. At the edge, a text encoder or a visual language model text branch is also used to encode each sub-problem, resulting in a sub-problem feature matrix. .

[0123] (S5) Allocate the budget based on the importance of the sub-problems. Let the total upload frame budget (i.e., the preset total number of upload frames) be... The minimum guaranteed frame rate (i.e., the preset minimum frame rate) is: The edge first assigns each sub-problem... A base budget is then allocated based on the frame, and scores are awarded according to importance. Calculate normalized weights , with the remaining budget The allocation is weighted among the subproblems. Since the frame number must be an integer, this invention uses the maximum remainder method to convert the ideal quota into an integer frame budget. .

[0124] In the budget allocation process, the normalized weights of each sub-problem are first calculated: Then calculate the remaining budget: .when If the total budget is insufficient to meet the minimum guarantee requirements of all subproblems, then the allocation can degenerate into a uniform distribution or a priority distribution based on importance from highest to lowest. Otherwise, calculate the ideal additional allocation for each subproblem: .Will Round down to the nearest integer. And calculate the decimal remainder. Finally, calculate the number of unallocated remaining frames. ,according to Select from largest to smallest Each sub-problem adds an extra frame to arrive at the final budget. or .

[0125] (S6) Parallel semantic matching and keyframe selection. The similarity matrix between the candidate frame feature matrix and the sub-problem feature matrix is ​​calculated at the edge. This yields the relevance score of each candidate frame to each sub-problem. For the ... Each sub-issue, the edge end according to the corresponding budget Select several highly correlated frames and combine them with the maximum marginal correlation strategy or the neighboring frame redundancy removal strategy to achieve a balance between correlation and diversity.

[0126] Specifically, the edge can obtain the similarity scores between all candidate frames and all sub-problems at once through batch matrix multiplication. For the ... The problem of individual frames, if candidate frames The similarity score is You can select from highest to lowest score. Alternatively, a maximum marginal relevance strategy can be used to iteratively select frames, ensuring that the currently selected frame simultaneously satisfies two conditions: high relevance to the subproblem and no repetition with already selected frames. Specifically, in each iteration, a frame is selected that satisfies these conditions. The largest candidate frame, of which This is a tradeoff coefficient between relevance and diversity.

[0127] Therefore, through the above-mentioned budget allocation and retrieval methods, the present invention does not simply select a few frames most similar to the original problem, but rather ensures that each logical sub-problem receives a retrieval budget that matches its importance, thereby improving the coverage of key evidence and reducing the risk of evidence chain breakage caused by semantic overload.

[0128] (S7) Keyframe set merging and uploading. The edge end merges and uploads the keyframe sets selected from each sub-problem. merged into Duplicate frames are deduplicated and sorted according to their original video timestamps to maintain the narrative order and the integrity of the evidence chain. Subsequently, the edge processing only processes the set... Keyframes, along with necessary timestamps, sub-question tags, and confidence information, are uploaded to the cloud instead of the complete video.

[0129] (S8) Cloud-based deep reasoning and output of the answer. The cloud-based large-scale multimodal reasoning module receives the keyframe set. Original user issues List of sub-problems and logical dependency graph Based on a multimodal large model, cross-frame evidence integration and logical reasoning are performed to generate the final answer, which is then returned to the device layer or the user terminal.

[0130] For example, the user inputs the original question via a mobile terminal or camera system. "Did the person turn off the lights after cooking?", and the system retrieves the corresponding long video. The device layer will address the issue. The video stream or cached video is uploaded to the cloud-based semantic query planner and provided to the edge layer. The cloud-based semantic query planner first identifies that the question contains multiple atomic visual evidence requirements such as "cooking activity", "state after cooking", "turning off the lights", and "sequence of action". Then, it generates a directed acyclic graph containing these nodes and assigns importance scores to each node.

[0131] For example, the cloud can generate sub-questions and weights as follows: v_1 = Recognize cooking activities, ; "Recognizes when cooking is finished or when someone leaves the kitchen." ; "Recognize changes in light status or the action of turning off the lights". ; "Confirm whether the action of turning off the lights occurred after cooking." The corresponding dependencies can include The above planning results constitute the observation plan at the edge, enabling the edge to not only retrieve frames related to "cooking", but also to actively retrieve key evidence related to "turning off the lights" and "action sequence".

[0132] Therefore, by implementing the above example of long video semantic analysis based on cloud-edge collaboration, only the set of key frames filtered by the edge end is uploaded, rather than the complete long video or a large number of redundant frames. This significantly reduces network transmission volume and cloud input scale, and lowers long video upload bandwidth and end-to-end latency. The final inference is still completed by a large-scale multimodal model in the cloud, avoiding the accuracy bottleneck of pure edge small models on complex problems and maintaining cloud-level complex inference capabilities. Complex problems are decomposed into multiple visually verifiable atomic sub-problems, enabling the edge end to retrieve frames according to logical evidence requirements, rather than according to a single global similarity, thus alleviating the semantic overwhelming problem. Frame budgets are dynamically allocated according to the importance of sub-problems, avoiding non-critical visual concepts from occupying too many upload resources and improving the effective evidence coverage under limited frame budgets. At the edge end, budget allocation mainly involves sorting operations, and the number of sub-problems is usually small. Similarity calculation can be performed in parallel through matrix multiplication, making it suitable for deployment on GPU or NPU edge devices with low edge-end overhead.

[0133] Based on the above embodiments, this application can be applied to a long video semantic analysis system. The long video semantic analysis system includes at least a cloud server and an edge terminal. First, the cloud server receives the original question sent by the terminal, breaks it down into multiple sub-questions, determines the importance score for each sub-question, and sends the multiple sub-questions and their corresponding importance scores to the edge terminal. Thus, the cloud server decomposes the received original question into multiple sub-questions and assigns an importance score to each sub-question as a basis for controlling the number of video frames selected, laying the foundation for cloud-edge collaborative processing. Then, the edge terminal receives the multiple sub-questions and their corresponding importance scores, receives the long video file uploaded by the terminal, and extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. Thus, through frame extraction preprocessing at the edge terminal, the amount of data to be filtered is initially reduced, avoiding bandwidth loss caused by the full transmission of the original long video. Next, the edge terminal, combined with the importance score corresponding to each sub-question, determines the number of frames to be uploaded for each sub-question. Thus, under limited communication bandwidth transmission resources, priority is given to ensuring... The number of frames for the core sub-problems is controlled at the source to manage the overall amount of uploaded data. Then, for each sub-problem, the edge device combines the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set to select the target video frames corresponding to the upload frame count for each sub-problem. This targeted filtering of related video frames by sub-problems retains only the effective images needed for each sub-problem, eliminating irrelevant and redundant frames, reducing the number of uploaded video frames under limited communication bandwidth, and improving subsequent recognition efficiency. Furthermore, the edge device generates a target video frame set based on the target video frames corresponding to the upload frame count for each sub-problem and uploads the target video frame set to the cloud server. This uploads only a small number of filtered key frames, significantly reducing transmission bandwidth consumption and shortening transmission latency. Finally, the cloud server uses a preset multimodal inference model to infer the target video frame set, the original problem, and multiple sub-problems to obtain the target inference result, and returns the target inference result to the terminal. Thus, by combining the advantages of cloud-based algorithms, efficient inference is achieved on the streamlined target video frame set, improving the overall recognition efficiency of long videos while ensuring recognition accuracy.

[0134] Therefore, compared to related technologies that upload original long videos or a large number of sampled video frames, which suffer from high communication bandwidth overhead and long transmission time due to factors such as high video resolution, long video duration, and weak network signal, resulting in reduced recognition efficiency for long videos, this application's cloud server breaks down the original question sent by the terminal into multiple sub-questions and determines the importance score of each sub-question. Multiple sub-questions and their corresponding importance scores are then distributed to the edge device. The edge device extracts frames from the long video uploaded by the terminal to obtain a candidate video frame set. Based on each sub-question and its corresponding importance score, a specific number of target video frames that meet each sub-question are selected from the candidate video frame set, thus forming a target video frame set containing key frames for multiple sub-questions. The edge device only needs to upload the target video frame set containing only key frames to the cloud server for recognition, without uploading the complete long video. This reduces the consumption of communication bandwidth under limited communication bandwidth conditions, lowers communication bandwidth overhead, improves the transmission efficiency of video resources, and thus enhances the recognition efficiency for long videos.

[0135] To facilitate better implementation of the cloud-edge collaborative long video semantic analysis method provided in this application, this application also provides a cloud-edge collaborative long video semantic analysis device based on the aforementioned method. The meanings of the terms used are the same as in the cloud-edge collaborative long video semantic analysis method described above, and specific implementation details can be found in the descriptions within the method embodiments.

[0136] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a cloud-edge collaborative long video semantic analysis device provided in an embodiment of this application. The cloud-edge collaborative long video semantic analysis device is integrated into the computer device of this application. This computer device is the edge terminal of the long video semantic analysis system, which also includes a cloud server. The cloud-edge collaborative long video semantic analysis device may include a receiving unit 401, a determining unit 402, a selecting unit 403, and an uploading unit 404.

[0137] The receiving unit 401 is used to receive multiple sub-problems and the importance score corresponding to each sub-problem, as well as a long video file uploaded by the receiving terminal, and to extract video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. Among them, multiple sub-problems are obtained by the cloud server based on the original problem sent by the terminal, and each importance score is generated by the cloud server for the corresponding sub-problem; Unit 402 is used to determine the number of upload frames for each sub-problem by combining the importance score corresponding to each sub-problem. The selection unit 403 is used to select the target video frame corresponding to the number of uploaded frames for each sub-problem from the candidate video frame set, based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set. The upload unit 404 is used to generate a target video frame set by combining the target video frames corresponding to the upload frame number of each sub-problem, and upload the target video frame set to the cloud server. The cloud server then uses a preset multimodal inference model to infer the target inference result based on the target video frame set, the original problem, and multiple sub-problems, and returns the target inference result to the terminal.

[0138] In some implementations, the determining unit is further configured to: Get the preset total number of uploaded frames and the preset minimum number of frames for each sub-question; Determine the number of problems corresponding to multiple sub-problems; By combining the preset total number of upload frames, the preset minimum number of frames, the number of questions, and the importance score corresponding to each sub-question, the number of upload frames corresponding to each sub-question is determined.

[0139] In some implementations, the determining unit is further configured to: The initial total number of frames occupied for multiple sub-questions is determined by combining the number of questions and the preset minimum number of frames, and the remaining number of allocated frames is determined by the difference between the preset total number of uploaded frames and the initial total number of frames occupied. When the remaining number of allocated frames is greater than or equal to the preset frame number threshold, the allocation weight corresponding to each sub-problem is determined by combining the importance score of each sub-problem. The remaining number of allocated frames is then distributed to multiple sub-problems based on the allocation weight of each sub-problem to obtain the redistributed frame number for each sub-problem. Finally, the number of uploaded frames corresponding to each sub-problem is obtained by combining the redistributed frame number of each sub-problem with the preset minimum frame number. When the remaining number of allocated frames is less than the preset frame number threshold, the importance score corresponding to each sub-problem is combined to determine the score ranking relationship between multiple sub-problems. Then, according to the preset minimum frame number, at least one sub-problem ranked first in the score ranking relationship is selected for frame allocation, so as to obtain the number of upload frames corresponding to each sub-problem.

[0140] In some implementations, the determining unit is further configured to: Based on the allocation weight corresponding to each sub-problem and the remaining allocation frames, the initial redistribution value of each sub-problem is determined, and the initial redistribution value of each sub-problem is rounded down to obtain the candidate redistribution frame number of each sub-problem. Calculate the total number of candidate redistribution frames for multiple sub-problems by combining the candidate redistribution frame counts for each sub-problem, and determine the redistribution surplus between the remaining redistribution frames and the total number of candidate redistribution frames; By combining the importance score corresponding to each sub-problem, the score ranking relationship between multiple sub-problems is determined, and the remaining redistribution amount is sequentially allocated to at least one sub-problem ranked first in the score ranking relationship, thus obtaining the number of secondary redistribution frames for each sub-problem. For each subproblem, the number of reassignment frames for each subproblem is determined by combining the corresponding candidate reassignment frame count and the secondary reassignment frame count.

[0141] In some implementations, the selection unit is also used for: Each candidate video frame in the candidate video frame set is encoded to obtain the corresponding frame feature matrix, and each sub-problem is encoded to obtain the corresponding sub-problem feature matrix. For each sub-problem, the feature matrix of the corresponding sub-problem is compared with the feature matrix of each frame to obtain the feature similarity between the current sub-problem and each candidate video frame. For each sub-problem, candidate video frames with the corresponding number of uploaded frames are selected from the candidate video frame set in descending order of feature similarity to obtain the target video frame for each sub-problem.

[0142] In some implementations, the selection unit is also used for: For each sub-problem, candidate video frames with the corresponding number of uploaded frames are selected from the candidate video frame set in descending order of feature similarity to obtain the initial video frame set for each sub-problem. By combining the initial video frame set of each sub-problem, deduplication is performed to obtain the candidate video frame set corresponding to each sub-problem; Determine the number of video frames in the candidate video frame set corresponding to each sub-problem, and determine the number of supplementary frames for each sub-problem based on the number of video frames corresponding to each sub-problem and the corresponding number of uploaded frames; Based on the number of supplementary frames for each sub-problem, select the supplementary video frames corresponding to each sub-problem from the candidate video frame set; Add the supplementary video frames corresponding to each sub-problem to the corresponding set of candidate video frames to obtain the target video frames for each sub-problem.

[0143] As shown above, the cloud server in this application breaks down the original question sent by the terminal into multiple sub-questions and determines the importance score of each sub-question. It then distributes these sub-questions and their corresponding importance scores to the edge device. The edge device extracts frames from the long video uploaded by the terminal to obtain a candidate video frame set. Based on each sub-question and its corresponding importance score, it selects a specific number of target video frames from the candidate video frame set that meet each sub-question, thus forming a target video frame set containing key frames for multiple sub-questions. The edge device only needs to upload the target video frame set containing only key frames to the cloud server for recognition, without needing to upload the complete long video. This reduces bandwidth usage and overhead under limited communication bandwidth conditions, improves the transmission efficiency of video resources, and ultimately enhances the recognition efficiency of long videos.

[0144] This application also provides a long video semantic analysis device based on the aforementioned cloud-edge collaborative approach. The meanings of the terms used are the same as in the aforementioned cloud-edge collaborative long video semantic analysis method; for specific implementation details, please refer to the description in the method embodiments.

[0145] The cloud-edge collaborative long video semantic analysis device is integrated into the computer equipment of this application. This computer equipment is a cloud server within the long video semantic analysis system, which also includes an edge device. The cloud-edge collaborative long video semantic analysis device may include a receiving unit and an inference unit.

[0146] The receiving unit is used to receive the original question sent by the terminal, break down the original question into multiple sub-questions, determine the importance score corresponding to each sub-question, and send the multiple sub-questions and the importance score corresponding to each sub-question to the edge end. The edge device receives multiple sub-problems and their corresponding importance scores, and receives long video files uploaded by the terminal. It extracts video frames from the long video files according to a preset frame rate to obtain a candidate video frame set. Combining the importance scores of each sub-problem, it determines the number of upload frames for each sub-problem. For each sub-problem, it selects the target video frames corresponding to the number of upload frames for each sub-problem from the candidate video frame set based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set. It then generates a target video frame set by combining the target video frames corresponding to the number of upload frames for each sub-problem and uploads the target video frame set to the cloud server. The inference unit is used to perform inference based on the target video frame set, the original question, and multiple sub-questions using a preset multimodal inference model to obtain the target inference result, and then return the target inference result to the terminal.

[0147] In some embodiments, the receiving unit is further configured to: Determine the logical order of events among multiple subproblems, and construct a directed acyclic graph (DAG) for the multiple subproblems according to the logical order of events. The DAG represents the temporal relationship between the multiple subproblems. Based on the directed acyclic graph, an importance score for each subproblem relative to the original problem is generated.

[0148] The specific implementation of each of the above units can be found in the previous embodiments, and will not be repeated here.

[0149] Figure 6 To implement the structural block diagram of a terminal in this embodiment of the application, the terminal 110 includes: a radio frequency (RF) circuit 510, a memory 515, an input unit 530, a display unit 540, a sensor 550, an audio circuit 560, a wireless fidelity (WiFi) module 570, a processor 580, and a power supply 590, among other components. Those skilled in the art will understand that the terminal 110 structure shown in the figures does not constitute a limitation on a mobile phone or computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0150] The RF circuit 510 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 580; in addition, it transmits uplink data to the base station.

[0151] The memory 515 can be used to store software programs and modules. The processor 580 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 515.

[0152] The input unit 530 can be used to receive input numeric or character information, and to generate key signal inputs related to the terminal's settings and function control. Specifically, the input unit 530 may include a touch panel 531 and other input devices 532.

[0153] The display unit 540 can be used to display input or provided information, as well as various menus of the terminal. The display unit 540 may include a display panel 541.

[0154] Audio circuit 560, speaker 561, and microphone 562 provide an audio interface.

[0155] In this embodiment, the processor 580 included in the terminal 110 can execute the long video semantic analysis method based on cloud-edge collaboration in the previous embodiment.

[0156] The terminal 110 in this application embodiment includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. This application embodiment can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0157] Figure 7 This is a partial structural block diagram of a server implementing this embodiment of the application. The server may refer to... Figure 1 The edge device 120 and cloud server 130 may differ significantly due to variations in configuration or performance, but they share similar components. Taking edge device 120 as an example, it may include one or more central processing units (CPUs) 622 (e.g., one or more processors) and memory 632, and one or more storage media 620 (e.g., one or more mass storage devices) for storing application programs 642 or data 644. The memory 632 and storage media 620 may be temporary or persistent storage. The program stored in storage media 620 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on edge device 120. Furthermore, the CPU 622 may be configured to communicate with storage media 620 and execute the series of instruction operations on storage media 620 on edge device 120.

[0158] Edge device 120 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0159] The central processing unit 622 in the edge terminal 120 can be used to execute the long video semantic analysis method based on cloud-edge collaboration according to the embodiments of this application.

[0160] This application also provides a computer-readable storage medium for storing program code, which is used to execute the cloud-edge collaborative long video semantic analysis method of the foregoing embodiments.

[0161] This application also provides a computer program product, which includes a computer program. The processor of a computer device reads and executes the computer program, causing the computer device to perform the aforementioned cloud-edge collaborative long video semantic analysis method.

[0162] Furthermore, the terms “comprising” and “including”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product or device.

[0163] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0164] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0165] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed between them may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0166] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0167] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0168] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0169] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.

[0170] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0171] The above is a detailed description of the embodiments of this application. However, this application is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A long video semantic analysis method based on cloud-edge collaboration, characterized in that, The method, applied to a long-video semantic analysis system, which includes at least a cloud server and an edge device, comprises: The cloud server receives the original question sent by the terminal, breaks down the original question into multiple sub-questions, determines the importance score corresponding to each sub-question, and sends the multiple sub-questions and their corresponding importance scores to the edge terminal. Specifically, determining the importance score for each sub-question includes: determining the logical order of events among the multiple sub-questions, constructing a directed acyclic graph (DAG) for the multiple sub-questions according to the logical order of events, whereby the DAG represents the temporal relationship between the multiple sub-questions; and generating an importance score for each sub-question relative to the original question based on the DAG. The edge device receives the multiple sub-questions and the importance score corresponding to each sub-question, and receives the long video file uploaded by the terminal, and extracts video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. The edge terminal combines the importance score corresponding to each sub-problem to determine the number of upload frames corresponding to each sub-problem; wherein, the step of combining the importance score corresponding to each sub-problem to determine the number of upload frames corresponding to each sub-problem includes: obtaining a preset total number of upload frames and a preset minimum number of frames for each sub-problem; determining the number of problems corresponding to multiple sub-problems; and combining the preset total number of upload frames, the preset minimum number of frames, the number of problems, and the importance score corresponding to each sub-problem to determine the number of upload frames corresponding to each sub-problem respectively. For each sub-problem, the edge terminal combines the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set to select the target video frame corresponding to the number of uploaded frames for each sub-problem from the candidate video frame set; The edge device generates a target video frame set by combining the target video frames corresponding to the number of uploaded frames for each sub-problem, and uploads the target video frame set to the cloud server; The cloud server uses a preset multimodal reasoning model to reason based on the target video frame set, the original question, and the multiple sub-questions to obtain the target reasoning result, and then returns the target reasoning result to the terminal.

2. The method according to claim 1, characterized in that, The process involves combining the preset total number of uploaded frames, the preset minimum number of frames, the number of questions, and the importance score corresponding to each sub-question to determine the number of uploaded frames for each sub-question, including: The initial total number of frames occupied for the multiple sub-questions is determined by combining the number of questions and the preset minimum number of frames, and the remaining number of allocated frames is determined based on the difference between the preset total number of uploaded frames and the initial total number of frames occupied. When the remaining number of allocated frames is greater than or equal to the preset frame number threshold, the allocation weight corresponding to each sub-problem is determined by combining the importance score corresponding to each sub-problem. The remaining number of allocated frames is then allocated to the multiple sub-problems to obtain the redistributed frame number for each sub-problem, and the uploaded frame number corresponding to each sub-problem is obtained by combining the redistributed frame number for each sub-problem and the preset minimum frame number. When the remaining number of allocated frames is less than the preset frame number threshold, the score ranking relationship between the multiple sub-questions is determined by combining the importance score corresponding to each sub-question, and at least one sub-question ranked first in the score ranking relationship is selected from the preset minimum frame number for frame allocation to obtain the number of upload frames corresponding to each sub-question.

3. The method according to claim 2, characterized in that, The step of allocating the remaining allocation frames to the multiple sub-problems by combining the allocation weights corresponding to each sub-problem to obtain the redistribution frame number for each sub-problem includes: Based on the allocation weight corresponding to each sub-problem and the remaining allocation frames, the initial redistribution value of each sub-problem is determined, and the initial redistribution value of each sub-problem is rounded down to obtain the candidate redistribution frame number of each sub-problem. Calculate the total number of candidate reassignment frames for the multiple sub-problems by combining the number of candidate reassignment frames for each sub-problem, and determine the remaining amount of reassignment between the remaining number of reassignment frames and the total number of candidate reassignment frames; By combining the importance score corresponding to each sub-problem, the score ranking relationship among the multiple sub-problems is determined, and the remaining redistribution amount is sequentially allocated to at least one sub-problem ranked first in the score ranking relationship to obtain the number of secondary redistribution frames for each sub-problem. For each sub-problem, the number of reassignment frames for each sub-problem is determined by combining the corresponding candidate reassignment frame number and the secondary reassignment frame number.

4. The method according to claim 1, characterized in that, For each sub-problem, the step of selecting the target video frame corresponding to the number of uploaded frames for each sub-problem from the candidate video frame set, based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set, includes: Each candidate video frame in the candidate video frame set is encoded to obtain the corresponding frame feature matrix, and each sub-problem is encoded to obtain the corresponding sub-problem feature matrix. For each sub-problem, the feature matrix of the corresponding sub-problem is compared with the feature matrix of each frame to obtain the feature similarity between the current sub-problem and each candidate video frame; For each sub-problem, candidate video frames with the corresponding number of uploaded frames are selected from the candidate video frame set in descending order of feature similarity to obtain the target video frame for each sub-problem.

5. The method according to claim 4, characterized in that, For each sub-problem, candidate video frames with the corresponding number of uploaded frames are selected from the candidate video frame set in descending order of feature similarity to obtain the target video frame for each sub-problem, including: For each sub-problem, candidate video frames with the corresponding number of uploaded frames are selected from the candidate video frame set in descending order of the feature similarity to obtain the initial video frame set for each sub-problem. By combining the initial video frame set of each sub-problem, deduplication is performed to obtain the candidate video frame set corresponding to each sub-problem; Determine the number of video frames in the candidate video frame set corresponding to each sub-problem, and determine the number of supplementary frames for each sub-problem based on the number of video frames corresponding to each sub-problem and the corresponding number of uploaded frames; Based on the number of supplementary frames for each sub-problem, supplementary video frames corresponding to each sub-problem are selected from the candidate video frame set; The supplementary video frames corresponding to each sub-problem are added to the corresponding set of candidate video frames to obtain the target video frames corresponding to each sub-problem.

6. A long video semantic analysis device based on cloud-edge collaboration, characterized in that, An edge device applied to a long-video semantic analysis system, the long-video semantic analysis system further including a cloud server, the device comprising: The receiving unit is used to receive multiple sub-problems and the importance score corresponding to each sub-problem, as well as a long video file uploaded by the receiving terminal, and to extract video frames from the long video file according to a preset frame rate to obtain a candidate video frame set. The multiple sub-problems are obtained by the cloud server from the original problem sent by the terminal, and each importance score is generated by the cloud server for the corresponding sub-problem. The process of generating the importance score is as follows: determining the logical order of events among the multiple sub-problems, and constructing a directed acyclic graph (DAG) for the multiple sub-problems according to the logical order of events, wherein the DAG represents the temporal relationship between the multiple sub-problems; and generating an importance score for each sub-problem relative to the original problem based on the DAG. The determining unit is used to determine the number of upload frames corresponding to each sub-problem by combining the importance score corresponding to each sub-problem. The determination of the number of upload frames corresponding to each sub-problem by combining the importance score corresponding to each sub-problem includes: obtaining a preset total number of upload frames and a preset minimum number of frames for each sub-problem; determining the number of problems corresponding to multiple sub-problems; and determining the number of upload frames corresponding to each sub-problem by combining the preset total number of upload frames, the preset minimum number of frames, the number of problems, and the importance score corresponding to each sub-problem. The selection unit is used to select the target video frame corresponding to the number of uploaded frames for each sub-problem from the candidate video frame set, based on the feature similarity between the sub-problem and each candidate video frame in the candidate video frame set. The upload unit is used to generate a target video frame set by combining the target video frames corresponding to the upload frame number of each sub-problem, and upload the target video frame set to the cloud server. The cloud server then uses a preset multimodal inference model to infer the target inference result based on the target video frame set, the original problem, and the multiple sub-problems, and returns the target inference result to the terminal.

7. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the long video semantic analysis method based on cloud-edge collaboration as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions adapted for loading by a processor to execute the long video semantic analysis method based on cloud-edge collaboration as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Long video understanding method capable of relieving time sequence illusion in video language large model

    CN121392714A

  • Long video question and answer enhancement processing method and system based on bidirectional audio visual alignment

    CN122090839A