A comprehensive analysis and processing method for audio and video system data based on virtual distributed architecture

Through the comprehensive analysis and processing method of audio and video system data based on a virtual distributed architecture, the efficiency and flexibility problems of the traditional centralized architecture in processing heterogeneous audio and video data are solved, efficient and accurate comprehensive analysis and processing of audio and video data are achieved, and system performance and stability are optimized.

CN120475226BActive Publication Date: 2025-09-09SHENZHEN ZIDOO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510979410.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-09
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Traditional audio and video data processing methods are based on a centralized architecture and are unable to effectively cope with sequences of encoded data blocks encapsulated by different transmission protocols. This results in limited data processing efficiency and flexibility, and computing and storage resources are unable to meet real-time processing requirements, affecting system performance.

Method used

A comprehensive analysis and processing method for audio and video system data based on a virtual distributed architecture is adopted. By deploying a multi-source audio and video acquisition component cluster in a virtualized resource pool, distributed task decomposition and feature extraction of heterogeneous audio and video data sets are performed. The real-time resource availability parameters of virtual computing nodes are used for dynamic allocation, and data fusion and resource regulation are performed through cross-node communication interfaces.

Benefits of technology

It achieves efficient and accurate comprehensive analysis and processing of audio and video data, improves the pertinence and efficiency of task processing, fully utilizes computing resources, optimizes system performance, and improves system stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475226B_ABST
    Figure CN120475226B_ABST
Patent Text Reader

Abstract

The present invention provides a method for comprehensive analysis and processing of audio and video system data based on a virtual distributed architecture. The method collects a heterogeneous audio and video data set of a target area through a cluster of multi-source audio and video acquisition components deployed in a virtualized resource pool, performs distributed task decomposition according to preset spatiotemporal correlation constraints, and generates a feature extraction task set. Subsequently, tasks are dynamically allocated to a processing unit group based on real-time resource availability parameters of virtual computing nodes, triggering a parallel feature extraction process. Then, intermediate feature data sets are fused through a cross-node communication interface to generate a global state description vector of the target area. Finally, a resource control instruction set is generated according to a mapping relationship between the global state description vector and a preset system configuration policy, and is transmitted to a scheduling management node of the virtualized resource pool, thereby achieving efficient and accurate comprehensive analysis and processing of audio and video data and improving the overall performance and stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio and video processing, and in particular to a method for comprehensive analysis and processing of audio and video system data based on a virtual distributed architecture. Background Art

[0002] In the field of audio and video systems, with the continuous development of technology and the increasing complexity of application scenarios, the demand for processing and analyzing audio and video data has become increasingly diverse and refined. Traditional audio and video data processing methods are often based on a centralized architecture, which faces many challenges when processing large-scale, heterogeneous audio and video data. For example, the centralized architecture cannot effectively handle sequences of encoded data blocks encapsulated by different transmission protocols, resulting in limited efficiency and flexibility in data processing. At the same time, as the amount of data continues to grow, the computing and storage resources of the centralized architecture cannot meet the needs of real-time processing, which can easily cause data processing delays and affect the overall performance of the system. In addition, traditional audio and video data processing methods lack effective mechanisms in terms of task decomposition, resource allocation, feature extraction, and data fusion, making it difficult to achieve efficient and accurate comprehensive analysis and processing of audio and video data. Summary of the Invention

[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for comprehensive analysis and processing of audio and video system data based on a virtual distributed architecture, the method comprising:

[0004] Collecting a heterogeneous audio and video data set of the target area through a multi-source audio and video acquisition component cluster deployed in a virtualized resource pool, wherein the heterogeneous audio and video data set includes a sequence of coded data blocks encapsulated by different transmission protocols;

[0005] Performing a distributed task decomposition operation on the heterogeneous audio and video data set according to preset spatiotemporal correlation constraints to generate a feature extraction task set matching different analysis dimensions;

[0006] Dynamically assigning the feature extraction task set to corresponding processing unit groups based on real-time resource availability parameters of virtual computing nodes, and triggering each processing unit to execute a parallel feature extraction process corresponding to the assigned feature extraction task;

[0007] Performing a fusion operation on the intermediate feature data sets generated by the parallel feature extraction process through a cross-node communication interface to generate a global state description vector of the target area;

[0008] A resource control instruction set is generated according to a mapping relationship between the global state description vector and a preset system configuration policy, and the resource control instruction set is transmitted to a scheduling management node of a virtualized resource pool.

[0009] On the other hand, an embodiment of the present invention also provides an audio and video system data comprehensive analysis and processing system based on a virtual distributed architecture, including a processor and a machine-readable storage medium, the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0010] Based on the above aspects, the embodiments of the present invention, through a cluster of multi-source audio and video acquisition components deployed in a virtualized resource pool, can efficiently acquire heterogeneous audio and video data sets in a target area, effectively handle sequences of encoded data blocks encapsulated by different transmission protocols, perform distributed task decomposition operations on the heterogeneous audio and video data sets according to preset spatiotemporal correlation constraints, generate feature extraction task sets matching different analysis dimensions, achieve refined task division, and improve the pertinence and efficiency of task processing. The feature extraction task sets are dynamically allocated to corresponding processing unit groups based on real-time resource availability parameters of virtual computing nodes, triggering parallel feature extraction processes, fully utilizing computing resources, achieving dynamic optimization of resource allocation, and improving overall system performance. The intermediate feature data sets generated by the parallel feature extraction processes are fused through a cross-node communication interface to generate a global state description vector for the target area, achieving deep data fusion and feature extraction. A resource control instruction set is generated based on the mapping relationship between the global state description vector and a preset system configuration policy, and the resource control instruction set is transmitted to the scheduling management node of the virtualized resource pool, achieving intelligent control and optimization of the system, and improving the stability and reliability of the system. Thus, efficient and accurate comprehensive analysis and processing of audio and video data is achieved through a virtual distributed architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a schematic diagram of the execution flow of the method for comprehensive analysis and processing of audio and video system data based on a virtual distributed architecture provided by an embodiment of the present invention.

[0012] Figure 2 It is a schematic diagram of exemplary hardware and software components of an audio and video system data comprehensive analysis and processing system based on a virtual distributed architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0013] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a method for comprehensive analysis and processing of audio and video system data based on a virtual distributed architecture provided by an embodiment of the present invention. The method for comprehensive analysis and processing of audio and video system data based on a virtual distributed architecture is introduced in detail below.

[0014] Step S110: collecting a heterogeneous audio and video data set of a target area through a multi-source audio and video collection component cluster deployed in a virtualized resource pool, wherein the heterogeneous audio and video data set includes a sequence of coded data blocks encapsulated by different transmission protocols.

[0015] In the actual scenario of this embodiment, the target area is a specific spatial range that requires audio and video monitoring and analysis. A virtualized resource pool is a collection of multiple physical resources integrated through virtualization technology, within which a cluster of multi-source audio and video acquisition components is deployed. These acquisition components, with different functions and features, are dispersed throughout the target area to collect audio and video data from all angles and perspectives.

[0016] The multi-source audio and video capture component cluster includes various types of capture devices, such as cameras and microphones of varying models. Because these devices may come from different manufacturers and utilize different technical standards and transmission protocols, the captured data is heterogeneous. For example, some cameras may use the H.264 encoding format and transmit data via the RTSP protocol, while others may use the AAC encoding format and transmit audio data via UDP. These sequences of encoded data blocks encapsulated by different transmission protocols together constitute a heterogeneous audio and video data set.

[0017] Step S111: Establish a bandwidth-adjustable data transmission channel for different audio and video acquisition components deployed in the virtualized resource pool through a programmable network interface, and integrate a real-time quality control unit in the data transmission channel to dynamically adjust data sampling rate parameters, compression ratio parameters and encoding format parameters.

[0018] To ensure efficient and stable transmission of heterogeneous audio and video data to subsequent processing nodes, dedicated data transmission channels must be established for each audio and video acquisition component. A programmable network interface allows flexible configuration of data transmission channel parameters based on actual needs. This interface allows for the allocation of different bandwidths to different acquisition components, and the bandwidth can be dynamically adjusted based on network conditions and data transmission requirements.

[0019] A real-time quality control unit is integrated into the data transmission channel to ensure data quality. This unit continuously monitors various metrics during data transmission and dynamically adjusts the data sampling rate, compression ratio, and encoding format parameters based on these metrics. For example, when network bandwidth is sufficient, the real-time quality control unit can appropriately increase the data sampling rate to obtain higher-quality audio and video data. When network bandwidth is limited, it can reduce the data sampling rate and adjust the compression ratio and encoding format to reduce data volume and ensure smooth data transmission.

[0020] Step S1111: monitor the current bandwidth usage indicator, end-to-end delay indicator, and data packet loss rate indicator of the data transmission channel in real time.

[0021] The real-time quality control unit first monitors the current bandwidth utilization of the data transmission channel. Bandwidth utilization refers to the ratio of network bandwidth currently occupied by data transmission to the total available bandwidth. By obtaining this metric in real time, we can understand network bandwidth usage and determine whether there is a bandwidth shortage. For example, if bandwidth utilization approaches 100%, it indicates that network bandwidth is very limited and measures may be needed to reduce data volume.

[0022] End-to-end latency refers to the time it takes for data to travel from the acquisition component to the receiving end. This metric is crucial for real-time audio and video data transmission. Excessive end-to-end latency can cause issues like stuttering and delays in audio and video playback. The real-time quality control unit continuously monitors this metric to promptly detect latency anomalies and take appropriate adjustments.

[0023] The packet loss rate is the ratio of the number of packets lost during data transmission to the total number of packets sent. Packet loss can cause incomplete audio and video data, affecting data quality. The real-time quality control unit closely monitors this metric. If the packet loss rate is too high, it attempts to resolve the issue by adjusting the transmission strategy or retransmitting lost packets.

[0024] Step S1112: establishing a dynamic adjustment model with data fidelity index, transmission efficiency index and energy consumption index as optimization targets.

[0025] In order to comprehensively consider factors such as data quality, transmission efficiency, and energy consumption, a dynamic adjustment model needs to be established. This dynamic adjustment model takes data fidelity index, transmission efficiency index, and energy consumption index as optimization targets.

[0026] Data fidelity measures how closely the collected audio and video data resembles the original data after processing and transmission. Data fidelity can be quantified by calculating metrics such as the peak signal-to-noise ratio (PSNR). Transmission efficiency focuses on data transmission speed and bandwidth utilization, expressed as the amount of data transmitted per unit time. Energy consumption considers the energy consumed by acquisition components and transmission equipment during data collection and transmission.

[0027] The dynamic adjustment model automatically adjusts data sampling rate parameters, compression ratio parameters, and encoding format parameters based on real-time monitoring of metrics such as bandwidth utilization, end-to-end latency, and packet loss rate, as well as preset optimization goals. For example, when data fidelity is critical, the model minimizes the compression ratio to ensure data integrity. When transmission efficiency is paramount, the model increases the compression ratio to reduce data volume and increase transmission speed. The model also considers energy consumption, minimizing device energy consumption while maintaining data quality and transmission efficiency.

[0028] Step S1113: performing a real-time update operation on the weight parameters of the dynamic adjustment model through an online learning algorithm.

[0029] An online learning algorithm is one that continuously updates model parameters based on new data. In this embodiment, an online learning algorithm is used to update the weight parameters of the dynamic adjustment model in real time. The weight parameters determine the importance of data fidelity, transmission efficiency, and energy consumption in the model optimization process.

[0030] For example, in the initial stages, each metric might be assigned a set weight based on experience. However, as data transmission progresses, the online learning algorithm dynamically adjusts these weights based on real-time monitoring of various metrics and actual optimization results. If the algorithm finds that the current data fidelity is low while the transmission efficiency is high, the algorithm may increase the weight of the data fidelity metric to improve data quality.

[0031] Step S1114: Calculate candidate parameter combinations based on the updated dynamic adjustment model, select the parameter combination that meets the data fidelity threshold and has the best transmission efficiency through the closed-loop feedback control algorithm, and generate data sampling rate adjustment parameters, compression ratio parameter combinations and encoding format switching strategies.

[0032] Based on the updated dynamic adjustment model, a series of candidate parameter combinations can be calculated, including different data sampling rates, compression ratios, and encoding formats. The closed-loop feedback control algorithm will evaluate and screen these candidate parameter combinations.

[0033] First, each candidate parameter combination is checked to see if it meets the data fidelity threshold. The data fidelity threshold is a pre-defined standard used to ensure that the transmitted data quality reaches a certain level. If the data fidelity generated by a candidate parameter combination falls below the threshold, the combination is eliminated.

[0034] Then, from the candidate parameter combinations that meet the data fidelity threshold, the combination with the best transmission efficiency is selected. Transmission efficiency can be measured by calculating metrics such as the amount of data transmitted per unit time or bandwidth utilization. Finally, based on the selected optimal parameter combination, the data sampling rate adjustment parameters, compression ratio parameter combinations, and encoding format switching strategies are generated.

[0035] Step S1115: Send the data sampling rate adjustment parameter, compression ratio parameter combination and encoding format switching strategy to the data processing module of the audio and video acquisition component to perform a configuration update operation.

[0036] The generated data sampling rate adjustment parameters, compression ratio parameter combinations, and encoding format switching strategies are sent to the data processing module of the audio and video acquisition component. The data processing module is the part of the acquisition component responsible for processing and encoding the collected raw data.

[0037] Upon receiving these configuration update policies, the data processing module adjusts its own configuration based on the parameters and policies. For example, if the received data sampling rate adjustment parameters require a higher sampling rate, the data processing module will increase the sampling frequency accordingly; if the encoding format switching policy requires switching from H.264 to H.265, the data processing module will update the encoding algorithm to implement the encoding format switch.

[0038] Step S1116: Feedback the adjusted data quality indicators to the dynamic adjustment model in real time for iterative optimization.

[0039] After the data processing module of the audio and video acquisition component performs the configuration update operation, it is necessary to feed back the adjusted data quality indicators to the dynamic adjustment model in real time. The data quality indicators include indicators in terms of data fidelity, transmission efficiency, and energy consumption.

[0040] The dynamic adjustment model evaluates and adjusts based on the feedback from data quality indicators. If the adjusted data quality indicators do not achieve the expected optimization results, the online learning algorithm updates the weight parameters again, recalculates the candidate parameter combinations, and selects the optimal parameter combination for configuration update. This iterative optimization method continuously improves the quality and efficiency of data transmission.

[0041] Step S112: performing an alignment and caching operation on the coded data blocks transmitted by each data transmission channel according to a time reference to form the heterogeneous audio and video data set with temporal and spatial consistency.

[0042] Because different audio and video acquisition components may have different clock synchronization accuracy and data acquisition rates, the collected coded data blocks may have different timings. To form a heterogeneous audio and video data set with temporal and spatial consistency, the coded data blocks transmitted by each data transmission channel need to be aligned and cached according to a time reference.

[0043] First, a unified time base is established for each data transmission channel. This can be achieved by synchronizing the clocks of various acquisition components through methods such as the Network Time Protocol (NTP). Each encoded data block is then timestamped based on this time base.

[0044] Next, the timestamped encoded data blocks are cached in a temporary storage area. During the caching process, the data blocks are sorted and aligned according to their timestamps, ensuring that data blocks collected at the same time are placed together. Ultimately, these aligned cached encoded data blocks form a heterogeneous audio and video data set with temporal and spatial consistency.

[0045] Step S120: performing a distributed task decomposition operation on the heterogeneous audio and video data set according to preset spatiotemporal correlation constraints to generate a feature extraction task set matching different analysis dimensions.

[0046] After acquiring a collection of heterogeneous audio and video data with spatiotemporal consistency, further processing and analysis are required. The preset spatiotemporal correlation constraints are pre-defined rules based on the characteristics of the target area and the analysis requirements, which are used to guide the distributed task decomposition of the data.

[0047] The purpose of distributed task decomposition is to break down large, heterogeneous audio and video data sets into multiple smaller tasks for parallel processing and improved efficiency. These smaller tasks are categorized and organized according to different analysis dimensions, forming a set of feature extraction tasks that match these dimensions. For example, analysis dimensions can include time, space, and semantics. Different analysis dimensions correspond to different feature extraction tasks. For example, in the time dimension, the temporal features of audio and video data may need to be extracted, while in the spatial dimension, the position and motion features of objects may need to be extracted.

[0048] Step S121: performing a unified timestamp calibration operation on each coded data block sequence in the heterogeneous audio and video data set to obtain a calibrated coded data block sequence.

[0049] Since different audio and video acquisition components may have errors in clock synchronization, resulting in inconsistent timestamps in the acquired coded data block sequences, it is necessary to perform a unified timestamp calibration operation on each coded data block sequence to ensure the accuracy of subsequent analysis.

[0050] Uniform timestamp calibration can be achieved by following the following steps. First, a reference time base is selected, such as the standard time provided by the Network Time Protocol (NTP). Then, for each coded data block sequence, the time difference between it and the reference time base is calculated. Based on this time difference, the timestamp of each coded data block is adjusted to align with the reference time base.

[0051] For example, assuming that the timestamp of a certain coded data block sequence is t seconds faster than the reference time base, then the timestamp of each coded data block in the sequence is subtracted by t seconds to obtain a calibrated coded data block sequence.

[0052] Step S122: performing a segmentation operation on the calibrated coded data block sequence based on a sliding time window mechanism to generate a set of data segments with temporal continuity.

[0053] The sliding time window mechanism is a commonly used method for processing time series data. In this embodiment, the sliding time window mechanism is used to perform segmented interception operations on the calibrated coded data block sequence.

[0054] First, define the size and step size of a sliding time window. The size of the sliding time window determines the duration of each data segment, and the step size determines the time interval between each sliding of the window.

[0055] Then, starting from the start of the calibrated coded data block sequence, a sliding time window is applied to the sequence. Each time, a coded data block within the window is intercepted to form a data segment. As the window slides, new data segments are continuously intercepted until the entire coded data block sequence is traversed.

[0056] For example, assume the sliding time window is T seconds long and the sliding step is t seconds. Starting from the 0th second of the coded data block sequence, the coded data blocks from 0th to Tth seconds are intercepted to form the first data segment. Then, the window is slided by t seconds, and the coded data blocks from tth to T+tth seconds are intercepted to form the second data segment, and so on. Ultimately, a set of data segments with temporal continuity is generated.

[0057] Step S123: performing a spatial coordinate system conversion operation on the multiple data streams in each data segment in the data segment set, and establishing a spatial alignment mapping relationship model across the data streams.

[0058] In the target area, different audio and video capture components may be installed at different positions and angles, resulting in spatial differences in the collected multi-channel data streams. To effectively analyze and fuse these multi-channel data streams, it is necessary to perform spatial coordinate system transformation operations on the multi-channel data streams within each data segment and establish a spatial alignment mapping relationship model across the data streams.

[0059] First, establish a unified spatial coordinate system as a reference coordinate system. This can be done by selecting a fixed point in the target area as the origin to establish a three-dimensional spatial coordinate system. Then, for each audio and video capture component, determine its position and orientation within the reference coordinate system.

[0060] Based on the position and orientation of the acquisition components, the multiple data streams within each data segment are transformed from their respective local coordinate systems to the reference coordinate system. This transformation allows the data streams collected by different acquisition components to be spatially aligned.

[0061] During the conversion process, a spatial alignment mapping model is established across the data streams. This spatial alignment mapping model describes the position and orientation relationships of different data streams in the reference coordinate system, and how to align them. For example, this spatial alignment mapping model can be used to match the position of objects in video data captured by one camera with the position of objects in video data captured by another camera.

[0062] Step S124: performing a spatial overlap region detection operation on the multiple data streams in the data segment according to the spatial alignment mapping relationship model, and identifying a target sub-region set with spatial consistency.

[0063] After establishing a spatial alignment mapping relationship model across data streams, the spatial alignment mapping relationship model is used to perform a spatial overlap region detection operation on multiple data streams within a data segment. Spatial overlap regions refer to regions that are spatially overlapped by different data streams.

[0064] By analyzing the position and range of multiple data streams in the reference coordinate system, the spatial overlap area between them can be determined. During the detection process, all multiple data streams within each data segment are traversed and compared.

[0065] For detected spatially overlapping regions, the data features and content are further analyzed to identify a set of spatially consistent target subregions. A spatially consistent target subregion is one where the information collected by different data streams is similar and relevant. For example, in a surveillance scenario, different cameras may capture the same object or event within spatially overlapping regions. These regions are considered spatially consistent target subregions.

[0066] Step S125: performing a spatiotemporal association binding operation on the target sub-region set and the corresponding data segments to generate a feature extraction task set containing multi-dimensional constraint conditions, and sorting the feature extraction task set according to a preset priority rule to generate a final task allocation queue.

[0067] After identifying a set of spatially consistent target sub-regions, these sub-regions are then temporally bound to the corresponding data segments. Temporal binding involves combining the spatial location information of the target sub-regions with the temporal information of the data segments to form a complete spatiotemporal information unit.

[0068] Through spatiotemporal binding operations, a feature extraction task set containing multidimensional constraints is generated. These multidimensional constraints include temporal constraints, spatial constraints, and semantic constraints. For example, a feature extraction task may require feature extraction from data within a specific spatial region within a specific time range.

[0069] The feature extraction task set is then sorted according to pre-set priority rules. Priority rules can be determined based on factors such as task importance, urgency, and resource requirements. For example, tasks with high real-time requirements can be given a higher priority; tasks with high resource consumption can be prioritized based on resource availability.

[0070] Finally, after the sorting operation, the final task allocation queue is generated, which facilitates the subsequent allocation of tasks to different processing units for processing.

[0071] Step S130: dynamically allocating the feature extraction task set to corresponding processing unit groups based on the real-time resource availability parameters of the virtual computing nodes, and triggering each processing unit to execute a parallel feature extraction process corresponding to the allocated feature extraction task.

[0072] After generating the final task allocation queue, the feature extraction task set needs to be dynamically assigned to the corresponding processing unit group. Virtual compute nodes are computing resources within the virtualized resource pool. Each virtual compute node has a certain amount of computing power and resources. Real-time resource availability parameters reflect the current resource usage of each virtual compute node, such as CPU utilization, memory utilization, and disk I / O.

[0073] Based on these real-time resource availability parameters, tasks in the feature extraction task set are dynamically assigned to appropriate processing unit groups. A processing unit group is a collection of multiple processing units, each of which can independently perform feature extraction tasks. By executing these tasks in parallel, processing efficiency can be improved.

[0074] Step S131: collecting real-time resource availability parameters of each computing node in the virtualized resource pool in real time.

[0075] In order to accurately assign feature extraction tasks to appropriate computing nodes, it is necessary to collect real-time resource availability parameters of each computing node in the virtualized resource pool. This can be achieved by installing a monitoring agent on each computing node.

[0076] The monitoring agent regularly collects various resource usage information of computing nodes, such as CPU usage, memory usage, disk I / O rate, etc. These real-time resource availability parameters are sent to a centralized management node for unified processing and analysis.

[0077] For example, at regular intervals (e.g., 1 minute), the monitoring agent collects the CPU usage of the compute nodes and sends it to the management node. The management node stores these real-time resource availability parameters in a database for subsequent task allocation decisions.

[0078] Step S132: Constructing a dynamic scheduling decision model with task processing delay threshold, multi-dimensional resource utilization balance and load difference minimization as constraints.

[0079] In order to achieve a reasonable allocation of feature extraction tasks, a dynamic scheduling decision model needs to be built. The model takes task processing delay threshold, multi-dimensional resource utilization balance, and load difference minimization as constraints.

[0080] The task processing delay threshold is the maximum processing time allowed for each feature extraction task. During task allocation, it is necessary to ensure that the processing time of each task does not exceed this threshold to ensure the real-time performance of the task.

[0081] Multi-dimensional resource utilization balancing means that when assigning tasks, the usage of multiple resource dimensions (such as CPU, memory, and disk) is considered to ensure balanced resource utilization across all compute nodes. This prevents overutilization of some compute nodes while leaving others idle.

[0082] Minimizing load differences means minimizing the load differences between computing nodes so that the amount of tasks undertaken by each computing node is relatively balanced.

[0083] The dynamic scheduling decision model will make optimal decisions on task allocation based on these constraints and the resource availability parameters of the computing nodes collected in real time.

[0084] Step S133: Convert each feature extraction task in the feature extraction task set into a task descriptor object.

[0085] To facilitate the management and allocation of feature extraction tasks, each feature extraction task in the feature extraction task set is converted into a task descriptor object. A task descriptor object is an object that contains detailed information about the task, which records various attributes and parameters of the task.

[0086] The information contained in a task descriptor object can include the task name, priority, input data, required computing resources, and expected processing time. For example, for a task to extract object motion features from audio and video data, the task descriptor object would record the task name as "Object Motion Feature Extraction," the priority information set according to preset rules, the input data as the corresponding audio and video data segment, the required computing resources such as a specific number of CPU cores and a certain amount of memory, and the expected processing time predicted based on experience or models. By converting each feature extraction task into a task descriptor object, task management and scheduling become more convenient and efficient.

[0087] Step S134: performing task matching calculation on the task descriptor object and the real-time resource availability parameters of the computing node through the dynamic scheduling decision model to generate a task allocation priority list.

[0088] The dynamic scheduling decision model considers all information in the task descriptor object and the real-time resource availability parameters of the compute nodes to perform task matching calculations. First, the model selects compute nodes that meet the task's resource requirements based on the real-time resource availability parameters. For example, if a task requires a specific number of CPU cores and a certain amount of memory, the model searches for compute nodes with the required number of CPU cores and memory.

[0089] The model then evaluates the selected compute nodes based on constraints such as task processing latency thresholds, multi-dimensional resource utilization balancing, and minimizing load variance. For each compute node, the model calculates its current load, resource utilization, and other metrics. Furthermore, the model assesses the node's suitability for the task based on the task's priority and expected processing time.

[0090] During the calculation process, different metrics are weighted. For example, task priority may be assigned a higher weight to ensure that high-priority tasks are processed first. Metrics such as multi-dimensional resource utilization balancing and load variability minimization are also weighted according to their importance. This weighted calculation yields a matching score for each task on each compute node.

[0091] Finally, tasks are sorted according to their matching scores to generate a task allocation priority list. In this task allocation priority list, tasks with high matching scores are ranked first and assigned to appropriate computing nodes first.

[0092] Step S135: Distribute the feature extraction task to the target computing node according to the task allocation priority list, and activate the parallel processing unit in the target computing node to execute the feature extraction process.

[0093] According to the generated task allocation priority list, the feature extraction tasks are distributed to the target computing nodes in sequence. The target computing nodes are determined based on the task matching calculation results and are computing nodes that can meet the task resource requirements and have a high matching score.

[0094] When a task is dispatched, a task descriptor object is sent to the target compute node. After receiving the task descriptor object, the target compute node prepares for the task based on the information contained in it. For example, it loads the required computing resources and reads the input data for the task.

[0095] The parallel processing units (PPUs) within the target compute node are then activated to execute the feature extraction process. PPUs can be computing resources such as multiple CPU cores and GPUs, enabling them to process tasks simultaneously, improving processing efficiency. During the feature extraction process, each PPU performs feature extraction operations on the input data based on the task's requirements. For example, for audio and video data, operations such as image feature extraction and audio feature extraction might be performed.

[0096] Step S136: monitor the task execution progress of the feature extraction task in real time and dynamically adjust the task allocation strategy according to changes in resource status.

[0097] During the feature extraction task execution process, it is necessary to monitor the task's progress in real time. This can be achieved by installing a monitoring program on the compute node. The monitoring program regularly collects task execution status information, such as the amount of work completed, the amount of work remaining, and the current processing speed.

[0098] At the same time, it is also necessary to monitor the resource status changes of computing nodes in real time. Since the resource usage of computing nodes may change as tasks are executed, for example, the CPU usage of a computing node may increase due to the execution of a task, resulting in a decrease in its resource availability.

[0099] Dynamically adjust task allocation strategies based on task execution progress and resource status changes. If a task's execution progress is slow, perhaps due to insufficient compute node resources or other reasons, the task can be migrated to a compute node with more abundant resources. Alternatively, if resource utilization on a compute node is too low, other tasks can be assigned to it to improve resource utilization.

[0100] Step S140: performing a fusion operation on the intermediate feature data set generated by the parallel feature extraction process through the cross-node communication interface to generate a global state description vector of the target area.

[0101] After each compute node completes the parallel feature extraction process, a series of intermediate feature datasets are generated. These intermediate feature datasets are the result of feature extraction on audio and video data on different compute nodes, reflecting the characteristics of different parts or aspects of the target area. To comprehensively and accurately describe the state of the target area, these intermediate feature datasets need to be fused through a cross-node communication interface.

[0102] The cross-node communication interface is an interface used for data transmission and communication between different computing nodes. It can ensure that the intermediate feature data set can be safely and efficiently transmitted from each computing node to an aggregation service node for fusion processing.

[0103] Step S141: After the target computing node completes the feature extraction process, extract the local feature vector set generated inside the target computing node.

[0104] After the target computing node completes the feature extraction process, it generates a series of local feature vectors within it. These local feature vectors are the result of feature extraction of the input audio and video data, which reflect the local feature information of the target area.

[0105] For example, for an audio or video data segment, we might extract color and texture features of the image, and frequency and volume features of the audio. These features are combined into a local feature vector. Within the target compute node, all local feature vectors are collected to form a local feature vector set.

[0106] Step S142: transmitting the local feature vector set to a predefined aggregation service node via a distributed messaging middleware.

[0107] To transmit the local feature vector sets generated by each target computing node to a predefined aggregation service node for fusion processing, a distributed messaging middleware is used. Distributed messaging middleware is a software system used for message transmission and communication within a distributed system. It provides a reliable message transmission mechanism, ensuring that the local feature vector sets are not lost or damaged during transmission. During transmission, the distributed messaging middleware encapsulates and encodes the local feature vector sets for transmission across the network. It also handles message routing and distribution, ensuring that the local feature vector sets are accurately transmitted to the predefined aggregation service node.

[0108] Step S143: constructing a multi-layer feature processing architecture in the aggregation service node, wherein the multi-layer feature processing architecture includes a temporal correlation analysis processing layer, a spatial topology relationship processing layer, and a semantic abstraction processing layer.

[0109] In the aggregation service node, a multi-layer feature processing architecture needs to be built to fuse the local feature vector set. This multi-layer feature processing architecture consists of three main processing layers: temporal correlation analysis processing layer, spatial topology relationship processing layer, and semantic abstraction processing layer.

[0110] The temporal correlation analysis layer analyzes the temporal relationships between sets of local feature vectors. Because audio and video data exhibit time series characteristics, features at different time points may exhibit certain correlations. This layer performs contextual correlation analysis on local features across consecutive time segments to uncover temporal patterns between these features.

[0111] The spatial topology processing layer analyzes the spatial topological relationships of local feature vector sets. Within the target region, features at different locations may have spatial dependencies. This layer constructs a graph-based feature propagation network model. Through graph embedding, it converts topological features into fixed-dimensional vectors, capturing spatial dependencies across subregions.

[0112] The semantic abstraction processing layer is primarily used to extract high-level semantic information from a set of local feature vectors. This high-level semantic information can more accurately describe the state and meaning of the target region. This processing layer deploys a multi-level feature abstraction network to gradually extract high-level semantic information.

[0113] Step S1431: Adopting an adaptive sliding window mechanism in the temporal association analysis processing layer to perform context association analysis operations on the local features of the continuous time segments.

[0114] In the time series correlation analysis processing layer, an adaptive sliding window mechanism is used to perform contextual correlation analysis on the local features of continuous time segments. The adaptive sliding window mechanism is a mechanism that automatically adjusts the window size based on the characteristics of the data and the analysis requirements.

[0115] First, define an initial sliding window size and sliding step size. Then, starting from the starting position of the local feature vector set, apply the sliding window to the local features of consecutive time segments. Each time, local features within the window are intercepted and the correlation between them is analyzed.

[0116] During the analysis process, the sliding window size is adaptively adjusted based on the changes in local features and the degree of correlation. If local features change dramatically but the degree of correlation is low, the sliding window size may be reduced to more accurately analyze the correlation between features. If local features change more slowly but the degree of correlation is high, the sliding window size may be increased to improve analysis efficiency.

[0117] Through this adaptive sliding window mechanism, the contextual association relationship between local features of consecutive time segments can be more accurately mined.

[0118] Step S1432: Construct a topological structure model based on a graph neural network in the spatial topological relationship processing layer to capture spatial correlation features across sub-regions.

[0119] In the spatial topology relationship processing layer, a topological structure model based on a graph neural network is constructed to capture the spatial correlation characteristics across sub-regions. A graph neural network is a neural network model specifically designed to process graph-structured data.

[0120] First, the set of local feature vectors is represented as a graph structure. The nodes in the graph represent different sub-regions, and the feature vectors of the nodes are the local feature vectors of the sub-regions. The edges in the graph represent the spatial relationship between sub-regions, and the weights of the edges can represent the degree of spatial correlation between sub-regions.

[0121] Then, a graph neural network is used to process the graph structure. The graph neural network transmits and aggregates information between nodes through a message passing mechanism, thereby learning the spatial correlation characteristics between nodes.

[0122] During processing, the graph neural network continuously updates the node feature vectors to better reflect the spatial correlation across sub-regions. Ultimately, a graph embedding operation converts the topological features into fixed-dimensional vectors, facilitating subsequent processing and analysis.

[0123] Step S1433: deploying a multi-level feature processing unit in the semantic abstraction processing layer, wherein the multi-level feature processing unit includes a primary feature filtering module, a mid-level feature fusion module, and a high-level semantic generation module.

[0124] In the semantic abstraction processing layer, a multi-level feature processing unit is deployed to gradually extract high-level semantic information. The multi-level feature processing unit consists of three main modules: a primary feature filtering module, an intermediate feature fusion module, and a high-level semantic generation module.

[0125] The primary feature filtering module performs a weighted filtering operation on the input local feature vector to suppress interference from noisy features. For example, a gated attention mechanism can be used to weight features based on their importance. The gated attention mechanism learns the weight of each feature, giving more weight to important features and less weight to noisy features. In this way, unimportant noisy features are filtered out, improving the accuracy of subsequent processing.

[0126] The primary function of the mid-level feature fusion module is to interactively fuse features from multiple sources. A feature dimension alignment layer precedes it, mapping multi-source features to a unified dimensional space through linear projection to address the mismatch between feature dimensions. The mid-level feature fusion module then uses a cross-attention mechanism to enable interaction and fusion between features from different sources, uncovering potential correlations between them.

[0127] The main function of the high-level semantic generation module is to perform temporal analysis on feature sequences and infer the semantic evolution path. For example, a recursive neural network model can be used to process feature sequences after primary and secondary processing. The recursive neural network model can capture the temporal information in the feature sequence and, by analyzing this information, infer the semantic evolution path, thereby generating high-level semantic information.

[0128] Step S1433-1: In the primary feature filtering module, a gated attention mechanism is used to perform weighted screening operations on the input features to suppress noise feature interference.

[0129] In the primary feature filtering module, a gated attention mechanism is used to perform weighted filtering on the input features. The gated attention mechanism is a mechanism that can automatically learn the importance of features.

[0130] First, the input local feature vector is fed into a gated attention mechanism. The gated attention mechanism calculates the weight of each feature through a neural network model. This neural network model learns the importance of each feature for the final semantic generation based on its content and context.

[0131] The input features are then weighted according to the calculated weights. Important features receive larger weights, while noise features receive smaller weights. This weighting operation suppresses the interference of noise features and highlights the role of important features.

[0132] Finally, the weighted features are filtered to retain only those with larger weights. This can reduce the amount of computation required for subsequent processing while improving processing accuracy.

[0133] Step S1433-2: Deploy a feature dimension alignment layer before the intermediate feature fusion module, map the multi-source features to a unified dimensional space through linear projection, and implement the multi-source feature interactive fusion operation through the cross-attention mechanism in the intermediate feature fusion module.

[0134] Before the mid-level feature fusion module, a feature dimension alignment layer is deployed to address the issue of mismatched feature dimensions from multiple sources. Local feature vectors from different sources may have different dimensions, which can affect subsequent feature fusion operations.

[0135] The Feature Dimension Alignment layer maps multiple source features to a unified dimensional space through linear projection. Linear projection is a linear transformation operation that learns a projection matrix based on the input feature dimensions and the target dimension. By multiplying the input features by the projection matrix, they are mapped to a unified dimensional space.

[0136] In the mid-level feature fusion module, a cross-attention mechanism is used to interactively fuse features from multiple sources. This mechanism allows features from different sources to interact and focus on each other. For example, it calculates attention scores between each feature and other features, and then weights and aggregates features based on these scores. In this way, potential connections between features from different sources are discovered, enabling feature fusion.

[0137] Step S1433-3: Use a recursive neural network model in the high-level semantic generation module to perform temporal analysis on the feature sequence and infer the semantic evolution path.

[0138] In the high-level semantic generation module, a recursive neural network model is used to perform temporal analysis on the feature sequences after primary and secondary processing. The recursive neural network model is a neural network model specifically designed for processing sequence data, which can capture the temporal information in sequence data.

[0139] When a feature sequence is input into a recurrent neural network model, the model updates the current hidden state based on the input features at the current moment and the hidden state at the previous moment. Through continuous iterative updates, the model learns the temporal patterns in the feature sequence.

[0140] Based on the learned temporal patterns, the semantic evolution path is inferred. The semantic evolution path describes the changes and development trends of semantics over time. By inferring the semantic evolution path, high-level semantic information can be generated more accurately.

[0141] Step S1433-4: Introduce a learnable feature weight coefficient matrix into each processing module to dynamically adjust the contribution of features at different levels.

[0142] To dynamically adjust the contribution of features at different levels, a learnable feature weight matrix is ​​introduced in each processing module. A feature weight matrix is ​​a matrix containing multiple weight coefficients, each corresponding to a feature or a group of features.

[0143] During training, the feature weight matrix is ​​learned and updated through the backpropagation algorithm. Based on the model's training objectives and loss function, the weight coefficients in the feature weight matrix are continuously adjusted to ensure that features at different levels contribute appropriately to the final output according to their importance.

[0144] For example, in the primary feature filtering module, the feature weight coefficient matrix can adjust the screening weights of different features; in the intermediate feature fusion module, it can adjust the fusion weights of features from different sources; in the high-level semantic generation module, it can adjust the temporal analysis weights of features at different times.

[0145] Step S1433-5: Set a semantic consistency verification module at the output end to perform a matching verification operation on the generated high-level semantics and the predefined knowledge rule base.

[0146] A semantic consistency verification module is set at the output end of the high-level semantic generation module to verify the consistency of the generated high-level semantics. The predefined knowledge rule base is a knowledge base that contains various semantic rules and constraints, which can be built based on domain knowledge and experience.

[0147] The generated high-level semantics are matched and verified against the predefined knowledge rule base. The semantic consistency verification module checks whether the generated high-level semantics conform to the rules and constraints in the knowledge rule base. For example, the knowledge rule base may specify logical relationships between certain semantics, and the semantic consistency verification module checks whether the generated high-level semantics meet these logical relationships.

[0148] If the generated high-level semantics do not match the rules in the knowledge rule base, it means there may be semantic errors or anomalies. At this time, the semantic consistency check module will mark these abnormal semantics and take corresponding measures.

[0149] Step S1433-6: Correct the abnormal semantics based on the verification results to ensure that the output results meet the domain constraints.

[0150] Based on the verification results of the semantic consistency check module, the abnormal semantics are corrected. If an abnormality is found in the generated high-level semantics, different correction strategies will be adopted according to the specific abnormality.

[0151] For example, if the abnormal semantics is caused by inaccurate feature extraction, you can re-extract the feature or adjust the feature extraction parameters; if the abnormal semantics is caused by unreasonable feature fusion, you can adjust the feature fusion method or weight.

[0152] By correcting abnormal semantics, we ensure that the output meets domain constraints. Domain constraints refer to the requirements for the rationality and validity of semantics in the target domain. Only semantic information that meets domain constraints can accurately describe the state and meaning of the target area.

[0153] Step S1433-7: Write the semantic information that has finally passed the verification into the corresponding field position of the global state description vector.

[0154] After semantic consistency verification and correction, the final verified semantic information is written into the corresponding field position of the global state description vector. The global state description vector is a vector used to comprehensively describe the state of the target area. It contains multiple fields, each of which corresponds to a specific feature or semantic information.

[0155] The verified semantic information is written into the corresponding field position of the global state description vector according to its corresponding features or semantic types. In this way, the global state description vector can accurately reflect the overall state and semantic information of the target area.

[0156] Step S1434: insert a feature normalization module between adjacent processing layers to perform normalization processing operations on the input features.

[0157] A feature normalization module is inserted between adjacent processing layers in a multi-layer feature processing architecture to normalize the input features. Feature normalization is an operation that uniformly scales and transforms feature data, which can make data of different features have the same scale and distribution.

[0158] In the feature normalization module, the mean and standard deviation of the input features are calculated. Then, the input features are normalized according to the mean and standard deviation, converting them into a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0159] Feature normalization can avoid computational errors and model training instability caused by scale differences between different features. It can also improve the convergence speed and generalization ability of the model.

[0160] Step S1435: Establish a residual connection channel between adjacent processing layers, and insert a dimension alignment module between the spatial topology relationship processing layer and the semantic abstraction processing layer to unify the dimensional space of the input features. Connect the dimensional projection module at the final output end to map the processed high-dimensional features to the low-dimensional semantic space.

[0161] Establishing residual connections between adjacent processing layers in a multi-layer feature processing architecture. Residual connections are a type of connection that passes input features directly to subsequent processing layers, addressing the vanishing and exploding gradient problems in deep learning models.

[0162] A dimension alignment module is inserted between the spatial topological relationship processing layer and the semantic abstraction processing layer to unify the dimensional space of the input features. Since the processing goals and methods of the spatial topological relationship processing layer and the semantic abstraction processing layer are different, the dimensions of the input features may be inconsistent. The dimension alignment module adjusts the feature dimensions output by the spatial topological relationship processing layer to a dimension that matches the input requirements of the semantic abstraction processing layer through a series of linear transformations and mapping operations. For example, it can learn a suitable transformation matrix to map high-dimensional spatial topological features to a dimensional space compatible with the semantic abstraction processing layer, to ensure that features can be smoothly transferred and processed between different processing layers.

[0163] The final output is connected to a dimensional projection module, which maps the processed high-dimensional features into a low-dimensional semantic space. While high-dimensional features contain rich information, in practical applications, excessively high dimensionality can increase computational complexity and storage costs, and may also introduce information redundancy. The dimensional projection module uses dimensionality reduction techniques, such as principal component analysis (although specific formulas are not used here), to identify the main components and key information in the high-dimensional features and then project them into a low-dimensional semantic space. This method preserves the key information of the features while reducing the dimensionality, making the final global state description vector more concise and easier to understand and apply.

[0164] Step S1436: Set a feature cache queue after the dimension projection module to temporarily store the processing results for subsequent operation calls.

[0165] A feature cache queue is set up after the dimension projection module to temporarily store feature results after dimension projection. Because these results may be used in multiple subsequent operations, such as further analysis, comparison, or fusion with other data, the feature cache queue avoids repeated calculations and improves system processing efficiency.

[0166] After the dimension projection module outputs processing results, it stores them sequentially in the feature cache queue. The queue's storage follows the first-in, first-out principle to ensure data order and consistency. Subsequent operations that require these feature results can retrieve them directly from the feature cache queue without requiring further processing such as dimension projection. Furthermore, to prevent the queue from growing indefinitely and causing excessive memory usage, a maximum queue capacity can be set. When the queue reaches maximum capacity, space is freed up according to a set policy (such as eliminating the oldest stored data).

[0167] Step S144: inputting the local feature vector set into the multi-layer feature processing architecture according to a predefined fusion priority order to perform layer-by-layer processing.

[0168] The predefined fusion priority order is pre-set based on the importance and relevance of different types of local feature vectors to the global state description. When the local feature vector set is input into the multi-layer feature processing architecture, it will be input in this priority order.

[0169] For example, local feature vectors closely related to the core information of the target area may be given higher priority and input into the multi-layer feature processing architecture for processing. Feature vectors of lesser importance, on the other hand, are input later in order. This ensures that important features are processed and analyzed first, improving the accuracy and efficiency of generating the global state description vector.

[0170] During the input process, each local feature vector first passes through the temporal correlation analysis layer. In this layer, contextual correlation analysis is performed using an adaptive sliding window mechanism to extract temporal correlation information. Next, after normalization by the feature normalization module, it enters the spatial topology relationship processing layer. In this layer, a topological structure model based on a graph neural network is constructed to capture spatial correlation features across sub-regions. After further processing by the feature normalization and dimension alignment modules, it enters the semantic abstraction processing layer. In this layer, it passes through the primary feature filtering module, the intermediate feature fusion module, and the high-level semantic generation module to gradually extract high-level semantic information.

[0171] Step S145: performing a context association analysis operation on the local features of the continuous time segments in the temporal association analysis processing layer.

[0172] In the temporal association analysis processing layer, it is very important to perform contextual association analysis on the local features of continuous time segments, because audio and video data have obvious time series characteristics, and there may be close connections between the features at different time points.

[0173] The adaptive sliding window mechanism slides over local features in consecutive time segments, capturing features within the window for analysis. The analysis considers the temporal order and changing trends of the features. For example, for the motion characteristics of an object in a video, the changes in its position, speed, and direction at different time points are analyzed to infer the object's trajectory and behavior pattern.

[0174] This contextual analysis can uncover hidden information within time series, providing richer and more accurate features for subsequent spatial topology analysis and semantic abstraction. Furthermore, the adaptive sliding window mechanism automatically adjusts the window size based on feature changes, making analysis more flexible and accurate.

[0175] Step S146: constructing a graph-structure-based feature propagation network model in the spatial topology relationship processing layer, converting topological features into fixed-dimensional vectors through graph embedding operations, and capturing spatial dependencies across sub-regions.

[0176] In the spatial topology processing layer, a graph-based feature propagation network model is constructed to better capture spatial dependencies across sub-regions. First, the set of local feature vectors is represented as a graph structure. Nodes in the graph represent different sub-regions, and the attributes of the nodes are the local feature vectors of the sub-region. Edges represent the spatial relationships between sub-regions, and the edge weights reflect the degree of spatial correlation between sub-regions.

[0177] The graph-based feature propagation network model propagates and aggregates information between nodes through a message-passing mechanism. Each node updates its feature vector based on information from its neighbors, allowing the node's feature vector to incorporate more cross-subregion information. For example, the features of one subregion may be influenced by the features of its neighboring subregions. The feature propagation network model can transmit and integrate this influence.

[0178] Graph embedding converts the topological features of a graph into fixed-dimensional vectors. This allows complex graph structures to be converted into a vector form that is easily processed by computers. During the graph embedding process, a mapping function is learned to map the node and edge information in the graph into a fixed-dimensional vector space. The resulting fixed-dimensional vectors can more effectively represent spatial dependencies across subregions, facilitating subsequent semantic abstraction.

[0179] Step S147: deploying a multi-level feature abstraction network in the semantic abstraction processing layer to gradually extract high-level semantic information.

[0180] The multi-level feature abstraction network in the semantic abstraction processing layer consists of a primary feature filtering module, an intermediate feature fusion module, and a high-level semantic generation module, which work together to gradually extract high-level semantic information.

[0181] The primary feature filtering module uses a gated attention mechanism to perform weighted filtering on the input local feature vectors. The gated attention mechanism learns the importance weights of each feature, assigning greater weights to important features and less weight to noisy features. This method filters out unimportant information, improving the accuracy and efficiency of subsequent processing.

[0182] After the feature dimension alignment layer maps multi-source features to a unified dimensional space, the mid-level feature fusion module uses a cross-attention mechanism to interactively fuse these features. This mechanism allows features from different sources to interact and focus on each other, uncovering potential connections between them. For example, the cross-attention mechanism can discover synchronization and semantic connections between image and audio features in a video.

[0183] The high-level semantic generation module uses a recurrent neural network model to perform temporal analysis on feature sequences after primary and secondary processing. The recurrent neural network model captures the temporal information in the feature sequence and uses this information to infer the evolutionary path of semantics, thereby generating high-level semantic information. For example, from a surveillance video, the high-level semantic generation module can infer the behavioral intentions of the characters and the development trends of events in the video.

[0184] Step S148: Generate a global state description vector with a unified dimensional space at the output end of the final processing layer, and write the global state description vector into a distributed storage system for persistent storage.

[0185] After layer-by-layer processing in the multi-layer feature processing architecture, a global state description vector with a unified dimensional space is generated at the output of the final processing layer. This global state description vector integrates information about the target area across multiple dimensions, including time, space, and semantics, and can comprehensively and accurately describe the state of the target area.

[0186] The unified dimensional space is achieved through the dimensional alignment module and the dimensional projection module. The dimensional alignment module ensures the consistency of feature dimensions between different processing layers, and the dimensional projection module maps high-dimensional features to a low-dimensional semantic space, making the global state description vector have a unified dimension.

[0187] The generated global state description vector is written to a distributed storage system for persistent storage. Distributed storage systems offer advantages such as high reliability, scalability, and high concurrent access capabilities, enabling secure and efficient storage of large numbers of global state description vectors. During the writing process, each global state description vector is assigned a unique identifier to facilitate subsequent query and use. Furthermore, to improve data access efficiency, global state description vectors can be indexed and stored in a categorized manner.

[0188] Step S150: generating a resource control instruction set according to the mapping relationship between the global state description vector and the preset system configuration policy, and transmitting the resource control instruction set to the scheduling management node of the virtualized resource pool.

[0189] The global state description vector comprehensively reflects the state of the target area, while the preset system configuration policy is a set of rules and parameters pre-defined based on the system's needs and objectives. By establishing a mapping relationship between the global state description vector and the preset system configuration policy, a corresponding set of resource control instructions can be generated based on the actual state of the target area.

[0190] For example, if the global state description vector indicates a sudden increase in the amount of audio and video data in the target area, the system may need to increase computing and storage resources to process and store this data, based on the preset system configuration policy. At this point, corresponding resource control instructions will be generated, such as increasing the number of computing nodes or expanding storage capacity.

[0191] The generated resource control instruction set is transmitted to the scheduling management node of the virtualized resource pool. The scheduling management node is responsible for the unified scheduling and management of the resources in the virtualized resource pool. Based on the received resource control instruction set, it will reasonably allocate and adjust resources to meet the needs of the target area.

[0192] Step S151: inputting the global state description vector into a pre-trained resource allocation decision model, wherein the resource allocation decision model comprises a resource demand prediction module, a strategy generation module and an instruction conversion module.

[0193] The pre-trained resource allocation decision model is a model trained on a large amount of data. It can generate reasonable resource control instructions based on the input global state description vector. The model consists of three main modules: resource demand prediction module, policy generation module, and instruction conversion module.

[0194] The global state description vector is input into the resource allocation decision model and first processed by the resource demand prediction module. This module analyzes the characteristic distribution patterns in the global state description vector and uses these patterns to predict resource requirements within a future time window. For example, by analyzing the traffic trends of audio and video data, it can predict the required computing and storage resources over a period of time.

[0195] Step S152: parsing the characteristic distribution pattern in the global state description vector by the resource demand prediction module to generate a resource demand prediction result in a future time window.

[0196] The resource demand prediction module conducts in-depth analysis of each feature in the global state description vector to uncover its distribution patterns. For example, for audio and video data, the module analyzes the periodic changes in traffic volume and sudden increases.

[0197] Based on the analyzed characteristic distribution patterns, combined with historical data and empirical models, a resource demand forecast for a future time window is generated. This resource demand forecast includes the quantity and timeframe of demand for different resource types (such as CPU, memory, and storage). For example, it is predicted that a certain number of CPU cores will be needed to process the increased audio and video data within a certain timeframe.

[0198] Step S153: The policy generation module combines the operating status parameters of the current virtualized resource pool to generate an optimization policy set including resource allocation ratio parameters, task scheduling path parameters, and node expansion parameters.

[0199] The strategy generation module generates a set of optimization strategies by combining resource demand forecasts with the current operational status parameters of the virtualized resource pool, which include resource usage and load balancing of each computing node.

[0200] Based on this information, the strategy generation module calculates reasonable resource allocation parameters and determines how to allocate limited resources to different tasks and computing nodes. For example, the number of CPU cores and memory size allocated to each task are determined based on the priority and resource requirements of each task.

[0201] At the same time, the strategy generation module also generates task scheduling path parameters and plans the scheduling order and path of tasks between different computing nodes to improve task execution efficiency and resource utilization. For example, some tasks can be scheduled to execute on computing nodes with lower resource utilization.

[0202] In addition, if the resource demand forecast results show that the current resources cannot meet future needs, the policy generation module will generate node expansion parameters to determine whether new computing nodes need to be added and the number and type of additional nodes.

[0203] Step S154: input the optimization strategy set into the instruction conversion module and convert it into an executable control instruction sequence.

[0204] The instruction conversion module converts the optimization strategy set into an executable control instruction sequence. First, it parses each strategy entry in the optimization strategy set and identifies the strategy operation type identifier, strategy parameters, and execution priority parameters.

[0205] The policy operation type identifier is used to distinguish different types of policy operations, such as resource allocation, task scheduling, and node expansion. Policy parameters are specific operation parameters, such as the amount of resource allocation and the path for task scheduling. The execution priority parameter is used to determine the execution order of each policy operation.

[0206] Based on the policy operation type identifier, the pre-stored instruction template library is matched and an intermediate representation instruction set compatible with the current system architecture is selected. The intermediate representation instruction set is a universal instruction representation that can be easily converted into local executable instructions for different target computing nodes.

[0207] The intermediate representation instruction set is converted into native executable instruction template instances of the target compute node through just-in-time compilation technology. Just-in-time compilation technology compiles the intermediate representation instruction set into native executable machine code based on the hardware architecture and operating system of the target compute node.

[0208] Mapping policy parameters to corresponding fields in the instruction template instance generates atomic control instruction units with defined operational semantics. For example, the resource allocation quantity parameter is filled into the corresponding field in the instruction template, giving the instruction a clear operational meaning.

[0209] The atomic control instruction units are logically sorted according to the execution priority parameters to form a complete instruction sequence. This ensures that the control instructions are executed in the correct order to avoid conflicts and errors.

[0210] Step S1541: Parse each policy entry in the optimization policy set to identify the policy operation type identifier, policy parameters, and execution priority parameters.

[0211] Perform detailed analysis of each policy entry in the optimization policy set, and identify key information by analyzing the text content and structure of the entry. The policy operation type identifier is a specific identifier or name used to clarify the operation type of the policy entry, such as "resource allocation", "task scheduling", "node expansion", etc. Policy parameters are specific numerical values ​​or configuration information used to guide the execution of policy operations, such as the specific number of resource allocation, the target node for task scheduling, etc. The execution priority parameter is a parameter that indicates the order in which the policy entries are executed. It can be a number or a priority level to ensure that the policies are executed in the correct order.

[0212] Step S1542: Match the pre-stored instruction template library according to the policy operation type identifier, select an intermediate representation instruction set compatible with the current system architecture, and convert the intermediate representation instruction set into a local executable instruction template instance of the target computing node through just-in-time compilation technology.

[0213] According to the identified policy operation type identifier, a matching search is performed in the pre-stored instruction template library. The instruction template library is a database containing instruction templates corresponding to various policy operations, and each template corresponds to a specific policy operation type.

[0214] Select an intermediate representation instruction set (IR) compatible with the current system architecture from the instruction template library. IR is a universal instruction representation that is independent of the specific hardware architecture and operating system, making it easy to convert and execute across different computing nodes.

[0215] Just-in-time compilation (JIT) technology is used to convert the intermediate representation instruction set into natively executable instruction template instances for the target compute node. Based on the hardware architecture and operating system characteristics of the target compute node, JIT compiles the intermediate representation instruction set into natively executable machine code. During the compilation process, code optimization and adjustments are performed to improve instruction execution efficiency and performance.

[0216] Step S1543: Map the policy parameters to corresponding field positions of the instruction template instance to generate an atomic control instruction unit with limited operational semantics.

[0217] Map the parsed policy parameters to the corresponding fields of the instruction template instance. The instruction template instance contains some reserved fields for receiving specific policy parameters. By filling these fields with policy parameters, the instruction template instance has clear operational semantics, becoming an atomic control instruction unit with a defined operational meaning.

[0218] For example, if the policy parameter is the quantity of resource allocation, the quantity value is filled into the field representing the quantity of resource allocation in the instruction template instance, so that the instruction specifies the specific quantity of resources to be allocated.

[0219] Step S1544: performing a logical sorting operation on the atomic control instruction unit according to the execution priority parameter to form a complete instruction sequence.

[0220] The atomic control instruction units are logically sorted according to their execution priority parameters. The execution priority parameter can be a number or a priority level. The smaller the number or the higher the priority level, the earlier the execution order of the atomic control instruction unit.

[0221] By sorting atomic control instruction units, we ensure that they are executed in the correct order. For example, in a resource allocation and task scheduling scenario, resource allocation may need to be performed before task scheduling. Therefore, the execution priority of the atomic control instruction unit for resource allocation will be higher than that of the atomic control instruction unit for task scheduling.

[0222] After the sorting is complete, these atomic control instruction units are connected in sequence to form a complete instruction sequence. This instruction sequence contains all policy operations and is arranged in the correct order, which can accurately guide the scheduling and management nodes of the virtualized resource pool to perform resource control operations.

[0223] Step S1545: After adding the check code field and the transaction tracking identification field at the starting position of the complete instruction sequence, the complete instruction sequence is encapsulated into a data packet format that complies with the network transmission protocol, and a timestamp mark and a source node identifier are added.

[0224] A checksum field and a transaction tracking identifier field are added to the beginning of the complete command sequence. The checksum field is used to verify whether the command sequence has been erroneous or corrupted during transmission. A checksum is generated by performing a specific calculation on the command sequence contents and added to the beginning of the command sequence. Upon receiving the command sequence, the receiver recalculates the checksum and compares it with the received checksum. If the two do not match, the command sequence may contain an error.

[0225] The transaction tracking identifier field is used to track the execution process and status of instruction sequences. Each instruction sequence has a unique transaction tracking identifier, which can be used to track the execution status of the instruction sequence throughout the system, such as whether it was executed successfully and the execution progress.

[0226] The complete instruction sequence, with the checksum and transaction tracking fields added, is encapsulated into a data packet format that complies with the network transmission protocol. Network transmission protocols specify the format and transmission rules for data packets. Encapsulating instruction sequences into protocol-compliant data packets ensures their secure and reliable transmission over the network.

[0227] Add a timestamp and source node identifier to the data packet. The timestamp records the time when the instruction sequence was generated, facilitating subsequent time analysis and sorting. The source node identifier identifies the sending node of the instruction sequence, allowing the receiver to know the source of the instruction.

[0228] Step S155: verify the control instruction sequence through the integrity check mechanism, mark the control instruction sequence that passes the verification as a pending execution state, and write it into the instruction queue to wait for processing by the scheduling management node.

[0229] The integrity check mechanism verifies the encapsulated control instruction sequence. First, the instruction sequence's contents are checked against the checksum field to detect errors or corruption during transmission. If the checksums do not match, the instruction sequence may have a problem and will be marked as invalid, with appropriate error handling.

[0230] If the verification passes, it means that the control instruction sequence is complete and has not been tampered with. At this time, the verified control instruction sequence will be marked as ready for execution. This marking process is performed in the system's task management module. By updating the status flag of the instruction sequence, the scheduling management node is informed that the instruction sequence can be executed.

[0231] The control instruction sequence marked as pending is then written to the instruction queue. The instruction queue is a first-in, first-out data structure that temporarily stores pending control instruction sequences. The scheduling management node retrieves the control instruction sequences sequentially from the instruction queue for processing. This ensures the correct execution order of the instructions, avoiding confusion and conflicts.

[0232] After receiving a sequence of control instructions from the instruction queue, the scheduling management node first parses the instruction sequence. For example, it identifies the individual atomic control instruction units within the instruction sequence, along with each unit's operation type, parameters, and execution priority. Based on this information, the scheduling management node allocates and schedules resources.

[0233] For resource allocation-type atomic control instruction units, the scheduling management node checks the available resources in the virtualized resource pool. For example, it can allocate the corresponding resources from the resource pool based on the resource type and quantity requested in the instruction. For example, if the instruction requires the allocation of a certain number of CPU cores and memory, the scheduling management node will search for currently idle compute nodes and allocate the required resources from these nodes. During the allocation process, the scheduling management node considers factors such as resource availability and load balancing to ensure the rational use of resources.

[0234] For atomic control instruction units with task scheduling, the scheduling management node assigns the task to the appropriate compute node for execution based on the task scheduling path parameters specified in the instruction. For example, it can consider factors such as the compute node load, resource utilization, and task priority to select the optimal scheduling solution. For example, if a task requires high computing resources, the scheduling management node will prioritize assigning it to a compute node with sufficient resources and a light load.

[0235] For atomic control instructions for node expansion, the scheduling management node launches new compute nodes based on the node expansion parameters specified in the instruction. This may involve creating new virtual machine instances and allocating new physical servers. During the node launch process, the scheduling management node ensures that the new node is correctly added to the virtualized resource pool and communicates and collaborates with other nodes.

[0236] During the execution of a control instruction sequence, the scheduling management node monitors the execution status of each atomic control instruction unit in real time. For example, it can record information such as the instruction's start time, end time, and execution result, and feed this information back to the system's monitoring module. If an atomic control instruction unit fails to execute, the scheduling management node will handle it according to the preset error handling mechanism. For example, it may attempt to re-execute the instruction or adjust resource allocation and scheduling to ensure the normal operation of the system.

[0237] At the same time, the scheduling management node dynamically adjusts the order of instructions in the instruction queue based on the execution status of the instruction sequence. If a high-priority instruction needs to be executed early, the scheduling management node will remove it from the queue and process it first. If an instruction cannot be executed temporarily due to insufficient resources or other reasons, the scheduling management node will pause it and move it to the end of the queue, waiting for the appropriate time to execute it.

[0238] Throughout the resource control process, audio and video data from the target area is continuously collected and analyzed to update the global state description vector. Based on this updated global state description vector, a new set of resource control instructions is generated using the pre-trained resource allocation decision model. The aforementioned instruction generation, verification, and execution processes are repeated, forming a closed-loop resource control system. This ensures that the virtualized resource pool can dynamically adjust resource allocation and scheduling based on the actual needs of the target area, improving system performance and efficiency.

[0239] Furthermore, the protection of privacy-sensitive data is crucial during data collection, processing, and transmission. For privacy-sensitive information in audio and video data, such as facial and voice features, a series of privacy protection and anti-leakage technologies are required. For example, during the data collection phase, strict permission management is implemented for collection devices. Only authorized collection devices can collect data within specific areas and time periods. Furthermore, the collected raw data is encrypted using symmetric encryption algorithms to ensure data security during transmission and storage. During the data processing phase, privacy-sensitive information is desensitized. For example, for facial features, blurring and deformation techniques are used to render them unrecognizable. For voice features, filtering and noise reduction are performed on the audio data to remove any potentially sensitive information. During feature extraction and analysis, only non-sensitive data after desensitization is used to prevent the leakage of private information. During data transmission, secure network transmission protocols, such as SSL / TLS, are used to encrypt data. Furthermore, the transmission process is monitored to detect any risks of data leakage. If any unusual network traffic or data transmission behavior is detected, timely measures will be taken to block and prevent it. During the data storage phase, encrypted audio and video data will be stored in a secure distributed storage system. The storage system will employ access control mechanisms to ensure that only authorized users can access and manipulate data. Furthermore, stored data will be regularly backed up and restored to ensure data reliability and availability.

[0240] The training process of pre-trained resource allocation decision models also requires strict management and control. The selection of training data ensures its legality and compliance, avoiding the use of data containing sensitive information or violating legal regulations. During training, privacy-preserving machine learning algorithms, such as differential privacy technology, are employed to process the training data, ensuring that no private information is leaked during the model learning process.

[0241] The model training steps are as follows: Collect a large amount of historical audio and video data, along with corresponding resource usage data. This data includes audio and video traffic over different time periods, compute node resource utilization, and task execution time. The collected data is then cleaned and preprocessed to remove noise and outliers, ensuring data quality and consistency.

[0242] Extract relevant features from cleaned historical audio and video data. For audio and video data, features such as traffic flow, frame rate, and resolution can be extracted. For resource usage data, features such as CPU usage, memory usage, and disk I / O can be extracted. Extracted features are normalized to ensure that different features have the same scale and range, facilitating model learning and training.

[0243] After feature extraction and normalization, the dataset is divided into training, validation, and test sets. The training set is used to train the model, the validation set is used to adjust the model parameters and evaluate the model's performance, and the test set is used to ultimately evaluate the model's generalization ability.

[0244] Build a resource allocation decision model consisting of a resource demand prediction module, a policy generation module, and an instruction conversion module. The resource demand prediction module can use deep learning models, such as long short-term memory (LSTM) or gated recurrent units (GRU), to predict resource demand within a future time window. The policy generation module can use a rule-based model or reinforcement learning model to generate an optimization policy set based on the resource demand prediction results and the current operating status of the virtualized resource pool. The instruction conversion module can use template matching and code generation techniques to convert the optimization policy set into an executable control instruction sequence.

[0245] The resource allocation decision model is trained using the training set. During training, the model parameters are updated using a backpropagation algorithm and an optimizer (such as stochastic gradient descent or the Adam optimizer) to ensure that the model's predictions and generated policies are as close to reality as possible. Simultaneously, the model's performance is evaluated using the validation set, and the model's parameters are adjusted based on the evaluation results to avoid overfitting and underfitting. A final evaluation of the trained resource allocation decision model is performed using the test set. Evaluation metrics include prediction accuracy, rationality of policy generation, and correctness of instruction conversion. If the model evaluation results do not meet the requirements, return to step S213, adjust the model's structure and parameters, and retrain and evaluate until the model performance reaches a satisfactory level.

[0246] Figure 2A schematic diagram illustrates exemplary hardware and software components of a virtual distributed architecture-based audio and video system data comprehensive analysis and processing system 100, which can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the virtual distributed architecture-based audio and video system data comprehensive analysis and processing system 100 to perform the functions described in the present application.

[0247] The audio and video system data comprehensive analysis and processing system 100 based on a virtual distributed architecture can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the audio and video system data comprehensive analysis and processing system 100 based on a virtual distributed architecture can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The audio and video system data comprehensive analysis and processing system 100 based on a virtual distributed architecture also includes an I / O interface 150 between the computer and other input and output devices.

[0248] In addition, an embodiment of the present invention also provides a readable storage medium, which has computer-executable instructions preset in the readable storage medium. When the processor executes the computer-executable instructions, the above-mentioned audio and video system data comprehensive analysis and processing method based on the virtual distributed architecture is implemented.

[0249] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A method for comprehensive analysis and processing of audio and video system data based on a virtual distributed architecture, characterized in that: The method comprises: Collecting a heterogeneous audio and video data set of the target area through a multi-source audio and video acquisition component cluster deployed in a virtualized resource pool, wherein the heterogeneous audio and video data set includes a sequence of coded data blocks encapsulated by different transmission protocols; Performing a distributed task decomposition operation on the heterogeneous audio and video data set according to preset spatiotemporal correlation constraints to generate a feature extraction task set matching different analysis dimensions; Dynamically assigning the feature extraction task set to corresponding processing unit groups based on real-time resource availability parameters of virtual computing nodes, and triggering each processing unit to execute a parallel feature extraction process corresponding to the assigned feature extraction task; Performing a fusion operation on the intermediate feature data sets generated by the parallel feature extraction process through a cross-node communication interface to generate a global state description vector of the target area; Generate a resource control instruction set according to the mapping relationship between the global state description vector and the preset system configuration policy, and transmit the resource control instruction set to the scheduling management node of the virtualized resource pool; The feature extraction task set is dynamically allocated to the corresponding processing unit group based on the real-time resource availability parameter of the virtual computing node, and each processing unit is triggered to execute the parallel feature extraction process corresponding to the allocated feature extraction task, specifically including: Real-time collection of real-time resource availability parameters of each computing node in the virtualized resource pool; Construct a dynamic scheduling decision model with task processing delay threshold, multi-dimensional resource utilization balance and load difference minimization as constraints; Convert each feature extraction task in the feature extraction task set into a task descriptor object; Performing task matching calculations on the task descriptor objects and the real-time resource availability parameters of the computing nodes through the dynamic scheduling decision model to generate a task allocation priority list; Distributing the feature extraction task to the target computing node according to the task allocation priority list, and activating the parallel processing unit in the target computing node to execute the feature extraction process; The task execution progress of the feature extraction task is monitored in real time and the task allocation strategy is dynamically adjusted according to changes in resource status.

2. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 1, characterized in that: The distributed task decomposition operation is performed on the heterogeneous audio and video data set according to the preset spatiotemporal correlation constraints to generate a feature extraction task set matching different analysis dimensions, specifically including: Performing a unified timestamp calibration operation on each coded data block sequence in the heterogeneous audio and video data set to obtain a calibrated coded data block sequence; Based on the sliding time window mechanism, the calibrated coded data block sequence is segmented and intercepted to generate a set of data segments with temporal continuity; Performing a spatial coordinate system conversion operation on multiple data streams in each data segment in the data segment set to establish a spatial alignment mapping relationship model across the data streams; Performing a spatial overlap region detection operation on multiple data streams within the data segment according to the spatial alignment mapping relationship model to identify a set of target sub-regions with spatial consistency; The target sub-region set is temporally and spatially associated with the corresponding data segments to generate a feature extraction task set containing multi-dimensional constraints, and the feature extraction task set is sorted according to preset priority rules to generate a final task allocation queue.

3. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 1, characterized in that: The step of performing a fusion operation on the intermediate feature data sets generated by the parallel feature extraction process through the cross-node communication interface to generate a global state description vector of the target area specifically includes: After the target computing node completes the feature extraction process, extracting the local feature vector set generated inside the target computing node; Transmitting the local feature vector set to a predefined aggregation service node through a distributed messaging middleware; Constructing a multi-layer feature processing architecture in the aggregation service node, the multi-layer feature processing architecture includes a temporal correlation analysis processing layer, a spatial topology relationship processing layer, and a semantic abstraction processing layer; Inputting the local feature vector set into the multi-layer feature processing architecture according to a predefined fusion priority order for layer-by-layer processing; Performing context association analysis on the local features of the continuous time segments in the temporal association analysis processing layer; A feature propagation network model based on a graph structure is constructed in the spatial topology relationship processing layer, and topological features are converted into fixed-dimensional vectors through graph embedding operations to capture spatial dependencies across sub-regions; Deploying a multi-level feature abstraction network in the semantic abstraction processing layer to gradually extract high-level semantic information; A global state description vector with a unified dimensional space is generated at the output end of the final processing layer, and the global state description vector is written into a distributed storage system for persistent storage.

4. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 3, characterized in that: The multi-layer feature processing architecture is constructed in the aggregation service node, specifically including: In the temporal correlation analysis processing layer, an adaptive sliding window mechanism is used to perform contextual correlation analysis on the local features of continuous time segments; A topological structure model based on graph neural network is constructed in the spatial topological relationship processing layer to capture the spatial correlation characteristics across sub-regions; Deploy a multi-level feature processing unit in the semantic abstraction processing layer, wherein the multi-level feature processing unit includes a primary feature filtering module, an intermediate feature fusion module, and a high-level semantic generation module; Insert feature normalization modules between adjacent processing layers to perform normalization operations on input features; A residual connection channel is established between adjacent processing layers, and a dimension alignment module is inserted between the spatial topology relationship processing layer and the semantic abstraction processing layer to unify the dimensional space of the input features. A dimension projection module is connected to the final output end to map the processed high-dimensional features to a low-dimensional semantic space. A feature cache queue is set after the dimension projection module to temporarily store the processing results for subsequent operation calls.

5. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 4, characterized in that: The multi-level feature processing unit is deployed in the semantic abstraction processing layer, specifically including: In the primary feature filtering module, a gated attention mechanism is used to perform weighted filtering operations on input features to suppress noise feature interference; A feature dimension alignment layer is deployed before the intermediate feature fusion module to map multi-source features to a unified dimensional space through linear projection. In the intermediate feature fusion module, a cross-attention mechanism is used to implement interactive fusion of multi-source features. In the high-level semantic generation module, a recurrent neural network model is used to perform temporal analysis on feature sequences and infer the semantic evolution path; Introducing a learnable feature weight coefficient matrix into each processing module to dynamically adjust the contribution of features at different levels; A semantic consistency verification module is set at the output end to match and verify the generated high-level semantics with the predefined knowledge rule base; Correct the abnormal semantics based on the verification results to ensure that the output results meet the domain constraints; The semantic information that has finally passed the verification is written into the corresponding field position of the global state description vector.

6. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 1, characterized in that: Generating a resource control instruction set according to a mapping relationship between the global state description vector and a preset system configuration policy specifically includes: Inputting the global state description vector into a pre-trained resource allocation decision model, wherein the resource allocation decision model includes a resource demand prediction module, a strategy generation module, and an instruction conversion module; Analyzing the feature distribution pattern in the global state description vector by the resource demand prediction module to generate a resource demand prediction result in a future time window; The strategy generation module combines the operating status parameters of the current virtualized resource pool to generate an optimization strategy set including resource allocation ratio parameters, task scheduling path parameters and node expansion parameters; Inputting the optimization strategy set into an instruction conversion module to convert it into an executable control instruction sequence; The control instruction sequence is verified through an integrity check mechanism, and the control instruction sequence that passes the verification is marked as a pending execution state and written into the instruction queue to wait for processing by the scheduling management node.

7. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 6, characterized in that: Inputting the optimization strategy set into the instruction conversion module and converting it into an executable control instruction sequence specifically includes: Parsing each policy entry in the optimization policy set to identify a policy operation type identifier, policy parameters, and execution priority parameters; Matching a pre-stored instruction template library according to the strategy operation type identifier, selecting an intermediate representation instruction set compatible with the current system architecture, and converting the intermediate representation instruction set into a local executable instruction template instance of the target computing node through a just-in-time compilation technology; Mapping the policy parameters to corresponding field positions of the instruction template instance to generate an atomic control instruction unit with limited operational semantics; Performing a logical sorting operation on the atomic control instruction units according to the execution priority parameters to form a complete instruction sequence; After adding a check code field and a transaction tracking identification field to the start position of the complete instruction sequence, the complete instruction sequence is encapsulated into a data packet format that complies with the network transmission protocol, and a timestamp mark and a source node identification are added.

8. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 1, characterized in that: The method of collecting heterogeneous audio and video data sets of the target area by using a multi-source audio and video collection component cluster deployed in the virtualized resource pool specifically includes: Establishing bandwidth-adjustable data transmission channels for different audio and video acquisition components deployed in the virtualized resource pool through a programmable network interface, and integrating a real-time quality control unit in the data transmission channel to dynamically adjust data sampling rate parameters, compression ratio parameters, and encoding format parameters; The coded data blocks transmitted by each data transmission channel are aligned and cached according to a time reference to form the heterogeneous audio and video data set with temporal and spatial consistency.

9. The method for comprehensive analysis and processing of audio and video system data based on virtual distributed architecture according to claim 8, characterized in that: The real-time quality control unit is integrated into the data transmission channel to dynamically adjust the data sampling rate parameters, compression ratio parameters and encoding format parameters, specifically including: Real-time monitoring of the current bandwidth utilization, end-to-end delay, and packet loss rate of the data transmission channel; Establish a dynamic adjustment model with data fidelity index, transmission efficiency index and energy consumption index as optimization targets; Performing real-time updating operations on the weight parameters of the dynamic adjustment model through an online learning algorithm; Candidate parameter combinations are calculated based on the updated dynamic adjustment model. A closed-loop feedback control algorithm is used to select a parameter combination that meets the data fidelity threshold and optimizes transmission efficiency. This generates data sampling rate adjustment parameters, compression ratio parameter combinations, and coding format switching strategies. Sending the data sampling rate adjustment parameter, compression ratio parameter combination and encoding format switching strategy to the data processing module of the audio and video acquisition component to perform a configuration update operation; The adjusted data quality indicators are fed back to the dynamic adjustment model in real time for iterative optimization.

Citation Information

Patent Citations

  • Big data hybrid scheduling model on private cloud condition

    CN105893158A

  • Edge-end-side heterogeneous computing power resource scheduling system and method

    CN119537018A