Server-edge-free broadcast data processing method, system, equipment and medium

By acquiring the temporal characteristics of broadcast data streams, identifying anchor point sets, and determining the boundaries of business semantic segments, semantic windowing and fusion processing is performed, solving the problems of latency and resource waste in traditional broadcast data processing, and achieving efficient and real-time semantic coherence processing.

CN121907784APending Publication Date: 2026-04-21山西广播电视无线管理中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
山西广播电视无线管理中心
Filing Date
2026-01-20
Publication Date
2026-04-21

Smart Images

  • Figure CN121907784A_ABST
    Figure CN121907784A_ABST
Patent Text Reader

Abstract

The invention relates to a server-edge-free broadcast data processing method, system and device and a medium. The method comprises the following steps: acquiring a broadcast data stream and extracting data time sequence characteristics; identifying a to-be-cut-in anchor point set by adopting a silent threshold boundary identification method based on data time sequence characteristics; based on the to-be-cut-in anchor point set and the data time sequence features, analyzing and determining a business semantic segment boundary through a time sequence feature mode; performing semantic windowing segmentation on the broadcast data stream based on the business semantic segment boundary to generate business semantic data slices; and performing semantic fusion processing on the service semantic data slices by adopting the server-free edge function instance to obtain an edge broadcast data processing result. By adopting the method, the real-time performance and reliability of broadcast data processing in the edge environment can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer technology, and in particular relates to a serverless edge broadcast data processing method, system, device and medium. Background Technology

[0002] With the rapid development of computer and communication technologies, especially the deep integration of edge computing and serverless architecture, serverless edge computing technology has emerged. This technology features distributed deployment, on-demand elastic scaling, and processing close to the data source, effectively reducing data transmission latency and improving computing resource utilization. It has been widely applied in real-time data processing scenarios such as broadcast data and IoT data streams. In traditional technologies, centralized processing transmits broadcast data streams generated by terminals to a cloud server, where the cloud handles the entire process, including data temporal feature extraction, data segmentation, and semantic processing. Traditional edge segmentation processing mechanically segments broadcast data streams at edge nodes using fixed time windows or fixed data volume thresholds, transmitting the segmented data fragments to edge functions for processing, without considering the temporal correlation characteristics and business semantic logic of the data itself.

[0003] However, current processing methods have many problems: centralized processing requires transmitting massive broadcast data streams over long distances to the cloud, which not only consumes a large amount of network bandwidth resources but also generates significant transmission delays, making it difficult to meet the real-time requirements of broadcast data processing; simple edge segmentation processing uses fixed window or fixed data volume segmentation methods without combining data temporal characteristics and business semantic associations, resulting in a lack of semantic coherence in the segmented data stream fragments, which easily leads to feature alignment deviations during subsequent semantic fusion processing, affecting the accuracy of the processing results; traditional processing methods do not design effective boundary recognition mechanisms for the temporal characteristics of broadcast data, making it impossible to accurately locate the boundaries of business semantic segments, resulting in data segmentation redundancy or loss of key semantic information; the matching degree between edge-side resource scheduling and data processing needs is low, and traditional edge function instance scheduling does not allocate resources in a targeted manner based on data semantic categories, resulting in resource waste or insufficient processing capacity, reducing the reliability and stability of broadcast data processing. Summary of the Invention

[0004] Therefore, it is necessary to provide a serverless edge broadcast data processing method, system, device, and medium that can solve the above problems.

[0005] Firstly, this application provides a serverless edge broadcast data processing method, including:

[0006] Acquire broadcast data streams and extract time-series features based on the broadcast data streams;

[0007] Based on the temporal characteristics of the data, the silent threshold boundary identification method is used to identify the set of anchor points to be cut in;

[0008] Based on the set of anchor points to be entered and the temporal characteristics of the data, the boundaries of the business semantic segments are determined through temporal characteristic pattern analysis.

[0009] Based on the business semantic segment boundary, the broadcast data stream is semantically windowed and segmented to generate business semantic data slices;

[0010] A serverless edge function instance is used to perform semantic fusion processing on business semantic data slices to obtain edge broadcast data processing results.

[0011] In one embodiment, based on data time-series characteristics, a silent threshold boundary identification method is used to identify the set of anchor points to be cut in, including:

[0012] Obtain data type information for the broadcast data stream;

[0013] The silent threshold of data time series features is determined by using a preset multimodal boundary threshold judgment rule and based on data type information.

[0014] Based on data time series characteristics and silent threshold, a sliding window method is used to filter out time windows in which the data time series characteristics are within the silent threshold range;

[0015] Based on the time window, the start and end times of the time window are extracted to obtain the silent boundary anchor point;

[0016] Integrate the silent boundary anchor points to obtain the set of anchor points to be entered.

[0017] In one embodiment, based on the set of anchor points to be entered and the temporal characteristics of the data, the boundaries of the business semantic segments are determined through temporal characteristic pattern analysis, including:

[0018] Based on each anchor point in the set of anchor points to be cut in, extract time series feature segments from the time series features of the data;

[0019] For each temporal feature segment, keyword sequences and temporal association identifiers are obtained through semantic information extraction;

[0020] Based on the keyword sequence, the semantic association similarity between each temporal feature segment is calculated;

[0021] Based on temporal association identifiers, the semantic association similarity is statistically analyzed to obtain the similarity distribution interval;

[0022] Based on semantic association similarity and similarity distribution interval, pattern clustering is performed on temporal feature segments to obtain feature pattern clusters;

[0023] Based on each feature pattern cluster, the boundary of the business semantic segment is determined according to the anchor point corresponding to the anchor point set to be cut in.

[0024] In one embodiment, the broadcast data stream is semantically segmented based on the business semantic segment boundary to generate business semantic data slices, including:

[0025] Based on the boundaries of business semantic segments, the time span data of each semantic segment is obtained by calculating the time difference between adjacent boundaries;

[0026] Based on the time series characteristics of the data, data density information is obtained, and based on the data density information, data density verification is performed on the time span data to obtain the data density verification result.

[0027] Based on the data density verification results, the semantic window range is determined for the business semantic segment boundary corresponding to the time span data that passed the verification.

[0028] The broadcast data stream is segmented into initial data slices according to the semantic window range.

[0029] Semantic annotation is performed on the initial data slices to generate business semantic data slices.

[0030] In one embodiment, based on data density information, data density verification is performed on the time span data to obtain the data density verification result, including:

[0031] Construct a spatiotemporal density matrix based on data density information;

[0032] Based on time span data, the spatiotemporal region to be verified is delineated from the spatiotemporal density matrix;

[0033] Calculate the density probability value of each spatiotemporal point within the spatiotemporal region to be verified;

[0034] Based on the density probability value, and according to the preset density fluctuation threshold, abnormal spatiotemporal points where the density probability value exceeds the density fluctuation threshold are identified.

[0035] Aggregate anomalous spatiotemporal points to form density anomaly regions;

[0036] The percentage of the total coverage of all density anomaly intervals in the spatiotemporal region to be verified is calculated. If the percentage is greater than or equal to the preset percentage threshold, the verification is deemed to have failed; otherwise, the verification is deemed to have passed, and the data density verification result is obtained.

[0037] In one embodiment, based on the boundaries of business semantic segments, the time span data of each semantic segment is obtained by calculating the time difference between adjacent boundaries, using the following formula:

[0038]

[0039] in, For the time span data of the i-th semantic segment, Let i be the time sequence of the boundary of the i-th business semantic segment. Let i be the time sequence of the (i+1)th business semantic segment boundary. The temporal association weight corresponding to the i-th semantic segment is... Let be the temporal offset correction coefficient for the i-th semantic segment. Here, m is the time-series decay constant, and m is the radius of the neighborhood boundary sampling window. Let be the temporal mean of the neighborhood of the boundary of the i-th semantic segment. Let be the temporal confidence coefficient of the k-th neighborhood boundary.

[0040] In one embodiment, a serverless edge function instance is used to perform semantic fusion processing on business semantic data slices to obtain edge broadcast data processing results, including:

[0041] Based on business semantic data slices, combined with keyword sequences, semantic category labels for the slices are obtained through semantic category mapping;

[0042] Based on slice semantic category labels, the corresponding serverless edge function instance cluster is matched through a preset edge function instance scheduling strategy;

[0043] Business semantic data slices with similar slice semantic category labels are input into a serverless edge function instance cluster, and semantic feature alignment is performed by combining them with temporal association identifiers to obtain an aligned semantic feature set.

[0044] Based on the aligned semantic feature set, the data is reorganized according to the temporal order corresponding to the boundaries of the business semantic segments to obtain the edge broadcast data processing results.

[0045] Secondly, this application also provides a serverless edge broadcast data processing system, comprising:

[0046] The data feature extraction module is used to acquire broadcast data streams and extract time-series features based on the broadcast data streams.

[0047] The anchor point set identification module is used to identify the set of anchor points to be cut in based on the data time series characteristics and the silent threshold boundary identification method.

[0048] The semantic boundary determination module is used to determine the boundaries of business semantic segments based on the set of anchor points to be entered and the temporal characteristics of the data, through temporal characteristic pattern analysis.

[0049] The semantic slice generation module is used to perform semantic windowing on the broadcast data stream based on the business semantic segment boundaries to generate business semantic data slices;

[0050] The semantic fusion processing module is used to perform semantic fusion processing on business semantic data slices using serverless edge function instances to obtain edge broadcast data processing results.

[0051] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described serverless edge broadcast data processing method.

[0052] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described serverless edge broadcast data processing method.

[0053] The aforementioned serverless edge broadcast data processing method, system, device, and medium acquire broadcast data streams and extract data temporal features to provide a data foundation for subsequent processing. Based on these temporal features, a silent threshold boundary recognition method is used to identify the set of anchor points to be cut, solving the problem of the lack of an effective boundary recognition mechanism in traditional processing. Combining the set of anchor points to be cut and the data temporal features, the boundaries of business semantic segments are determined through temporal feature pattern analysis, avoiding data segmentation redundancy or loss of key semantic information. Based on the boundaries of the business semantic segments, the broadcast data stream is semantically segmented into business semantic data slices, ensuring the semantic coherence of the slices and reducing the alignment deviation of subsequent fusion features. Serverless edge function instances are used to perform semantic fusion processing on the business semantic data slices, eliminating the need for long-distance transmission of massive amounts of data, reducing transmission latency, improving resource matching, and solving problems such as insufficient real-time performance of centralized processing and poor accuracy of simple edge segmentation. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a flowchart of a serverless edge broadcast data processing method according to the present invention;

[0056] Figure 2 This is a structural diagram of a serverless edge broadcast data processing system according to the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0058] In one embodiment, such as Figure 1 As shown, a serverless edge broadcast data processing method is provided. This embodiment illustrates the application of this method to a distributed edge computing architecture. It is understood that this method can also be applied to cloud-edge collaborative systems, and further to an integrated architecture including a broadcast data acquisition terminal, an edge node cluster, a serverless edge function instance pool, and a data storage module, achieved through interaction between the various hardware components. The hardware in the implementation environment of this method includes: broadcast data acquisition terminals (such as IoT sensing devices, live streaming terminals, emergency broadcast transmission equipment, etc.) and distributed edge nodes (integrating data processing units and communication modules), which are interconnected via a network. Application scenarios include: when there is a need for low-latency processing of massive amounts of broadcast data, avoiding the loss of key semantics, and optimizing resource utilization (such as real-time monitoring of IoT device status broadcasts, high-definition live data stream parsing, and semantic processing of emergency broadcast signals), the acquisition terminal transmits the broadcast data stream to the nearest edge node. The edge node extracts the temporal features of the data and accurately identifies the business semantic boundaries. By scheduling a cluster of function instances matching the semantic categories, it performs fusion processing on the data slices. The processing results can be fed back to the terminal in real time or uploaded to the cloud as needed, achieving efficient local data processing and meeting the needs of real-time business response and reliable processing.

[0059] In this embodiment, the method includes the following steps:

[0060] S01, acquire the broadcast data stream, and extract the data time sequence features based on the broadcast data stream.

[0061] Broadcast data streams typically refer to continuous data sequences generated in real time from distributed data sources such as IoT sensing devices, live streaming terminals, or emergency broadcasting equipment. Their characteristics include high throughput, time-series dependence, and multimodality. Data time-series features are a set of features reflecting the dynamic changes of data over time, such as quantifying data trends, periodicity, and abrupt changes through statistical indicators (e.g., mean, variance), frequency domain transformations (e.g., Fourier coefficients), or time-series patterns (e.g., autocorrelation functions). In implementation, broadcast data streams can be acquired through the network interface between edge computing nodes and data acquisition hardware. Based on the data stream, time-series feature extraction techniques, such as sliding window analysis, wavelet transform, or deep learning models (e.g., recurrent neural networks), are applied to capture key time-series attributes, thus obtaining the data time-series features.

[0062] S02, based on the time-series characteristics of the data, the silent threshold boundary identification method is used to identify the set of anchor points to be cut in.

[0063] The silent threshold boundary identification method is a technique that adaptively sets thresholds based on the dynamic characteristics of data to identify low-change or silent regions in a data stream. Its core lies in analyzing time-series characteristics (such as fluctuation amplitude, frequency components, or statistical indicators) to detect relatively stable intervals and reduce irrelevant components. The anchor point set refers to a series of time points or window boundaries used to identify candidate locations in the data stream where business semantics may change. This can be achieved through multimodal adaptation mechanisms (such as dynamically adjusting thresholds based on data type) or algorithms (such as sliding window analysis, peak detection, or machine learning models). In implementation, the silent threshold can be determined using preset rules (such as statistical quantiles or entropy calculations) or real-time learning (such as online clustering). Combined with window sliding, boundary aggregation, or anomaly detection techniques, time anchor points that meet the conditions are selected, achieving robustness under different broadcast scenarios (such as audio streams, video streams, or sensor data).

[0064] S03. Based on the set of anchor points to be entered and the temporal characteristics of the data, the boundaries of the business semantic segments are determined through temporal characteristic pattern analysis.

[0065] Temporal feature pattern analysis is a technique based on temporal data fragments for feature extraction, similarity calculation, and cluster analysis. It is used to identify intervals with consistent semantic patterns in a data stream. The business semantic segment boundary refers to the temporal position of different business units (such as dialogue segments, event segments, or topic transition points) in the data stream. It can be implemented through algorithms such as fragment clustering, statistical distribution analysis, or machine learning models. In practice, temporal feature fragments can be extracted from the temporal features of the data based on the anchor point set to be entered. The semantic correlation between fragments is analyzed by extracting semantic information (such as keyword sequence generation or temporal association identifier calculation). Then, the fragments are grouped into feature pattern clusters using similarity measures (such as cosine similarity or dynamic time warping) and pattern clustering (such as K-means or hierarchical clustering). The business semantic segment boundary is mapped according to the anchor points within the cluster. This allows it to adapt to dynamic changes in data under different broadcast scenarios (such as real-time audio and video streams or IoT data), avoid semantic breaks caused by fixed window segmentation, provide reliable input for subsequent semantic windowing, and meet the requirements of high real-time performance and low resource consumption under serverless architecture.

[0066] S04. Based on the business semantic segment boundary, semantic windowing is performed on the broadcast data stream to generate business semantic data slices.

[0067] Semantic windowing is a data segmentation process that dynamically adjusts the size and position of windows based on business logic boundaries (such as dialogue turning points or event transition moments). This ensures that data within each window possesses complete semantic unitality. Business semantic data slices refer to data segments with accompanying semantic annotations (such as topic tags or time sequence identifiers) after segmentation, facilitating subsequent processing. In implementation, the temporal difference between adjacent boundaries can be calculated based on the boundaries of business semantic segments (e.g., through formulaic time span calculations or adaptive window adjustments). This difference is then combined with data density information for verification (e.g., constructing a spatiotemporal density matrix and identifying abnormal regions) to determine the semantic window range. Windowing techniques (such as fixed-step sliding, variable-window segmentation, or machine learning-driven dynamic segmentation) are used to divide the broadcast data stream, generating initial data slices. Semantic annotations (such as keyword mapping and category label attachment) enhance the semantic information of the slices, ensuring their coherence and processability. Semantic segmentation effectively avoids information fragmentation caused by mechanical windows, improving the accuracy of subsequent semantic fusion.

[0068] S05, using a serverless edge function instance to perform semantic fusion processing on the business semantic data slices, to obtain the edge broadcast data processing results.

[0069] Serverless edge function instances are computing units that are elastically deployed on edge nodes on demand to support serverless function execution. Semantic fusion processing refers to the operations of feature extraction, alignment, and reorganization of data slices to ensure semantic coherence. In implementation, serverless edge function instance clusters can be dynamically allocated based on the semantic category labels of business semantic data slices (e.g., obtained through keyword sequence mapping) using a preset scheduling strategy (e.g., load balancing or priority matching). Slices of the same type are input into the cluster and semantic feature alignment is performed using temporal association identifiers (e.g., through similarity calculation or sequence modeling). This generates an aligned semantic feature set, which is then reorganized according to the temporal order corresponding to the boundaries of the business semantic segments to obtain structured edge broadcast data processing results, improving the accuracy of data processing and resource utilization.

[0070] In one embodiment, based on data time-series characteristics, a silent threshold boundary identification method is used to identify the set of anchor points to be cut in, including:

[0071] S11, obtain the data type information of the broadcast data stream;

[0072] S12, using a preset multimodal boundary threshold determination rule, determines the silent threshold of data time series characteristics based on data type information;

[0073] S13. Based on the data time series characteristics and the silent threshold, a sliding window method is used to filter out the time windows in which the data time series characteristics are within the silent threshold range.

[0074] S14, Based on the time window, extract the start and end times of the time window to obtain the silent boundary anchor point;

[0075] S15, integrate silent boundary anchor points to obtain the set of anchor points to be cut in.

[0076] For example, data type information (such as audio data, video data, sensor monitoring data, etc.) can be parsed from the broadcast data stream through the header identifier, metadata carried by the data transmission protocol, or a preset data type configuration table. The preset multimodal boundary threshold determination rule is a threshold determination rule set in advance for the temporal characteristic fluctuation characteristics (such as audio volume fluctuation, video frame rate change, and sensor data numerical fluctuation range) corresponding to different data types. Based on the parsed data type information, the determination rule of the corresponding type is called to calculate the silence threshold corresponding to the temporal characteristic of the data (e.g., for audio data, the feature threshold corresponding to the volume being lower than the preset base value for a certain duration is used as the silence threshold; for sensor data, the feature threshold corresponding to the numerical fluctuation amplitude being less than the set fluctuation range is used as the silence threshold). The process employs a sliding mode with a fixed window size (preset based on data type, e.g., 100ms for audio data and 500ms for sensor data) or an adaptive window size. Data time-series features are sequentially extracted using the sliding window. Each window's time-series features are checked against a silent threshold, and all time windows meeting this condition are selected. The start and end times of each selected time window are extracted using timestamps or time-series indexes to identify silent boundary anchors that define the data stream's silent state. All extracted silent boundary anchors are sorted chronologically, and duplicate or out-of-range anchors, or those with excessively small time intervals with other anchors, are removed. The remaining valid silent boundary anchors are then integrated to obtain the set of anchors to be entered.

[0077] In one embodiment, based on the set of anchor points to be entered and the temporal characteristics of the data, the boundaries of the business semantic segments are determined through temporal characteristic pattern analysis, including:

[0078] S21, Based on each anchor point in the set of anchor points to be cut in, extract time series feature segments from the time series features of the data;

[0079] S22, for each temporal feature segment, keyword sequences and temporal association identifiers are obtained through semantic information extraction;

[0080] S23, Calculate the semantic association similarity between each temporal feature segment based on the keyword sequence;

[0081] S24. Based on the temporal association identifier, the semantic association similarity is statistically analyzed to obtain the similarity distribution interval;

[0082] S25. Based on semantic association similarity and similarity distribution interval, perform pattern clustering on temporal feature segments to obtain feature pattern clusters;

[0083] S26. Based on each feature pattern cluster, the boundary of the business semantic segment is determined according to the anchor point corresponding to the anchor point set to be cut in.

[0084] Specifically, each anchor point in the set of anchor points to be cut can be used as a time reference, and two adjacent anchor points can be used as start and end boundaries, respectively. Feature data within the corresponding time interval can be extracted from the extracted temporal features of the data to obtain multiple independent temporal feature segments. For each temporal feature segment, semantic information extraction techniques (such as TF-IDF algorithm, BERT pre-trained model, or LDA topic model) can be used to mine core semantic elements from the segment, generate a keyword sequence composed of key terms, and assign a unique temporal association identifier to each segment (including the segment's start timestamp, duration, and temporal dependency index with other segments). The keyword sequence of each temporal feature segment is converted into a high-dimensional vector representation (such as word embedding through Word2Vec or GloVe models), and algorithms such as cosine similarity, Jaccard similarity, or Pearson correlation coefficient are used to calculate the semantic association similarity between any two temporal feature segments, quantizing the segments. The semantic fit between segments is assessed. Based on the temporal dependencies between segments as reflected by the temporal association identifiers, all calculated semantic association similarities are grouped and statistically analyzed according to the degree of association. By calculating the maximum, minimum, median, and standard deviation of the similarity for each group, the similarity distribution interval corresponding to each group is determined. Pattern clustering algorithms such as K-means clustering, hierarchical clustering, or density clustering (DBSCAN) can be used. Based on whether the semantic association similarity of each temporal feature segment is within a reasonable range of the similarity distribution interval of the corresponding group, temporal feature segments with close semantic association and consistent feature patterns are aggregated into multiple feature pattern clusters. For each feature pattern cluster, the anchor points to be cut into all temporal feature segments within the cluster are extracted. The common starting anchor point of the segments within the cluster is determined as the starting boundary of the business semantic segment, and the common ending anchor point is determined as the ending boundary of the business semantic segment. Thus, the boundary of the business semantic segment that can divide complete business semantic units is delineated.

[0085] In one embodiment, the broadcast data stream is semantically segmented based on the business semantic segment boundary to generate business semantic data slices, including:

[0086] S31, based on the boundaries of business semantic segments, the time span data of each semantic segment is obtained by calculating the time difference between adjacent boundaries;

[0087] S32, based on the time series characteristics of the data, obtain the data density information, and based on the data density information, perform data density verification on the time span data to obtain the data density verification result;

[0088] S33, based on the data density verification results, determine the semantic window range for the business semantic segment boundary corresponding to the verified time span data;

[0089] S34. The broadcast data stream is segmented into windows according to the semantic window range to obtain the initial data slice;

[0090] S35 performs semantic annotation on the initial data slices to generate business semantic data slices.

[0091] For example, based on the determined boundaries of business semantic segments, the time span data corresponding to each semantic segment can be calculated by calculating the difference in time sequence between the boundaries of two adjacent business semantic segments and substituting it into a preset time span calculation formula. Data density information represents the density of data distribution within a unit of time for the data time sequence characteristics (e.g., obtained through indicators such as the number of data points within a unit time window and the frequency of data fluctuations in the statistical data time sequence characteristics). Based on this data density information, data density verification can be performed on the time span data. By constructing a spatiotemporal density matrix, a spatiotemporal region to be verified, encompassing the entire spatiotemporal range of the semantic segment, is delineated from the spatiotemporal density matrix according to the time span data. The density probability value of each spatiotemporal point within this region is calculated. Based on a preset density fluctuation threshold, abnormal spatiotemporal points with density probability values ​​exceeding the threshold are identified. These abnormal spatiotemporal points are clustered to form density anomaly intervals, and the statistical results are then analyzed. The proportion of the total coverage area of ​​the density anomaly interval within the spatiotemporal region to be verified is determined. If this proportion is greater than or equal to a preset percentage threshold, the verification is deemed to have failed; otherwise, the verification is deemed to have passed, and the data density verification result is obtained. Based on this data density verification result, the business semantic segment boundaries corresponding to the verified time span data are selected. Based on the temporal range of these boundaries and combined with the distribution pattern of data temporal characteristics, the semantic window range corresponding to each semantic segment is determined. According to the determined semantic window range, the broadcast data stream is segmented into windows using a temporal truncation method, and complete data fragments within each window range are extracted to obtain initial data slices. The initial data slices are semantically annotated, and the annotation content includes the keyword sequence, semantic category label, temporal association identifier, and other information related to business semantics, generating business semantic data slices with complete semantic information and semantic coherence.

[0092] In one embodiment, based on data density information, data density verification is performed on the time span data to obtain the data density verification result, including:

[0093] S41, constructing a spatiotemporal density matrix based on data density information;

[0094] S42, based on time span data, delineate the spatiotemporal region to be verified from the spatiotemporal density matrix;

[0095] S43, Calculate the density probability value of each spatiotemporal point in the spatiotemporal region to be verified;

[0096] S44, based on the density probability value, identifies abnormal spatiotemporal points where the density probability value exceeds the density fluctuation threshold according to the preset density fluctuation threshold;

[0097] S45 aggregates anomalous spatiotemporal points to form density anomaly regions;

[0098] S46. Calculate the proportion of the total coverage of all density anomaly intervals in the spatiotemporal region to be verified. If the proportion is greater than or equal to the preset proportion threshold, the verification is deemed to have failed. Otherwise, the verification is deemed to have passed, and the data density verification result is obtained.

[0099] Specifically, when constructing a spatiotemporal density matrix based on data density information, the time dimension is used as the horizontal axis and the corresponding spatial dimension of the data (if the broadcast data has spatial attributes, such as the deployment location of IoT sensing devices; otherwise, only the time dimension is extended) is used as the vertical axis. The density data corresponding to each spatiotemporal coordinate is filled into the corresponding cells of the matrix to form a complete spatiotemporal density matrix. Based on the semantic time range defined by the time span data, a region containing this time range and the corresponding data distribution space is delineated in the spatiotemporal density matrix as the spatiotemporal region to be verified. The density probability value of each spatiotemporal point in the spatiotemporal region to be verified can be calculated using the kernel density estimation method. The density probability value of each spatiotemporal point is obtained by calculating the ratio of the average density of the spatiotemporal point within a preset neighborhood radius (based on the data sampling frequency, such as when the broadcast data is sampled 10 times per second, the neighborhood radius is set to the time interval corresponding to 3 sampling points) to the overall average density of the spatiotemporal region to be verified. The density probability value of each spatiotemporal point is then compared with a preset density fluctuation threshold. (Based on the density fluctuation statistics of historical broadcast data of the same type, such as set to ±25% of the overall average density) comparison is performed to screen out spatiotemporal points with density probability values ​​exceeding the threshold range as abnormal spatiotemporal points; the DBSCAN clustering algorithm can be used to aggregate abnormal spatiotemporal points that are spatially less than the preset threshold (such as the physical distance threshold for data transmission coverage) and temporally continuous, forming one or more continuous density abnormal intervals; the total coverage area of ​​all density abnormal intervals (the product of time span and spatial coverage area, or only the sum of time spans when there is no spatial attribute) is calculated, and the proportion of this total coverage area to the total area of ​​the spatiotemporal region to be verified (the product of time span and spatial coverage area, or only the time span to be verified when there is no spatial attribute) is calculated. If this proportion is greater than or equal to the preset proportion threshold (set based on the business requirements for data density uniformity, such as 12%), the corresponding time span data is determined to have failed the data density verification; otherwise, it is determined to have passed, and the data density verification result is obtained.

[0100] In one embodiment, S51, based on the boundaries of the business semantic segments, the time span data of each semantic segment is obtained by calculating the time difference between adjacent boundaries, and is achieved through the following formula:

[0101]

[0102] in, For the time span data of the i-th semantic segment, Let i be the time sequence of the boundary of the i-th business semantic segment. Let i be the time sequence of the (i+1)th business semantic segment boundary. The temporal association weight corresponding to the i-th semantic segment is... Let be the temporal offset correction coefficient for the i-th semantic segment. Here, m is the time-series decay constant, and m is the radius of the neighborhood boundary sampling window. Let be the temporal mean of the neighborhood of the boundary of the i-th semantic segment. Let be the temporal confidence coefficient of the k-th neighborhood boundary.

[0103] For example, the time span data of the i-th semantic segment is calculated. At that time, the time sequence of the boundary of the i-th business semantic segment can be extracted. The time sequence of the (i+1)th business semantic segment boundary And perform the following calculations: Temporal correlation weights The autocorrelation coefficient is determined based on the time-series characteristics of the data within the i-th semantic segment. The autocorrelation coefficient reflects the strength of the time-series correlation of the data; the stronger the correlation, the higher the correlation. The closer the value is to 1, the better the range should be set between 0.5 and 1.0. For example, if the autocorrelation coefficient is 0.9... Set to 0.95; Timing offset correction factor The calibration is based on the degree of offset of the temporal characteristics of the semantic segment data. A value of 0 is used for no offset, 0.1-0.3 for slight offset, and 0.3-0.6 for significant offset, with a value range of 0-1.0; temporal decay constant. Based on the preset sampling frequency and transmission rate of the broadcast data, with a sampling frequency of 10Hz-20Hz and a transmission rate of 1Mbps-10Mbps... The value is taken from 5 to 10 seconds; the radius m of the neighborhood boundary data collection window is a positive integer, set based on the average distribution of business semantic segments, usually 3-5; the neighborhood temporal mean of the i-th semantic segment boundary. The temporal confidence coefficient of the k-th neighboring boundary is obtained by calculating the average of the temporal times of each of the m neighboring boundaries before and after the i-th boundary. Based on the reliability assignment of the neighborhood boundary, the boundary transformed from a high-confidence silent anchor point is set to 0.8-1.0, and the boundary transformed from a normal anchor point is set to 0.4-0.8. Calculations using the above formula can ensure... It can accurately reflect the actual time span of semantic segments and adapt to subsequent data density verification and semantic windowing requirements.

[0104] In one embodiment, a serverless edge function instance is used to perform semantic fusion processing on business semantic data slices to obtain edge broadcast data processing results, including:

[0105] S61, based on business semantic data slicing, combined with keyword sequences, obtains slice semantic category labels through semantic category mapping;

[0106] S62, based on slice semantic category labels, matches the corresponding serverless edge function instance cluster through a preset edge function instance scheduling strategy;

[0107] S63, input the business semantic data slices with the same slice semantic category label into the serverless edge function instance cluster, combine them with the temporal association identifier, perform semantic feature alignment processing, and obtain the aligned semantic feature set;

[0108] S64, based on the aligned semantic feature set, is reorganized according to the temporal order corresponding to the business semantic segment boundary to obtain the edge broadcast data processing result.

[0109] Specifically, when mapping semantic categories based on business semantic data slices and corresponding keyword sequences, a dictionary containing common semantic categories in broadcast business scenarios (such as "emergency notification" for emergency broadcasts, "user interaction" for live data streams, and "device status" for IoT broadcasts) can be used. Each semantic category in the dictionary is associated with a set of core keywords. By calculating the Jaccard similarity (i.e., the ratio of the number of intersection elements to the number of union elements) between the keyword sequence of the business semantic data slice and the core keyword set of each semantic category, the semantic category with the highest similarity exceeding a preset threshold (e.g., 0.6) is determined as the [specific category]. If the semantic category labels of the slices do not exceed the threshold, they are marked as "general category". When matching serverless edge function instance clusters based on slice semantic category labels, according to the preset functional adaptation relationship between edge function instance clusters and semantic categories (e.g., the "video parsing" label corresponds to clusters deploying H.264 / H.265 codec functions, and the "data statistics" label corresponds to clusters deploying statistical functions such as summation and mean calculation), the load status of each cluster (including CPU utilization, memory usage, and task queue length) is collected in real time. A minimum load priority scheduling strategy can be adopted (prioritizing clusters with CPU utilization <70% and memory usage <70%). Clusters with occupancy rates < 60% and task queue lengths < 50 are selected from suitable clusters using a round-robin algorithm to identify the target serverless edge function instance cluster. After inputting business semantic data slices with semantic category labels of the same type of slice into this cluster, the keyword sequence of each slice is transformed into a 128-dimensional semantic feature vector using the Word2Vec model (a shallow neural network model used to compress high-dimensional sparse word representations into low-dimensional dense word vectors). This extended feature vector is then supplemented with the timestamp information corresponding to the time-series association identifier, and dynamic time warping is applied based on the timestamp of the time-series association identifier. The DTW algorithm calculates the temporal distance between extended feature vectors of similar slices, adjusts the arrangement order of feature vectors of each slice, eliminates feature misalignment caused by temporal offset, and selects effective feature vectors with a temporal distance less than a preset threshold (such as 0.3) to form an aligned semantic feature set. According to the temporal order of the business semantic segment boundary, the effective feature vectors in the aligned semantic feature set are sequentially reorganized into a continuous feature sequence. The structured edge broadcast data processing result containing semantic category, time span, core semantic features and key information summary is generated through feature sequence reconstruction, realizing the semantic coherence and practicality of the processing result.

[0110] The aforementioned serverless edge broadcast data processing method acquires broadcast data streams and extracts data temporal features. Based on these features, a multimodal silent threshold boundary recognition method adapted to data types is used to obtain a set of anchor points to be entered, establishing a targeted boundary recognition mechanism to solve the problem of traditional processing's difficulty in locating potential semantic transition points. Combining the set of anchor points to be entered with data temporal features, the method uses temporal feature pattern analysis to determine the boundaries of business semantic segments, dividing semantic units to avoid data segmentation redundancy or loss of key semantic information, thus solving the semantic breakage problem caused by traditional fixed boundary recognition. Based on the boundaries of business semantic segments, the method calculates temporal differences and data density... The verification and semantic annotation complete semantic window segmentation, generating business semantic data slices, ensuring the semantic coherence of the slices, reducing feature alignment deviations during subsequent fusion, and solving the problem of semantic fragmentation in simple edge segmentation; the serverless edge function instance cluster is scheduled and matched according to the slice semantic category label, and semantic feature alignment and temporal reorganization are performed by combining temporal association identifiers, eliminating the need for long-distance transmission of massive data, reducing transmission latency, improving the matching degree of edge-side resource scheduling and processing needs, solving the problems of insufficient real-time performance, resource waste, or insufficient processing capacity in centralized processing, and significantly improving the real-time performance, reliability, and accuracy of broadcast data processing in edge environments.

[0111] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0112] Based on the same inventive concept, this application also provides a serverless edge broadcast data processing system for implementing the serverless edge broadcast data processing method described above. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations of one or more embodiments of the serverless edge broadcast data processing system provided below can be found in the limitations of the serverless edge broadcast data processing method described above, and will not be repeated here.

[0113] In one exemplary embodiment, such as Figure 2 As shown, a serverless edge broadcast data processing system is provided, comprising:

[0114] The data feature extraction module 101 is used to acquire broadcast data streams and extract time-series features of the data based on the broadcast data streams;

[0115] Anchor point set identification module 102 is used to identify the anchor point set to be cut in based on data time series characteristics and using the silent threshold boundary identification method.

[0116] The semantic boundary determination module 103 is used to determine the boundary of the business semantic segment based on the set of anchor points to be cut in and the temporal characteristics of the data, through temporal characteristic pattern analysis.

[0117] The semantic slice generation module 104 is used to perform semantic windowing on the broadcast data stream based on the business semantic segment boundary to generate business semantic data slices;

[0118] The semantic fusion processing module 105 is used to perform semantic fusion processing on business semantic data slices using serverless edge function instances to obtain edge broadcast data processing results.

[0119] In one embodiment, the anchor point set identification module 102 is further configured to:

[0120] Obtain data type information for the broadcast data stream;

[0121] The silent threshold of data time series features is determined by using a preset multimodal boundary threshold judgment rule and based on data type information.

[0122] Based on data time series characteristics and silent threshold, a sliding window method is used to filter out time windows in which the data time series characteristics are within the silent threshold range;

[0123] Based on the time window, the start and end times of the time window are extracted to obtain the silent boundary anchor point;

[0124] Integrate the silent boundary anchor points to obtain the set of anchor points to be entered.

[0125] In one embodiment, the semantic boundary determination module 103 is further configured to:

[0126] Based on each anchor point in the set of anchor points to be cut in, extract time series feature segments from the time series features of the data;

[0127] For each temporal feature segment, keyword sequences and temporal association identifiers are obtained through semantic information extraction;

[0128] Based on the keyword sequence, the semantic association similarity between each temporal feature segment is calculated;

[0129] Based on temporal association identifiers, the semantic association similarity is statistically analyzed to obtain the similarity distribution interval;

[0130] Based on semantic association similarity and similarity distribution interval, pattern clustering is performed on temporal feature segments to obtain feature pattern clusters;

[0131] Based on each feature pattern cluster, the boundary of the business semantic segment is determined according to the anchor point corresponding to the anchor point set to be cut in.

[0132] In one embodiment, the semantic segmentation generation module 104 is further configured to:

[0133] Based on the boundaries of business semantic segments, the time span data of each semantic segment is obtained by calculating the time difference between adjacent boundaries;

[0134] Based on the time series characteristics of the data, data density information is obtained, and based on the data density information, data density verification is performed on the time span data to obtain the data density verification result.

[0135] Based on the data density verification results, the semantic window range is determined for the business semantic segment boundary corresponding to the time span data that passed the verification.

[0136] The broadcast data stream is segmented into initial data slices according to the semantic window range.

[0137] Semantic annotation is performed on the initial data slices to generate business semantic data slices.

[0138] In one embodiment, the semantic segmentation generation module 104 is further configured to:

[0139] Construct a spatiotemporal density matrix based on data density information;

[0140] Based on time span data, the spatiotemporal region to be verified is delineated from the spatiotemporal density matrix;

[0141] Calculate the density probability value of each spatiotemporal point within the spatiotemporal region to be verified;

[0142] Based on the density probability value, and according to the preset density fluctuation threshold, abnormal spatiotemporal points where the density probability value exceeds the density fluctuation threshold are identified.

[0143] Aggregate anomalous spatiotemporal points to form density anomaly regions;

[0144] The percentage of the total coverage of all density anomaly intervals in the spatiotemporal region to be verified is calculated. If the percentage is greater than or equal to the preset percentage threshold, the verification is deemed to have failed; otherwise, the verification is deemed to have passed, and the data density verification result is obtained.

[0145] In one embodiment, the semantic segment generation module 104 is further configured to obtain the time span data of each semantic segment based on the business semantic segment boundaries by calculating the time difference between adjacent boundaries using the following formula:

[0146]

[0147] in, For the time span data of the i-th semantic segment, Let i be the time sequence of the boundary of the i-th business semantic segment. Let i be the time sequence of the (i+1)th business semantic segment boundary. The temporal association weight corresponding to the i-th semantic segment is... Let be the temporal offset correction coefficient for the i-th semantic segment. Here, m is the time-series decay constant, and m is the radius of the neighborhood boundary sampling window. Let be the temporal mean of the neighborhood of the boundary of the i-th semantic segment. Let be the temporal confidence coefficient of the k-th neighborhood boundary.

[0148] In one embodiment, the semantic fusion processing module 105 is further configured to:

[0149] Based on business semantic data slices, combined with keyword sequences, semantic category labels for the slices are obtained through semantic category mapping;

[0150] Based on slice semantic category labels, the corresponding serverless edge function instance cluster is matched through a preset edge function instance scheduling strategy;

[0151] Business semantic data slices with similar slice semantic category labels are input into a serverless edge function instance cluster, and semantic feature alignment is performed by combining them with temporal association identifiers to obtain an aligned semantic feature set.

[0152] Based on the aligned semantic feature set, the data is reorganized according to the temporal order corresponding to the boundaries of the business semantic segments to obtain the edge broadcast data processing results.

[0153] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the serverless edge broadcast data processing method described above.

[0154] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0155] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0156] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A serverless edge broadcast data processing method, characterized in that, The method includes: Acquire broadcast data streams and extract data temporal features based on the broadcast data streams; Based on the aforementioned data time-series characteristics, a silent threshold boundary identification method is used to identify the set of anchor points to be cut in; Based on the set of anchor points to be cut in and the data time sequence characteristics, the boundaries of business semantic segments are determined through time sequence characteristic pattern analysis. Based on the boundaries of the business semantic segments, the broadcast data stream is semantically segmented into windows to generate business semantic data slices; The business semantic data slices are semantically fused using a serverless edge function instance to obtain the edge broadcast data processing results.

2. The method according to claim 1, characterized in that, Based on the data time-series characteristics, a silent threshold boundary identification method is used to identify the set of anchor points to be entered, including: Obtain the data type information of the broadcast data stream; Using a preset multimodal boundary threshold determination rule, the silent threshold of the data time series characteristics is determined based on the data type information; Based on the data time series characteristics and the silence threshold, a sliding window method is used to filter out time windows in which the data time series characteristics fall within the range of the silence threshold. Based on the time window, the start and end times of the time window are extracted to obtain the silent boundary anchor point; By integrating the silent boundary anchor points, the set of anchor points to be cut in is obtained.

3. The method according to claim 2, characterized in that, The step of determining the boundaries of business semantic segments based on the set of anchor points to be entered and the data temporal characteristics through temporal characteristic pattern analysis includes: Based on each anchor point in the set of anchor points to be cut in, extract time series feature fragments from the time series features of the data; For each of the aforementioned temporal feature segments, keyword sequences and temporal association identifiers are obtained through semantic information extraction; Based on the keyword sequence, the semantic association similarity between each temporal feature segment is calculated; Based on the temporal association identifier, the semantic association similarity is statistically analyzed to obtain the similarity distribution interval; Based on the semantic association similarity and the similarity distribution interval, the temporal feature segments are clustered to obtain feature pattern clusters; Based on each of the aforementioned feature pattern clusters, the boundary of the business semantic segment is determined according to the anchor points corresponding to the set of anchor points to be cut in.

4. The method according to claim 1, characterized in that, The step of semantically segmenting the broadcast data stream based on the business semantic segment boundary to generate business semantic data slices includes: Based on the boundaries of the business semantic segments, the time span data of each semantic segment is obtained by calculating the time difference between adjacent boundaries; Based on the data time series characteristics, data density information is obtained, and based on the data density information, data density verification is performed on the time span data to obtain data density verification results. Based on the data density verification results, the semantic window range is determined for the business semantic segment boundary corresponding to the verified time span data. According to the semantic window range, the broadcast data stream is windowed and segmented to obtain initial data slices; Semantic annotation is performed on the initial data slices to generate the business semantic data slices.

5. The method according to claim 4, characterized in that, The step of performing data density verification on the time span data based on the data density information to obtain the data density verification result includes: Based on the data density information, a spatiotemporal density matrix is ​​constructed; Based on the time span data, the spatiotemporal region to be verified is delineated from the spatiotemporal density matrix. Calculate the density probability value of each spatiotemporal point within the spatiotemporal region to be verified; Based on the density probability value, and according to a preset density fluctuation threshold, abnormal spatiotemporal points where the density probability value exceeds the density fluctuation threshold are identified. The anomalous spatiotemporal points are aggregated to form density anomaly regions; The proportion of the total coverage of all density anomaly intervals in the spatiotemporal region to be verified is calculated. If the proportion is greater than or equal to a preset proportion threshold, the verification is deemed to have failed; otherwise, the verification is deemed to have passed, and the data density verification result is obtained.

6. The method according to claim 4, characterized in that, Based on the boundaries of the business semantic segments, the time span data of each semantic segment is obtained by calculating the time difference between adjacent boundaries, which is achieved through the following formula: in, For the time span data of the i-th semantic segment, Let i be the time sequence of the boundary of the i-th business semantic segment. Let i be the time sequence of the (i+1)th business semantic segment boundary. The temporal association weight corresponding to the i-th semantic segment is... Let be the temporal offset correction coefficient for the i-th semantic segment. Here, m is the time-series decay constant, and m is the radius of the neighborhood boundary sampling window. Let be the temporal mean of the neighborhood of the boundary of the i-th semantic segment. Let be the temporal confidence coefficient of the k-th neighborhood boundary.

7. The method according to claim 3, characterized in that, The step of using a serverless edge function instance to perform semantic fusion processing on the business semantic data slices to obtain edge broadcast data processing results includes: Based on the business semantic data slices and the keyword sequences, semantic category labels for the slices are obtained through semantic category mapping. Based on the slice semantic category label, the corresponding serverless edge function instance cluster is matched through a preset edge function instance scheduling strategy; Business semantic data slices with similar slice semantic category labels are input into the serverless edge function instance cluster, and semantic feature alignment is performed in combination with the temporal association identifier to obtain the aligned semantic feature set. Based on the aligned semantic feature set, the data is reorganized according to the temporal order corresponding to the boundaries of the business semantic segments to obtain the edge broadcast data processing result.

8. A serverless edge broadcast data processing system, characterized in that, The system includes: The data feature extraction module is used to acquire broadcast data streams and extract data time-series features based on the broadcast data streams; Anchor point set identification module is used to identify the anchor point set to be cut in based on the data time sequence characteristics and using the silent threshold boundary identification method; The semantic boundary determination module is used to determine the boundary of the business semantic segment based on the set of anchor points to be cut in and the data temporal characteristics through temporal characteristic pattern analysis. The semantic segment generation module is used to perform semantic windowing on the broadcast data stream based on the business semantic segment boundary to generate business semantic data slices; The semantic fusion processing module is used to perform semantic fusion processing on the business semantic data slices using serverless edge function instances to obtain edge broadcast data processing results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.