Heterogeneous information cross-domain cooperative processing method based on multi-dimensional feature adaptation

By constructing a cross-domain collaborative processing method for heterogeneous information based on multi-dimensional feature adaptation, and using task state vectors to parse the semantics of collaborative tasks and dynamically modulate the temporal alignment strategy, this method solves the problem that static migration and alignment methods in existing technologies cannot adapt to real-time changes in task objectives. It achieves a closed-loop linkage between efficient task semantic understanding and streaming data processing, and improves the targeting and reliability of collaborative processing.

CN121980179APending Publication Date: 2026-05-05LIANYUNGANG CITY PLANNING EXHIBITION CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIANYUNGANG CITY PLANNING EXHIBITION CENT
Filing Date
2026-01-20
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies cannot effectively combine high-level collaborative task semantic understanding with low-level streaming data processing, resulting in an inability to adapt to real-time changes in task objectives in open environments, leading to resource waste and performance deviations.

Method used

By constructing a cross-domain collaborative processing method for heterogeneous information based on multi-dimensional feature adaptation, the method utilizes task state vectors to parse the semantics of collaborative tasks, dynamically modulates temporal alignment strategies, evaluates dynamic feature trust modulus and performs weighted fusion, thereby achieving closed-loop linkage between task semantic understanding and streaming heterogeneous data processing.

Benefits of technology

It improves the targeting and reliability of cross-domain collaborative processing, ensuring that the collaborative processing process always focuses on the core task objectives, meets the low latency requirements, reasonably controls resource utilization, and enhances the engineering practicality and robustness of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980179A_ABST
    Figure CN121980179A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous information cross-domain cooperative processing method based on multi-dimensional feature adaptation, and relates to the technical field of heterogeneous information process.The method comprises the steps that a cooperative task request is received, semantic analysis is conducted, and the task state vector is coded; in response to a task request, receiving a multi-domain streaming heterogeneous data block, dynamically modulating an increment time sequence alignment strategy based on a task state vector, completing time sequence deviation correction and alignment, and outputting a preliminary alignment feature; based on the task state vector and the preliminary alignment feature, evaluating a dynamic feature trust modulus, and performing weighted fusion to generate a task adaptive fusion feature; and generating a collaborative service result, generating a feedback signal according to the difference between the real-time effect index and the expected target, and updating the task state vector code. According to the method, dynamic adaptation of task semantic understanding and streaming data processing is achieved through a double-circulation driving architecture, pertinence, reliability and real-time performance of cooperative processing are improved, and the method is suitable for heterogeneous information cross-domain cooperative processing requirements under multiple scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heterogeneous information processing technology, specifically to a method for cross-domain collaborative processing of heterogeneous information based on multi-dimensional feature adaptation. Background Technology

[0002] With the deep integration of big data and artificial intelligence technologies, data generated by various information systems exhibits typical characteristics of massive scale, heterogeneous sources, streaming formats, and sparse value. Against this backdrop, mining and utilizing heterogeneous information scattered across different fields and systems, and achieving knowledge transfer and value enhancement through cross-domain collaborative processing, has become a key path to improve the efficiency of big data resource services. For example, in smart city governance, effectively coordinating real-time traffic flow data, historical meteorological data, and social media event information is of great significance for achieving accurate anomaly warnings and decision support. However, the inherent differences in data patterns, temporal characteristics, and semantic levels of heterogeneous information pose some challenges to cross-domain collaborative processing.

[0003] Existing technologies have proposed several solutions for different aspects of cross-domain collaboration of heterogeneous information. One approach focuses on improving the accuracy and robustness of cross-domain feature or knowledge transfer. For example, patent publication number CN115757529B discloses a cross-domain commonality transfer recommendation method based on multivariate auxiliary information fusion. This method extracts common features between user domains through a variational autoencoder and uses a self-attention mechanism to generate embedded features for recommendation. The core contribution of this method is to enhance common information transfer by minimizing the mutual information between common and individual features, thereby alleviating the negative transfer problem. However, such methods are usually trained and modeled offline based on historical static datasets. Their feature adaptation and transfer strategies are fixed after model deployment, lacking the ability to perceive and respond to dynamically evolving real-time collaborative task intentions. Another approach can solve the alignment problem of multi-source heterogeneous data at the structural level. For example, patent publication number CN116050374A discloses a cross-domain and cross-source data alignment method, which innovatively integrates the text representation and visual position information of tabular data and determines the alignment result by calculating the semantic distance of multimodal vector expressions. This method improves the accuracy of alignment, but its processing paradigm is essentially geared towards static, batch-processed tabular data. Its alignment process is independent of the specific collaborative task objectives at the upper layer and does not consider streaming scenarios where data arrives continuously and at high speed, making it difficult to support the need for low-latency real-time collaboration.

[0004] In summary, a current technical bottleneck in cross-domain collaborative processing of heterogeneous information lies in the failure of existing solutions to organically integrate and achieve closed-loop linkage between high-level, dynamic semantic understanding of collaborative tasks and low-level, dynamic streaming data processing. Specifically, static migration and alignment methods cannot adapt to the real-time changes in task objectives in open environments; while streaming processing lacking task semantic guidance makes it difficult to ensure that the collaborative process remains focused on the core objective, easily leading to resource waste and performance deviations. Therefore, a new technical solution is needed that enables the system to dynamically understand and adapt to the constantly evolving collaborative task objectives while continuously processing streaming heterogeneous data.

[0005] In view of this, the present invention aims to solve the above-mentioned technical problem of the separation between task awareness and data stream processing. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a cross-domain collaborative processing method for heterogeneous information based on multi-dimensional feature adaptation. By parsing the semantics of collaborative tasks to generate task state vectors, constructing a dual-loop driven architecture, dynamically modulating the temporal alignment strategy, evaluating the dynamic trust modulus of features and weighted fusion, it can achieve closed-loop linkage between task semantic understanding and streaming heterogeneous data processing, adapt to dynamically evolving task objectives, overcome the shortcomings of existing technologies that are static and lack task guidance, and improve the pertinence, reliability and real-time performance of cross-domain collaborative processing.

[0007] To address the aforementioned technical problems, this invention provides the following technical solution: a cross-domain collaborative processing method for heterogeneous information based on multi-dimensional feature adaptation, comprising the following steps:

[0008] Step 1: Receive a collaborative task request, perform semantic parsing on the collaborative task request, extract task semantic elements, and encode the task semantic elements into a task state vector, wherein the task state vector is a differentiable multidimensional vector.

[0009] Step 2: In response to the collaborative task request, receive streaming heterogeneous data blocks from at least two different domains in real time. Based on the task state vector, dynamically modulate the incremental timing alignment strategy for the streaming heterogeneous data blocks. According to the modulated incremental timing alignment strategy, perform timing deviation correction and alignment on the streaming heterogeneous data blocks, and output preliminary alignment features.

[0010] Step 3: Based on the task state vector and the preliminary alignment features, evaluate the dynamic feature trust modulus of the preliminary alignment features, and perform weighted fusion of the preliminary alignment features from different domains according to the dynamic feature trust modulus to generate task adaptation fusion features.

[0011] Step 4: Generate collaborative service results based on the task adaptation and fusion features, and generate task utility feedback signals based on the real-time effect indicators of the collaborative service results and the expected goals of the collaborative task requests. Use the task utility feedback signals to update the encoding of the task state vector.

[0012] Furthermore, in step one, the collaborative task request is semantically parsed to extract task semantic elements, specifically including: identifying and extracting the task subject, associated domain, task objective and task constraints in the collaborative task request. The task state vector at least contains implicit expressions of data freshness requirements, feature reliability requirements and domain weight bias.

[0013] Furthermore, in step two, the dynamic modulation of the incremental timing alignment strategy based on the task state vector specifically includes: decoding the real-time bias coefficient and the accuracy bias coefficient from the task state vector; selecting a target alignment strategy from a set of preset alignment strategies based on the ratio between the real-time bias coefficient and the accuracy bias coefficient; wherein the preset alignment strategies include at least an aggressive interpolation strategy that prioritizes low latency and a buffered fine alignment strategy that prioritizes alignment accuracy.

[0014] Furthermore, in step three, the dynamic feature trust modulus of the preliminary alignment feature is evaluated, specifically calculated using the following formula:

[0015]

[0016] in, This represents the dynamic feature trust modulus of the i-th feature at time t. This represents the sigmoid activation function. Represents the task state vector at time t. Let represent the initial aligned feature vector of the i-th feature at time t. It is the task relevance moderating coefficient. This represents the quantized value of the data uncertainty of the i-th feature at time t. This represents the quantized value of the model uncertainty for the i-th feature at time t. and These are the weighting coefficients for data uncertainty and model uncertainty, respectively.

[0017] The dynamic feature trust modulus is positively adjusted by the correlation between the task state vector and the feature, and negatively adjusted by the data and model uncertainty of the feature itself.

[0018] Furthermore, in step three, the preliminary alignment features are weighted and fused according to the dynamic feature trust modulus. Specifically, the dynamic feature trust modulus of each feature is normalized to obtain a fusion weight, and the fusion weight is used to perform a weighted summation on the corresponding preliminary alignment features to generate the task adaptation fusion feature.

[0019] Furthermore, in step four, the task utility feedback signal is used to update the encoding of the task state vector. This is specifically achieved through the following strategy: the task utility feedback signal and the task state vector are input together into a task encoder neural network. The task encoder neural network takes minimizing the negative value of the task utility feedback signal as its optimization objective and adjusts its network parameters through a backpropagation algorithm, so that the task encoder neural network outputs an updated task state vector for the same or similar collaborative task requests.

[0020] The task encoder neural network adopts a multilayer perceptron structure with an attention mechanism.

[0021] Furthermore, in step two, the modulation of the incremental temporal alignment strategy and the evaluation and weighted fusion of the dynamic feature trust modulus in step three constitute a data flow adaptive inner loop.

[0022] The generation of the task state vector in step one and the updating of the task state vector based on the task utility feedback signal in step four constitute a task strategy optimization outer loop.

[0023] The adaptive inner loop of the data flow is dynamically controlled by the outer loop of the task strategy optimization. The outer loop of the task strategy optimization is optimized based on the running effect of the adaptive inner loop of the data flow, and the two constitute a dual-loop driven architecture.

[0024] Furthermore, in step two, the timing deviation correction and alignment specifically includes: attaching a high-precision timestamp to each of the streaming heterogeneous data blocks; and estimating missing or delayed data points based on the high-precision timestamp and the current system timing using a lightweight timing prediction model, wherein the lightweight timing prediction model is a linear Kalman filter.

[0025] Furthermore, the different domains include data sources with different physical origins, different data formats, or different logical concepts;

[0026] The streaming heterogeneous data block includes at least two of the following: text data, numerical sequence data, and image data.

[0027] Furthermore, the real-time performance metrics of the collaborative service results include at least one of prediction accuracy, response latency, and resource utilization.

[0028] The task utility feedback signal is a quantified function value of the difference between the real-time performance indicator and the expected target.

[0029] Compared with existing technologies, this heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation has the following advantages:

[0030] I. This invention performs semantic parsing on collaborative task requests, extracts task semantic elements, and encodes them into differentiable task state vectors. It constructs a dual-loop driven architecture consisting of an adaptive data flow inner loop and a task strategy optimization outer loop. The task strategy optimization outer loop dynamically updates the task state vector based on the collaborative service effect, thereby regulating the temporal alignment strategy modulation and feature fusion process of the adaptive data flow inner loop. This achieves the organic integration and closed-loop linkage between high-level collaborative task semantic understanding and low-level streaming data processing. It effectively solves the problems that existing static migration and alignment methods cannot adapt to real-time changes in task objectives and that streaming processing without task semantic guidance is prone to resource waste and effect deviation. It ensures that the collaborative processing process always focuses on the core task objective and significantly improves the pertinence and effectiveness of cross-domain collaborative processing of heterogeneous information.

[0031] Second, this invention designs a dynamic feature trust modulus evaluation mechanism, which comprehensively considers the correlation between the task state vector and the features, as well as the data and model uncertainties of the features themselves. It performs weighted fusion of preliminary aligned features from different domains, which can accurately filter low-quality features and improve the quality of task-adapted fused features. At the same time, it adopts a lightweight temporal prediction model to correct temporal deviations. Combined with an efficient task encoder neural network and a collaborative service model, it effectively meets the low latency requirements of streaming data processing while ensuring the reliability of collaborative service results. It also reasonably controls the system resource utilization rate, enhances the engineering practicality and robustness of the method, and enables it to stably adapt to the cross-domain collaborative processing needs of heterogeneous information in multiple scenarios.

[0032] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0034] Figure 1This is a diagram illustrating the method steps of the present invention;

[0035] Figure 2 This is a schematic diagram of the dual-loop drive architecture of the present invention;

[0036] Figure 3 This is a schematic diagram of the multi-domain streaming heterogeneous data processing flow of the present invention. Detailed Implementation

[0037] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0038] Example

[0039] like Figures 1 to 3 As shown, this embodiment aims to elaborate in detail the complete implementation process of the heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation. This embodiment takes the smart city traffic anomaly early warning collaborative service scenario as a specific application scenario. In this scenario, it is necessary to collaboratively process streaming heterogeneous data from traffic monitoring domain, meteorological monitoring domain, and social media domain to achieve accurate early warning of abnormal events such as urban road congestion and traffic accidents.

[0040] In this embodiment, the data sources for different domains are specifically set as follows: the traffic monitoring domain outputs numerical sequence data, including traffic flow, average vehicle speed, and lane occupancy rate for each road segment every 5 minutes; the meteorological monitoring domain outputs heterogeneous data combining text data and numerical sequence data, with the text data being meteorological warning information and the numerical sequence data being rainfall, visibility, and wind force every 10 minutes; the social media domain outputs text data, including user-posted complaints, requests for help, and event broadcasts related to road conditions. The collaborative task request is "to monitor abnormal traffic events in the core urban area within the next 30 minutes in real time, requiring a warning response delay of no more than 2 minutes, a prediction accuracy of no less than 85%, and a resource utilization rate controlled below 70%."

[0041] This embodiment uses task state vectors to run through the entire collaborative processing flow, constructing a dual-loop driven architecture of an adaptive data flow inner loop and a task strategy optimization outer loop, to achieve the organic integration of task semantic understanding and streaming data processing. The specific implementation process of each step will be described in detail below.

[0042] In this embodiment, step one, the specific implementation of task state vector generation, is as follows:

[0043] The core of this step is to receive the collaborative task request, extract the semantic elements of the task through semantic parsing, and encode them into a differentiable multidimensional task state vector, providing a core basis for subsequent temporal alignment strategy modulation and dynamic feature trust modulus evaluation.

[0044] Collaborative task request received:

[0045] In this embodiment, collaborative task requests are received through the interface of the smart city collaborative service platform. This interface adopts a RESTful architecture and supports high-concurrency task request transmission. The receiving module has a built-in request validity verification mechanism to verify the format and signature information of the request, filter invalid requests, and ensure the stability of subsequent processing. After successful verification, the collaborative task request is transmitted to the semantic parsing module.

[0046] Task semantic element extraction:

[0047] The semantic parsing module adopts a Transformer-based semantic understanding model, which has been trained on a large-scale Chinese task request corpus. This model can accurately identify and extract the task semantic elements in collaborative task requests, including the task subject, related domain, task objective, and task constraints.

[0048] Task subject: refers to the core object targeted by the collaborative task. In this embodiment, the collaborative task request "real-time monitoring of abnormal traffic events in the core urban area within the next 30 minutes, requiring early warning response delay not to exceed 2 minutes, prediction accuracy not to be less than 85%, and resource utilization rate controlled below 70%" is analyzed by the model, and the task subject is extracted as "abnormal traffic events in the core urban area within the next 30 minutes".

[0049] Related domains: refer to the data source domains that are required to complete collaborative tasks. In this embodiment, the model identifies the related domains as traffic monitoring domain, meteorological monitoring domain, and social media domain.

[0050] Task objective: refers to the core effect that the collaborative task is expected to achieve. In this embodiment, the task objective is "to achieve real-time early warning of traffic anomalies".

[0051] Task constraints refer to the specific requirements for the collaborative service results. In this embodiment, task constraints include response latency not exceeding 2 minutes, prediction accuracy not less than 85%, and resource utilization controlled below 70%.

[0052] During the extraction process, the model employs a collaborative approach involving sub-modules such as entity recognition, relation extraction, and keyword extraction. The entity recognition module identifies entities such as "core urban area," "30 minutes," and "traffic anomaly events." The relation extraction module establishes the relationships between entities and constraints, such as the constraint relationship between "traffic anomaly event warning" and "response delay," "prediction accuracy," and "resource utilization." The keyword extraction module extracts core keywords such as "real-time monitoring," "warning," "response delay," "prediction accuracy," and "resource utilization," providing a foundation for subsequent task state vector encoding.

[0053] Task state vector encoding:

[0054] The extracted task semantic elements are input into the task encoder neural network, which adopts a multilayer perceptron structure with an attention mechanism and outputs a differentiable multidimensional task state vector.

[0055] The specific structure of the task encoder neural network is as follows: the input layer dimension is the dimension of the extracted task semantic feature, which is 64 dimensions after feature encoding in this embodiment; three hidden layers are set, with 128, 256 and 128 neurons in each layer, respectively, and the ReLU function is used for activation; the attention mechanism layer adopts the scaled dot product attention mechanism, which can assign different attention weights to task semantic features of different importance; the output layer dimension is 32 dimensions, that is, the dimension of the task state vector is 32 dimensions, and the output layer activation function adopts the linear activation function to ensure the differentiability of the task state vector.

[0056] Each dimension of the task state vector corresponds to an implicit expression of the task semantic elements, including at least implicit expressions of data freshness requirements, feature reliability requirements, and domain weight bias. In this embodiment, the first 8 dimensions of the task state vector correspond to the implicit expression of data freshness requirements. The larger the value, the higher the requirement for data freshness. Since this collaborative task requires a response delay of no more than 2 minutes, the requirement for data freshness is relatively high. Therefore, the values ​​of the first 8 dimensions are generally in the range of 0.7-0.9. The middle 12 dimensions correspond to the implicit expression of feature reliability requirements. The larger the value, the higher the requirement for feature reliability. This task requires a prediction accuracy of no less than 85%, the requirement for feature reliability is relatively high. Therefore, the values ​​of the middle 12 dimensions are generally in the range of 0.75-0.95. The last 12 dimensions correspond to the implicit expression of domain weight bias, which correspond to the weight bias of the traffic monitoring domain, the meteorological monitoring domain, and the social media domain, respectively. Since the data in the traffic monitoring domain is directly related to traffic anomalies, its corresponding dimension value is the highest, in the range of 0.8-0.95. The meteorological monitoring domain is second, in the range of 0.6-0.8. The social media domain is relatively low, in the range of 0.4-0.6.

[0057] The training process of the task encoder neural network is as follows: Stochastic gradient descent optimization is used, with a learning rate of 0.001, a batch size of 32, and 1000 training iterations. The training dataset consists of historical collaborative task requests and their corresponding labeled task state vectors. The labeled task state vectors are manually annotated by domain experts based on task semantic elements. During training, the optimization objective is to minimize the mean squared error between the predicted task state vector and the labeled task state vector. The network parameters are continuously adjusted using the backpropagation algorithm until the model converges.

[0058] In this embodiment, step two, the specific implementation of streaming heterogeneous data time-series alignment, is as follows:

[0059] This step responds to collaborative task requests, receives streaming heterogeneous data blocks from multiple different domains in real time, dynamically modulates incremental timing alignment strategies based on task state vectors, completes timing deviation correction and alignment, and outputs preliminary alignment features.

[0060] Streaming heterogeneous data reception:

[0061] The Kafka streaming data receiving framework is used to achieve real-time reception of heterogeneous streaming data blocks from traffic monitoring, weather monitoring, and social media domains. Kafka topics are divided according to different domains, with separate topics created for traffic monitoring, weather monitoring, and social media. Each topic has 8 partitions and 3 replicas to ensure high throughput and high reliability of data reception.

[0062] The numerical sequence data in the traffic monitoring domain is transmitted in JSON format, with each data entry containing fields such as road segment identification, collection time, traffic flow, average vehicle speed, and lane occupancy rate. The text data in the meteorological monitoring domain is transmitted in XML format, containing fields such as monitoring station identification, release time, weather warning type, and warning level. The numerical sequence data is transmitted in CSV format, containing fields such as monitoring station identification, collection time, rainfall, visibility, and wind force. The text data in the social media domain is transmitted in JSON format, containing fields such as posting user identification, posting time, text content, and geolocation tags.

[0063] The data receiving module performs preliminary preprocessing on the received streaming heterogeneous data blocks, including data format parsing, field validation, and data cleaning. The data format parsing module calls the corresponding parser to parse the data based on the data format type of different fields; the field validation module checks whether data fields are complete and whether field types are correct, filtering out data with missing fields or incorrect types; the data cleaning module removes duplicate and abnormal data to ensure the validity of data for subsequent processing.

[0064] Incremental timing alignment strategy modulation:

[0065] The incremental timing alignment strategy based on dynamic modulation of the task state vector is as follows:

[0066] First, decode the real-time emphasis coefficient and the accuracy emphasis coefficient from the task status vector. The decoding process is implemented by a fully connected neural network. The input of this fully connected neural network is a 32-dimensional task status vector, and the output is a 2-dimensional vector, corresponding to the real-time emphasis coefficient and the accuracy emphasis coefficient respectively. The number of neurons in the hidden layer of the fully connected neural network is 64, the ReLU function is used as the activation function, and the sigmoid activation function is used in the output layer to ensure that the values of the real-time emphasis coefficient and the accuracy emphasis coefficient are within the range of 0-1.

[0067] The decoding principle is as follows: The dimensional features related to the data freshness requirement in the task status vector are mapped through the fully connected neural network and transformed into the real-time emphasis coefficient; the dimensional features related to the feature reliability requirement are transformed into the accuracy emphasis coefficient through mapping. In this embodiment, the decoded real-time emphasis coefficient is 0.7, and the accuracy emphasis coefficient is 0.8.

[0068] Then, according to the proportional relationship between the real-time emphasis coefficient and the accuracy emphasis coefficient, select the target alignment strategy from a variety of preset alignment strategies. The preset alignment strategies include the active interpolation strategy that gives priority to ensuring low latency and the buffered fine alignment strategy that gives priority to ensuring alignment accuracy. The core parameters of the two strategies are as follows:

[0069] Active interpolation strategy: Use the linear interpolation method to fill in the missing or delayed data points. The interpolation window size is set to 5 data points, and the data caching time is set to 10 seconds. That is, when the data delay exceeds 10 seconds, linear interpolation is directly used for filling to ensure low latency in the alignment process.

[0070] Buffered fine alignment strategy: Use the cubic spline interpolation method to fill in the missing or delayed data points. The interpolation window size is set to 10 data points, and the data caching time is set to 30 seconds. That is, the data is allowed to be cached for 30 seconds, and fine interpolation is performed after the delayed data arrives to ensure alignment accuracy.

[0071] The selection rule for the alignment strategy is: Calculate the ratio R of the real-time emphasis coefficient to the accuracy emphasis coefficient. When R≥1.2, select the active interpolation strategy; when R≤0.8, select the buffered fine alignment strategy; when 0.8<R<1.2, perform alignment according to the weighted combination of the two strategies, and the weighting coefficients are the real-time emphasis coefficient and the accuracy emphasis coefficient respectively.

[0072] In this embodiment, the real-time weighting coefficient is 0.7, the accuracy weighting coefficient is 0.8, and the ratio R = 0.7 / 0.8 = 0.875, which is in the range of 0.8 < R < 1.2. Therefore, a weighted combination of the two strategies is used for alignment. The weighting coefficient of the active interpolation strategy is 0.7, and the weighting coefficient of the buffered fine alignment strategy is 0.8. The interpolation window size of the alignment strategy after weighted combination is 5×0.7 + 10×0.8 = 11.5, which is rounded up to 12 data points; the data caching time is 10×0.7 + 30×0.8 = 31 seconds.

[0073] Timing deviation correction and alignment:

[0074] The core of timing deviation correction and alignment is to attach a high-precision timestamp to each streaming heterogeneous data block. Based on the high-precision timestamp and the current system timing, a lightweight timing prediction model is used to estimate missing or delayed data points and output preliminary alignment features.

[0075] First, attach a high-precision timestamp to each streaming heterogeneous data block. The Network Time Protocol (NTP) is used to achieve time synchronization with a time synchronization accuracy reaching the millisecond level. When the data receiving module receives each data block, it records the current system time after NTP synchronization as the high-precision timestamp of this data block. The timestamp format is YYYY-MM-DD HH:MM:SS.fff, where fff represents milliseconds.

[0076] Then, perform timing deviation detection. Calculate the difference between the high-precision timestamp of each data block and the current system timing. This difference is the timing deviation. When the difference is positive, it means the data block arrives early without delay; when the difference is negative, it means the data block arrives late, and the delay time is the absolute value of the difference; when the data block corresponding to a certain moment is not received, it is determined as data missing.

[0077] For delayed or missing data points, a linear Kalman filter is used for estimation. The linear Kalman filter is a lightweight timing prediction model with low computational complexity and good real-time performance, suitable for real-time processing of streaming data.

[0078] [[ID=^18]]The state equation and observation equation of the linear Kalman filter are as follows:

[0079] State equation: ;

[0080] Observation equation: ;

[0081] [[ID=^31]]Where, : the system state vector at time k. In this embodiment, the system state vector includes the data value and its change rate, that is where The data value at time k. Let k be the rate of change of the data value at time k.

[0082] The state transition matrix describes the state transition from time k-1 to time k. In this embodiment... ,in For the data collection cycle, the traffic monitoring domain Seconds (5 minutes), numerical sequence data of meteorological monitoring domain Seconds (10 minutes), text data Adjusted dynamically based on the release interval.

[0083] : Input control matrix. In this embodiment, there is no external input control, therefore .

[0084] The external input vector at time k-1, in this embodiment .

[0085] The process noise at time k-1 follows a mean of 0 and a variance of . Gaussian distribution, The process noise covariance matrix is ​​shown in this embodiment. This value is determined through statistical analysis of historical data and can reflect the uncertainty of data changes.

[0086] : The observation value at time k, that is, the actual data value received at time k.

[0087] The observation matrix is ​​used to map the system state vector to the observation space. In this embodiment... That is, the observed values ​​are only related to the data values ​​in the system state vector.

[0088] The observation noise at time k follows a pattern with a mean of 0 and a variance of . Gaussian distribution, In this embodiment, to observe the noise covariance matrix, This value is determined based on factors such as sensor accuracy and data transmission error.

[0089] The iterative process of a linear Kalman filter is as follows:

[0090] Prediction steps:

[0091]

[0092]

[0093] in, This is the predicted state vector at time k based on the observations at time k-1. Let be the covariance matrix of the predicted state vector at time k. Let be the optimal estimated state vector at time k-1. Let be the covariance matrix of the optimal estimated state vector at time k-1.

[0094] Update steps:

[0095]

[0096]

[0097]

[0098] in, The Kalman gain at time k is used to balance the confidence of the predicted state vector and the observation. It is the identity matrix. Let be the optimal estimated state vector at time k. Let be the covariance matrix of the optimal estimated state vector at time k.

[0099] For delayed data points, once the delayed data arrives, it is used as the observation value. Substituting the values ​​into the Kalman filter update step, the previously predicted state vector is corrected; for missing data points, the values ​​obtained in the prediction step are used. Data values ​​in As an estimate.

[0100] Through the aforementioned time-series deviation correction and alignment process, heterogeneous streaming data blocks from different domains and different acquisition periods are aligned to a unified time axis. The time granularity of the time axis is set to 1 minute, meaning each time point corresponds to a preliminary alignment feature. The dimension of the preliminary alignment feature is the sum of the dimensions of the data features from each domain. In this embodiment, the feature dimension of the traffic monitoring domain is 3, the feature dimension of the meteorological monitoring domain is 4, and the feature dimension of the social media domain is 5. Therefore, the total dimension of the preliminary alignment feature is 3 + 4 + 5 = 12 dimensions. Each dimension corresponds to a specific feature indicator, and each feature indicator corresponds one-to-one with a time point on the unified time axis.

[0101] In this embodiment, step three, the specific implementation of task adaptation fusion feature generation, is as follows:

[0102] This step evaluates the dynamic feature trust modulus of each feature based on the task state vector and the initial alignment features, and then performs weighted fusion of the initial alignment features according to the dynamic feature trust modulus to generate task-adaptive fusion features.

[0103] Dynamic feature trust modulus evaluation:

[0104] The dynamic feature trust modulus is calculated using the following formula:

[0105]

[0106] in, : The dynamic feature trust modulus of the i-th feature at time t, used to quantify the degree of trust of the i-th feature in the current collaborative task at time t. The value ranges from 0 to 1. The larger the value, the higher the contribution of the feature to the collaborative task and the stronger the credibility.

[0107] The sigmoid activation function has the following expression: Its function is to map the input value to the range of 0-1, so that the value of the dynamic feature trust modulus conforms to the probability distribution characteristics, which facilitates the subsequent weighted fusion calculation.

[0108] The task relevance adjustment coefficient is used to adjust the weight of the correlation between the task state vector and the initial alignment feature vector. Its value is determined based on the type of collaborative task and the importance of the features. This embodiment verifies this value through historical data experiments. .when When the value is large, the correlation between the task state vector and the features has a more significant impact on the dynamic feature trust modulus; when When the value is small, the effect is relatively weakened.

[0109] : The task state vector at time t, i.e. the 32-dimensional differentiable vector generated in step (I). The meaning and value range of each dimension have been explained in detail in step (I), and will not be repeated here.

[0110] : The initial aligned feature vector of the i-th feature at time t. In this embodiment, the total dimension of the initial aligned features is 12, so the value of i ranges from 1 to 12. It is a 1-dimensional vector, and its value is determined according to the timing alignment result in step (II). For example, when i=1, This represents the traffic flow characteristic value of the traffic monitoring domain at time t.

[0111] The data uncertainty weighting coefficient is used to adjust the negative adjustment of the data uncertainty quantification value on the dynamic feature trust modulus. Its value ranges from 0 to 1 and is determined through historical data statistical analysis and experimental verification. In this embodiment... . The larger the value, the stronger the inhibitory effect of data uncertainty on the dynamic feature trust modulus; conversely, the weaker the effect.

[0112] : The data uncertainty quantification value of the i-th feature at time t, used to quantify the data quality uncertainty of the i-th feature at time t. The value ranges from 0 to 1; a larger value indicates worse data quality and higher uncertainty. Data uncertainty mainly stems from factors such as data acquisition errors, transmission noise, and missing data. Its calculation method is as follows:

[0113]

[0114] Where n is the size of the statistical window, and in this embodiment n=10; Let tj be the actual observed value of the i-th feature; Let be the predicted value of the i-th feature at time tj (obtained through a linear Kalman filter); This is the minimum value, taking the value 1e-6, used to avoid the denominator being 0; Let be the maximum value between the actual observed value and the predicted value of the i-th feature at time tj. This is calculated using the formula... It can reflect the degree of fluctuation and reliability of feature data; the greater the fluctuation, the lower the reliability. The larger the value, the better.

[0115] The model uncertainty weight coefficient is used to adjust the negative adjustment of the model uncertainty quantification value on the dynamic feature trust modulus. Its value ranges from 0 to 1 and is determined through historical data statistical analysis and experimental verification. In this embodiment... . The larger the value, the stronger the suppression effect of model uncertainty on the dynamic feature trust modulus; conversely, the weaker the effect.

[0116] : The model uncertainty quantification value for the i-th feature at time t, used to quantify the model prediction uncertainty of the i-th feature at time t. The value ranges from 0 to 1; a larger value indicates lower prediction reliability and higher uncertainty for that feature. Model uncertainty mainly stems from factors such as the randomness of model parameters and limitations of the model structure. Its calculation method is as follows:

[0117]

[0118] in, The first row and first column element of the covariance matrix of the optimal estimated state vector obtained by the linear Kalman filter for the i-th feature at time t reflects the variance of the model's estimation of the feature data value. The threshold value is 1e-4, used to prevent excessively small variance from causing... Too large; This is the minimum value, taking the value 1e-6, to avoid a denominator of 0. The result is obtained through this formula. It can reflect the reliability of the model's estimation of feature data; the larger the variance, the lower the reliability. The larger the value, the better.

[0119] The following section uses the first feature at time t=2024-05-2010:00:00 in this embodiment as an example to explain in detail the calculation process of the dynamic feature trust modulus:

[0120] Given: , This is a 32-dimensional task state vector, with the following specific values:

[0121] [0.82,0.85,0.78,0.81,0.79,0.83,0.80,0.77,0.88,0.92,0.85,0.90,0.87,0.91,0.86,0.89,0.84,0.76,0.72,0.68,0.75,0.79,0.65,0.71,0.58,0.52,0.49,0.55,0.47,0.53,0.42,0.48];

[0122] (Traffic flow characteristic value of the first feature at time t);

[0123] (Obtained through data uncertainty calculation methods);

[0124] (Obtained through model uncertainty calculation methods);

[0125] First calculate :because It is a 32-dimensional vector. Since it is a scalar, therefore It is a 32-dimensional vector, with each element being... corresponding element multiplied ,Right now:

[0126] [0.82×120,0.85×120,...,0.48×120]=

[0127] [98.4,102.0,93.6,97.2,94.8,99.6,96.0,92.4,105.6,110.4,102.0,108.0,104.4,109.2,103.2,106.8,100.8,91.2,86.4,81.6,90.0,94.8,78.0,85.2,69.6,62.4,58.8,66.0,56.4,63.6,50.4,57.6];

[0128] Then calculate Multiply each element of the above 32-dimensional vector by ,get

[0129] [98.4×0.8,102.0×0.8,...,57.6×0.8]=

[0130] [78.72,81.6,74.88,77.76,75.84,79.68,76.8,73.92,84.48,88.32,81.6,86.4,83.52,87.36,82.56,85.44,80.64,72.96,69.12,65.28,72.0,75.84,62.4,68.16,55.68,49.92,47.04,52.8,45.12,50.88,40.32,46.08];

[0131] Next calculation ;

[0132] Then calculate Subtracting 0.045 and 0.02 from each element of the above 32-dimensional vector yields:

[0133] [78.72-0.045-0.02,81.6-0.045-0.02,...,46.08-0.045-0.02]=

[0134] [78.655,81.535,74.815,77.715,75.775,79.615,76.735,73.855,84.415,88.255,81.535,86.335,83.455,87.295,82.495,85.375,80.575,72.895,69.055,65.215,71.935,75.775,62.335,68.095,55.615,49.855,46.975,52.735,45.055,50.815,40.255,45.995];

[0135] Since the dynamic feature trust modulus is a scalar, the above 32-dimensional vector needs to be aggregated using a weighted summation method, with the weights being the normalized values ​​of each dimension in the task state vector. The normalization process of the task state vector is as follows:

[0136]

[0137] in, The j-th element of the task state vector. Let be the normalized weight of the j-th element.

[0138] calculate =

[0139] 0.82+0.85+0.78+0.81+0.79+0.83+0.80+0.77+0.88+0.92+0.85+0.90+0.87+0.91+0.86+0.89+0.84+0.76+0.72+0.68+0.75+0.79+0.65+0.71+0.58+0.52+0.49+0.55+0.47+0.53+0.42+0.48=24.36;

[0140] Then the normalized weights of each dimension for For example, the weights of the first dimension The weights of the second dimension And so on.

[0141] Multiply each element of the 32-dimensional vector by its corresponding normalized weight, then sum them to obtain the aggregated value. :

[0142]

[0143] Substituting the numerical values, we obtain ;

[0144] Finally, Substituting into the sigmoid activation function, we obtain the dynamic feature trust modulus. (Because the value of x is relatively large, the output of the sigmoid function approaches 1.0)

[0145] Following the same method described above, the dynamic feature confidence modulus of the 2nd to 12th features at time t is calculated sequentially. Assume the calculation results are as follows: .

[0146] Preliminary alignment feature weighted fusion:

[0147] The initial alignment features are weighted and fused based on the dynamic feature trust modulus. The specific process is as follows:

[0148] First, the dynamic feature trust modulus of each feature is normalized to obtain the fusion weights. The normalization process uses the softmax function, as shown in the following formula:

[0149]

[0150] in, Let be the fusion weight of the i-th feature at time t. In this embodiment, the total dimension of the initial alignment features is used. , Let be the dynamic feature trust modulus of the j-th feature at time t.

[0151] Taking the dynamic feature trust modulus at time t in this embodiment as an example, the fusion weights are calculated as follows:

[0152]

[0153]

[0154] ;

[0155] Calculate the values ​​for each item:

[0156]

[0157]

[0158]

[0159] Summation yields ≈

[0160] 2.7183+2.6645+2.5857+2.3396+2.2703+2.1815+2.1170+2.0544+1.9744+1.9155+1.8586+1.7860≈25.4668;

[0161] The fusion weights for each feature are then:

[0162] ;

[0163] ;

[0164] ;

[0165] ;

[0166] ;

[0167] ;

[0168] ;

[0169] ;

[0170] ;

[0171] ;

[0172] ;

[0173] .

[0174] The sum of the fusion weights is ≈

[0175] 0.1067+0.1046+0.1015+0.0919+0.0891+0.0856+0.0831+0.0807+0.0775+0.0752+0.0730+0.0701≈1.0, which meets the normalization requirements.

[0176] Then, the corresponding preliminary alignment features are weighted and summed using fusion weights to generate task-adaptive fusion features. The dimension of the task-adaptive fusion features can be set according to actual needs; in this embodiment, it is set to 8 dimensions. The 12-dimensional preliminary alignment features are mapped to the 8-dimensional task-adaptive fusion features through a fully connected neural network, and the fusion formula is as follows:

[0177]

[0178] in, The task at time t is adapted and fused with 8 dimensions; This is a weight matrix with dimensions 8×12; It is the bias vector with 8 dimensions.

[0179] weight matrix and bias vector The training dataset consisted of historical preliminary alignment features, fusion weights, and corresponding annotation task-adapted fusion features. These annotation task-adapted fusion features were manually labeled by domain experts based on the collaborative task objectives. The training process employed a stochastic gradient descent optimization algorithm with a learning rate of 0.001, a batch size of 32, and 500 training iterations. The optimization objective was to minimize the mean squared error between the prediction task-adapted fusion features and the annotation task-adapted fusion features.

[0180] In this embodiment, the weight matrix is ​​determined through training. The possible values ​​are as follows (for simplicity, only the elements in the first 3 rows and 3 columns are listed):

[0181]

[0182] bias vector The possible values ​​are as follows:

[0183]

[0184] Substituting the fusion weights and preliminary alignment features at time t into the above formula, the task adaptation fusion features at time t can be calculated. For example, assuming the initial alignment feature vector at time t is [120, 50, 0.8, 15, 500, 3, 2, 0.9, 0.7, 0.6, 0.5, 0.4] (corresponding to the values ​​of 12 features respectively), then the weighted initial alignment feature vector is [0.1067×120, 0.1046×50, 0.1015×0.8, 0.0919×15, 0.0891×500, 0.0856]. ×3, 0.0831×2, 0.0807×0.9, 0.0775×0.7, 0.0752×0.6, 0.0730×0.5, 0.0701×0.4]=[12.804, 5.23, 0.0812, 1.3785, 44.55, 0.2568, 0.1662, 0.0726, 0.0543, 0.0451, 0.0365, 0.0280]

[0185] The weighted initial alignment feature vector and weight matrix are then compared. Multiply, and add the bias vector To obtain task adaptation and fusion features Its specific value is

[0186] [1.85,1.52,1.28,1.05,0.92,0.78,0.65,0.52] (The specific values ​​are calculated based on the complete values ​​of the weight matrix and the bias vector).

[0187] In this embodiment, step four, the generation of collaborative service results and the updating of task state vectors, are implemented as follows:

[0188] This step generates collaborative service results based on task adaptation and fusion features. It generates task utility feedback signals by comparing real-time performance indicators with expected goals, and uses these feedback signals to update the encoding of the task state vector, thus forming a closed-loop optimization.

[0189] Generation of collaborative service results:

[0190] The task-adapted fusion features are input into the collaborative service model to generate collaborative service results. In this embodiment, the collaborative service model is a traffic anomaly early warning model, constructed using the random forest algorithm. This algorithm is characterized by strong robustness, good generalization ability, and ease of interpretation, making it suitable for classification tasks.

[0191] The core parameters of the random forest algorithm are set as follows: 100 decision trees; maximum depth of each decision tree is set to 10; maximum number of features considered when splitting each node is 4; minimum number of samples in a leaf node is 5; and random seed is set to 42 to ensure the repeatability of model training.

[0192] The training process of the Random Forest algorithm is as follows: The training dataset consists of historical task-adapted fusion features and corresponding traffic anomaly event annotations. The training dataset is divided into a training set and a validation set in a 7:3 ratio. The training set is used to train the model, and the validation set is used to validate the model's performance. During training, multiple sample sets are extracted from the training set using the bootstrap sampling method. Each sample set is used to train a decision tree. Each decision tree is constructed using the CART algorithm, and the Gini coefficient is used as the splitting criterion when splitting nodes, selecting the feature with the smallest Gini coefficient for splitting. After training, the model is evaluated using the validation set. The model training is considered complete when the prediction accuracy on the validation set reaches 85% or higher.

[0193] Adapt and fuse features for the task at time t The data is input into a trained random forest model, which outputs a traffic anomaly warning result, either "abnormal" or "normal". In this embodiment, the model output is assumed to be "abnormal", indicating that a traffic anomaly event is predicted to occur in the core urban area within the next 30 minutes.

[0194] Real-time performance metrics calculation:

[0195] The real-time performance metrics for collaborative service results include prediction accuracy, response latency, and resource utilization, calculated as follows:

[0196] Prediction accuracy: This refers to the ratio of correctly predicted traffic anomalies to the total number of predicted incidents within a given time window. The time window is set to 1 hour, meaning predictions within one hour are statistically analyzed. The calculation formula is:

[0197]

[0198] Wherein, TP represents true positives, i.e., the number of actual traffic anomalies that the model predicts as "abnormal"; TN represents true negatives, i.e., the number of actual traffic anomalies that the model predicts as "normal"; FP represents false positives, i.e., the number of actual traffic anomalies that the model predicts as "abnormal"; and FN represents false negatives, i.e., the number of actual traffic anomalies that the model predicts as "normal". In this embodiment, the prediction results within 1 hour are statistically analyzed. TP=12, TN=35, FP=3, FN=2, then the prediction accuracy is... That is, 90.38%.

[0199] Response latency refers to the total time from receiving a collaborative task request to outputting a collaborative service result, including data reception time, time-series alignment time, feature fusion time, and model prediction time. The system timer records the timestamps of each stage, and the response latency is the timestamp of the output collaborative service result minus the timestamp of receiving the collaborative task request. In this embodiment, the timestamp of receiving the collaborative task request is 2024-05-2009:58:30.123, and the timestamp of outputting the collaborative service result is 2024-05-2009:59:45.678, therefore the response latency is 75.555 seconds, or 1.26 minutes.

[0200] Resource utilization rate: refers to the ratio of system resources used during collaborative processing to the total system resources. CPU utilization rate is collected in real time by the operating system's performance monitoring tools, and the average CPU utilization rate during collaborative processing is statistically analyzed. Memory utilization rate is calculated as the ratio of the amount of memory used during collaborative processing to the total system memory. In this embodiment, the system has a total of 16 CPU cores, and an average of 8 cores are used during collaborative processing, so the CPU utilization rate is 8 / 16 = 0.5, or 50%. The total system memory is 64GB, and an average of 25.6GB is used during collaborative processing, so the memory utilization rate is 25.6 / 64 = 0.4, or 40%. The resource utilization rate is the average of the CPU utilization rate and the memory utilization rate, i.e., (50% + 40%) / 2 = 45%.

[0201] Task utility feedback signal generation:

[0202] The task utility feedback signal is a quantitative function value of the difference between the real-time performance indicator and the expected goal. The expected goal is the indicator requirement specified in the collaborative task request. In this embodiment, the expected goal is a prediction accuracy of not less than 85%, a response delay of not more than 2 minutes, and a resource utilization rate controlled below 70%.

[0203] The task utility feedback signal is calculated using a weighted summation method, as shown in the following formula:

[0204]

[0205]

[0206] in, : Task utility feedback signal, with a value range between 0 and 1. The larger the value, the better the collaborative service result meets the expected goal.

[0207] The weighting coefficient for prediction accuracy, with a value of 0.5, reflects the importance of prediction accuracy in collaborative tasks.

[0208] The weighting coefficient for response latency, with a value of 0.3, reflects the importance of response latency in collaborative tasks.

[0209] The weighting coefficient for resource utilization rate is 0.2, reflecting the importance of resource utilization rate in collaborative tasks.

[0210] The actual calculated prediction accuracy is 90.38% in this embodiment.

[0211] The expected prediction accuracy is 85% in this embodiment.

[0212] The actual calculated response latency is 75.555 seconds in this embodiment.

[0213] : The expected response delay is 120 seconds in this embodiment.

[0214] The actual calculated resource utilization rate is 45% in this embodiment.

[0215] The expected resource utilization rate is 70% in this embodiment.

[0216] The maximum value function is set to 0 when x is negative, ensuring that the contribution of each indicator to the task utility feedback signal is non-negative.

[0217] Substitute the values ​​from this example into the formula to calculate:

[0218]

[0219]

[0220]

[0221] but

[0222] .

[0223] Task state vector update:

[0224] The task state vector encoding is updated using the task utility feedback signal, specifically through a task encoder neural network. This neural network is the same one used to generate the task state vector in step (I), except that the input data includes the task utility feedback signal.

[0225] The update process is as follows: The task utility feedback signal and the current task state vector are input together into the task encoder neural network. The dimension of the input vector is 32 + 1 = 33 dimensions (32-dimensional task state vector + 1-dimensional task utility feedback signal). The task encoder neural network optimizes by minimizing the negative value of the task utility feedback signal, i.e., the optimization objective function is: .

[0226] The stochastic gradient descent optimization algorithm is used, with a learning rate of 0.0005, a batch size of 16, and 100 iterations per update. Network parameters are adjusted using backpropagation, specifically by calculating the objective function. By taking the partial derivatives of the parameters of each layer of the network, and based on the direction and magnitude of the partial derivatives, the parameters are gradually adjusted according to the learning rate to make the objective function... minimize.

[0227] The updated task state vector can better adapt to the actual needs of the current collaborative task and reflect the effectiveness of the collaborative service results. For example, in this embodiment, the task utility feedback signal is 0.3618, indicating that the collaborative service results basically meet the expected goals, but there is still room for optimization. Through updating, the dimensions related to data freshness requirements and feature reliability requirements in the task state vector will be fine-tuned according to the feedback signal, making the subsequent time-series alignment strategy and feature fusion process more aligned with the collaborative task objectives.

[0228] In this embodiment, the specific implementation of the dual-loop driven architecture operation mechanism is as follows:

[0229] The data flow adaptive inner loop and the task strategy optimization outer loop of this technical solution form a dual-loop driven architecture, ensuring dynamic adaptation and continuous optimization of the entire collaborative processing process.

[0230] The specific implementation of the adaptive inner loop for data flow is as follows:

[0231] The adaptive inner loop of the data flow consists of the modulation of the incremental temporal alignment strategy, the evaluation and weighted fusion of the dynamic feature trust modulus. Its core function is to adapt the processing of streaming heterogeneous data in real time according to the current task state vector, so as to ensure that the initial alignment features and task adaptation fusion features can accurately match the collaborative task objectives.

[0232] In this embodiment, the operation cycle of the adaptive inner loop of the data stream is synchronized with the acquisition cycle of the streaming data. Specifically, the operation cycle for the traffic monitoring domain is 5 minutes, the operation cycle for the numerical sequence data in the meteorological monitoring domain is 10 minutes, and the operation cycle for the social media domain is 1 minute. Within each operation cycle, the inner loop dynamically modulates the incremental temporal alignment strategy based on the current task state vector, performs temporal deviation correction and alignment on newly received heterogeneous streaming data blocks, evaluates the dynamic feature trust modulus of the preliminary alignment features, completes weighted fusion, and outputs task-adaptive fusion features.

[0233] The adaptive adjustment of the inner loop is reflected in the following ways: when the value of data freshness requirement in the task state vector increases, the incremental temporal alignment strategy will tilt towards the active interpolation strategy to reduce data caching time and improve real-time performance; when the value of feature reliability requirement increases, the calculation of dynamic feature trust modulus will focus more on the negative adjustment of data uncertainty and model uncertainty to filter low-quality features; when the domain weight bias changes, the dynamic feature trust modulus of each domain feature will be adjusted accordingly, and the fusion weight will also change accordingly to ensure that features of important domains dominate the fusion process.

[0234] Task strategy optimization outer loop:

[0235] The outer loop of task strategy optimization consists of the generation of task state vectors and the updating of task state vectors based on task utility feedback signals. Its core function is to adapt the operation effect of the inner loop according to the data flow, optimize the encoding of task state vectors, and provide accurate guidance signals for the inner loop.

[0236] In this embodiment, the outer loop of the task strategy optimization operates on a 1-hour cycle. Every hour, based on the real-time performance metrics of the collaborative service results from the previous hour, a task utility feedback signal is generated, and the task state vector is updated. The optimization process of the outer loop is reflected in the following ways: when the prediction accuracy is lower than the expected target, the value of the feature reliability requirement in the task state vector increases, guiding the inner loop to pay more attention to feature reliability in subsequent processing and filter low-quality data; when the response latency exceeds the expected target, the value of the data freshness requirement increases, guiding the inner loop to adopt a more aggressive time-series alignment strategy to reduce latency; when the resource utilization exceeds the expected target, the weight bias of each domain in the task state vector is adjusted, reducing the feature weights of domains with high resource consumption to improve resource utilization efficiency.

[0237] Dual-circulation linkage mechanism:

[0238] The data flow adaptive inner loop is dynamically controlled by the task strategy optimization outer loop, and the task strategy optimization outer loop is optimized based on the running effect of the data flow adaptive inner loop. The two form a closed loop linkage.

[0239] In this embodiment, initially, the outer loop generates an initial task state vector to guide the inner loop in streaming data processing. The inner loop outputs task adaptation and fusion features based on the initial task state vector, generating collaborative service results. The outer loop generates a task utility feedback signal based on the real-time performance indicators of the collaborative service results, updating the task state vector. The updated task state vector is fed back to the inner loop, guiding it to adjust its processing strategy. The inner loop optimizes the processing based on the new task state vector, outputting better task adaptation and fusion features and collaborative service results. This process is repeated continuously to achieve continuous optimization of the entire collaborative processing system.

[0240] For example, in the initial operation phase, because the task state vector is generated based on historical experience, it may deviate from the actual collaborative task requirements, resulting in a prediction accuracy of 86%, a response latency of 1.5 minutes, and a resource utilization rate of 60% for the collaborative service results. The outer loop generates task utility feedback signals based on these real-time performance indicators, updates the task state vector, increases the value of feature reliability requirements, and decreases the value of data freshness requirements. After receiving the updated task state vector, the inner loop adjusts the incremental temporal alignment strategy, increases data caching time, improves alignment accuracy, and strengthens the negative adjustment of data uncertainty and model uncertainty in the calculation of dynamic feature trust modulus. After the adjustment, the prediction accuracy of the collaborative service results in the next cycle improves to 90%, the response latency is 1.3 minutes, and the resource utilization rate is 55%, which is closer to the expected target. The outer loop updates the task state vector again based on the new real-time performance indicators, guiding the inner loop to further optimize and forming a virtuous cycle.

[0241] This embodiment, through the detailed implementation process described above, achieves cross-domain collaborative processing of heterogeneous information based on multi-dimensional feature adaptation, and has the following beneficial effects compared with the prior art:

[0242] This approach achieves an organic integration of task semantic understanding and streaming data processing, solving the technical problems of existing static migration and alignment methods being unable to adapt to real-time changes in task objectives, and the potential for resource waste and performance deviations in streaming processing lacking task semantic guidance. By using task state vectors throughout the entire collaborative processing flow and dynamically modulating temporal alignment strategies and feature fusion weights, the processing process remains focused on the core objectives of the collaborative task, improving the accuracy and relevance of the collaborative service results.

[0243] Adopting a dual-loop driven architecture, the data flow adaptive inner loop can adapt to the dynamic changes of streaming heterogeneous data in real time, while the task strategy optimization outer loop can continuously optimize the task state vector based on the collaborative service effect. The two work together to achieve continuous adaptive optimization of the system, improve the system's robustness and adaptability, and can cope with the collaborative task requirements in different scenarios.

[0244] The calculation of dynamic feature trust modulus comprehensively considers task relevance, data uncertainty, and model uncertainty. It can accurately quantify the contribution and credibility of features to collaborative tasks, provide a scientific basis for feature fusion, effectively filter low-quality features, improve the quality of task-adaptive fusion features, and thus enhance the reliability of collaborative service results.

[0245] By employing a lightweight time-series prediction model (linear Kalman filter) and an efficient feature fusion and collaborative service model (random forest algorithm), the system's real-time performance and efficiency are ensured, meeting the low-latency requirements of streaming data processing while controlling resource utilization and improving the system's engineering practicality.

[0246] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for cross-domain collaborative processing of heterogeneous information based on multi-dimensional feature adaptation, characterized in that, The method includes the following steps: Step 1: Receive a collaborative task request, perform semantic parsing on the collaborative task request, extract task semantic elements, and encode the task semantic elements into a task state vector, wherein the task state vector is a differentiable multidimensional vector. Step 2: In response to the collaborative task request, receive streaming heterogeneous data blocks from at least two different domains in real time. Based on the task state vector, dynamically modulate the incremental timing alignment strategy for the streaming heterogeneous data blocks. According to the modulated incremental timing alignment strategy, perform timing deviation correction and alignment on the streaming heterogeneous data blocks, and output preliminary alignment features. Step 3: Based on the task state vector and the preliminary alignment features, evaluate the dynamic feature trust modulus of the preliminary alignment features, and perform weighted fusion of the preliminary alignment features from different domains according to the dynamic feature trust modulus to generate task adaptation fusion features. Step 4: Generate collaborative service results based on the task adaptation and fusion features, and generate task utility feedback signals based on the real-time effect indicators of the collaborative service results and the expected goals of the collaborative task requests. Use the task utility feedback signals to update the encoding of the task state vector.

2. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, In step one, the collaborative task request is semantically parsed to extract task semantic elements, specifically including: identifying and extracting the task subject, associated domain, task objective and task constraints in the collaborative task request. The task state vector contains at least an implicit expression of data freshness requirements, feature reliability requirements and domain weight bias.

3. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, In step two, the incremental timing alignment strategy is dynamically modulated based on the task state vector, specifically including: decoding the real-time bias coefficient and the accuracy bias coefficient from the task state vector; selecting a target alignment strategy from a set of preset alignment strategies based on the ratio of the real-time bias coefficient to the accuracy bias coefficient; wherein the preset alignment strategies include at least an aggressive interpolation strategy that prioritizes low latency and a buffered fine alignment strategy that prioritizes alignment accuracy.

4. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, In step three, the dynamic feature trust modulus of the preliminary alignment feature is evaluated, specifically calculated using the following formula: in, This represents the dynamic feature trust modulus of the i-th feature at time t. This represents the sigmoid activation function. Represents the task state vector at time t. Let represent the initial aligned feature vector of the i-th feature at time t. It is the task relevance moderating coefficient. This represents the quantized value of the data uncertainty of the i-th feature at time t. This represents the quantized value of the model uncertainty for the i-th feature at time t. and These are the weighting coefficients for data uncertainty and model uncertainty, respectively. The dynamic feature trust modulus is positively adjusted by the correlation between the task state vector and the feature, and negatively adjusted by the data and model uncertainty of the feature itself.

5. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, In step three, the preliminary alignment features are weighted and fused according to the dynamic feature trust modulus. Specifically, the dynamic feature trust modulus of each feature is normalized to obtain a fusion weight, and the corresponding preliminary alignment features are weighted and summed using the fusion weight to generate the task adaptation fusion feature.

6. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, In step four, the task utility feedback signal is used to update the encoding of the task state vector. Specifically, this is achieved through the following strategy: the task utility feedback signal and the task state vector are input together into a task encoder neural network. The task encoder neural network takes minimizing the negative value of the task utility feedback signal as the optimization objective and adjusts its network parameters through a backpropagation algorithm, so that the task encoder neural network outputs an updated task state vector for the same or similar collaborative task requests. The task encoder neural network adopts a multilayer perceptron structure with an attention mechanism.

7. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, In step two, the modulation of the incremental time alignment strategy and the evaluation and weighted fusion of the dynamic feature trust modulus in step three constitute an adaptive inner loop of the data flow. The generation of the task state vector in step one and the updating of the task state vector based on the task utility feedback signal in step four constitute a task strategy optimization outer loop. The adaptive inner loop of the data flow is dynamically controlled by the outer loop of the task strategy optimization. The outer loop of the task strategy optimization is optimized based on the running effect of the adaptive inner loop of the data flow, and the two constitute a dual-loop driven architecture.

8. The method for cross-domain collaborative processing of heterogeneous information based on multi-dimensional feature adaptation according to claim 1, characterized in that, In step two, the timing deviation correction and alignment specifically includes: attaching a high-precision timestamp to each of the streaming heterogeneous data blocks; and estimating missing or delayed data points based on the high-precision timestamp and the current system timing using a lightweight timing prediction model, wherein the lightweight timing prediction model is a linear Kalman filter.

9. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, The different domains include data sources with different physical origins, different data formats, or different logical concepts; The streaming heterogeneous data block includes at least two of the following: text data, numerical sequence data, and image data.

10. The heterogeneous information cross-domain collaborative processing method based on multi-dimensional feature adaptation according to claim 1, characterized in that, The real-time performance metrics of the collaborative service results include at least one of prediction accuracy, response latency, and resource utilization. The task utility feedback signal is a quantified function value of the difference between the real-time performance indicator and the expected target.

Citation Information

Patent Citations

  • A cross-domain common transfer recommendation method and system based on multi-source auxiliary information fusion

    CN115757529B

  • Cross-domain and cross-source data alignment method and system and electronic equipment

    CN116050374A