Perception collaborative decision method and system based on multi-modal heterogeneous data fusion
By using a perception-cooperative decision-making method that integrates multimodal heterogeneous data, the problem of data silos in multimodal data processing is solved, enabling intelligent perception and rapid response to complex environments, and improving the system's real-time performance and decision execution efficiency.
Patent Information
- Application Number
- CN202511947275.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-12-22
AI Technical Summary
Existing technologies suffer from data silos in multimodal heterogeneous data processing, making it difficult to uncover deep semantic relationships between modalities and ensuring real-time performance, resulting in poor system performance in fast-response scenarios.
By constructing a perception-cooperative decision-making method that integrates multimodal heterogeneous data, images, text, time-series data, and graph structure data are collected, preprocessed, parsed, and feature extracted to generate a global environmental situational awareness map. Decision optimization is then performed through a multi-objective collaborative decision-making model, and finally, real-time scheduling and control are achieved through a distributed communication architecture.
It achieves unified representation of multimodal data, improves data adaptability and reliability, strengthens spatial perception and feature learning capabilities in complex environments, and ensures rapid response and efficient execution of artificial intelligence decision-making.
Smart Images

Figure CN121365368B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a perception-based collaborative decision-making method and system based on multimodal heterogeneous data fusion. Background Technology
[0002] As artificial intelligence (AI) technology evolves from single-modal processing to multimodal fusion and collaborative decision-making, intelligent systems are increasingly being applied in areas such as government services, AI-powered animal husbandry, smart homes, and digital agriculture. These systems need to simultaneously process image, text, time-series data, and graph-structured data from multiple sensors, forming a typical multimodal heterogeneous data environment. However, related technologies have significant shortcomings at the data processing level, severely limiting system performance.
[0003] First, due to the prominent problem of data silos caused by modal heterogeneity, the spatial features of images, the semantic information of text, the dynamic patterns of time-series data, and the topological relationships of graph structure data belong to different feature spaces and lack a unified representation model. Related technologies mostly adopt shallow fusion strategies of simple splicing or weighted averaging after processing each modal data independently, which makes it difficult to explore the deep semantic relationships between modalities. This may lead to insufficient discriminative power of fused features and an inability to accurately characterize the state of complex systems.
[0004] Secondly, the data processing pipeline is lengthy and real-time performance is difficult to guarantee. Traditional batch processing mode cannot adapt to the continuous high-speed data flow. There is a delay in the entire process from data acquisition to decision output, which makes the system make decisions based on outdated information and performs poorly in application scenarios that require rapid response. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a perception and collaborative decision-making method and system based on multimodal heterogeneous data fusion. By constructing a linkage mechanism of multimodal heterogeneous data fusion, dynamic environment perception and collaborative decision execution, intelligent perception and rapid response to complex environments can be achieved.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0007] Firstly, a perceptual collaborative decision-making method based on multimodal heterogeneous data fusion, the method comprising:
[0008] Collect multimodal heterogeneous data, including images, text, time-series data, and graph structure data; preprocess the multimodal heterogeneous data to obtain a standardized data stream;
[0009] The standardized data stream is parsed and its features are extracted, aligned, and encoded to obtain multi-source feature vectors. These multi-source feature vectors are then weighted and fused to obtain a preliminary indicator set. Based on this preliminary indicator set, predefined key areas are gridded, and the dynamic correlation characteristics of each spatial grid cell are analyzed to obtain a spatial dynamic weight matrix. This spatial dynamic weight matrix is then weighted and fused with the preliminary indicator set, and optimized comprehensive state indicators are obtained through spatial context calibration. Finally, a global environmental situational awareness map is generated based on the optimized comprehensive state indicators.
[0010] The overall environmental situational awareness map is input into a pre-trained multi-objective collaborative decision-making model. The multi-objective collaborative decision-making model analyzes the resource entities and their relationships in the overall environmental situational awareness map to form a multi-objective decision feature set. Based on the multi-objective decision feature set, decision optimization is performed to obtain a comprehensive collaborative scheduling scheme.
[0011] The comprehensive collaborative scheduling scheme is parsed and encapsulated to obtain an executable instruction sequence. The executable instruction sequence is then distributed in parallel to the corresponding decision nodes and control terminals through a distributed communication architecture, thereby completing the real-time scheduling of resources and the collaborative release of control instructions.
[0012] Secondly, a perception-based collaborative decision-making system based on multimodal heterogeneous data fusion includes:
[0013] The multi-source acquisition module is used to acquire multimodal heterogeneous data, including images, text, time-series data, and graph structure data; it preprocesses the multimodal heterogeneous data to obtain a standardized data stream.
[0014] The fusion analysis module is used to parse and extract features from the standardized data stream, and then align and encode it to obtain multi-source feature vectors. These multi-source feature vectors are then weighted and fused to obtain a preliminary indicator set. Based on this preliminary indicator set, predefined key areas are gridded, and the dynamic correlation characteristics of each spatial grid cell are analyzed to obtain a spatial dynamic weight matrix. This spatial dynamic weight matrix is then weighted and fused with the preliminary indicator set, and optimized using spatial context calibration to obtain a comprehensive state indicator. Finally, based on the optimized comprehensive state indicator, a global environmental situational awareness map is generated.
[0015] The collaborative decision-making module is used to input the global environmental situational awareness map into the pre-trained multi-objective collaborative decision-making model. The multi-objective collaborative decision-making model forms a multi-objective decision feature set by analyzing the resource entities and their relationships in the global environmental situational awareness map. Based on the multi-objective decision feature set, decision optimization is performed to obtain a comprehensive collaborative scheduling scheme.
[0016] The execution module is used to parse and encapsulate the comprehensive collaborative scheduling scheme into an executable instruction sequence; through a distributed communication architecture, the executable instruction sequence is distributed in parallel to the corresponding decision nodes and control terminals to complete the real-time scheduling of resources and the collaborative release of control instructions.
[0017] Thirdly, a computing device includes:
[0018] One or more processors;
[0019] A storage device for storing one or more programs, and methods implemented by one or more processors when the programs are executed.
[0020] Fourthly, a computer-readable storage medium storing a program, the method of which is implemented when the program is executed by a processor.
[0021] The above-described solution of the present invention has at least the following beneficial effects:
[0022] First, by collecting and preprocessing multimodal heterogeneous data, the heterogeneity barrier of multimodal data is broken down. By integrating scattered image, text, time-series data, and graph structure data into a standardized data stream with a unified format, data format differences, redundant information, and abnormal interference can be eliminated, enhancing the consistency and reliability of input data. Then, by parsing, extracting, aligning, encoding, and weighting the standardized data stream, this invention mines the deep correlations between multimodal data, generating multi-source feature vectors that are both complete and discriminative. Through gridding processing and dynamic correlation characteristic analysis, a spatial dynamic weight matrix is constructed to achieve precise binding between features and spatial locations. Combined with spatial context calibration, feature optimization is completed, forming a structured comprehensive state index and a global environmental situational awareness map, which can strengthen the expression of spatial correlation features and improve spatial perception and feature learning capabilities in complex environments. Next, this invention will... The pre-trained model inputs a domain environment situational awareness map, analyzes resource entities and their relationships to form a multi-objective decision feature set, enabling the transformation of visualized spatial information into structured decision features. Decision deduction is completed through a multi-objective collaborative optimization algorithm, constructing a multi-objective overall decision framework. This promotes the evolution from single-objective decision-making to multi-objective collaborative decision-making, achieving efficient collaborative reasoning and improving the systematic nature of AI in handling complex decision-making tasks. Finally, this invention constructs a bridge between abstract decisions and concrete execution instructions by parsing and encapsulating the collaborative scheduling scheme, ensuring the executability of decision outputs. Parallel instruction issuance through a distributed communication architecture leverages the advantages of distributed processing to improve execution efficiency. Combined with state feedback and dynamic adjustment mechanisms, a closed-loop system of decision-making, execution, and feedback is constructed, strengthening adaptability and dynamic self-adaptation capabilities, enabling AI decision-making to quickly respond to real-time changes in the actual scenario. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating a perception-based collaborative decision-making method based on multimodal heterogeneous data fusion, provided as an embodiment of the present invention.
[0024] Figure 2 This is a schematic diagram of a perception-cooperative decision-making system based on multimodal heterogeneous data fusion, provided as an embodiment of the present invention. Detailed Implementation
[0025] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0026] Embodiments of this invention propose a perceptual collaborative decision-making method based on multimodal heterogeneous data fusion, such as... Figure 1 The diagram shown is a flowchart illustrating a perception-based collaborative decision-making method based on multimodal heterogeneous data fusion, provided by an embodiment of the present invention. The method includes the following steps:
[0027] Step 100: Collect multimodal heterogeneous data, including images, text, time-series data, and graph structure data; preprocess the multimodal heterogeneous data to obtain a standardized data stream.
[0028] Step 200: The standardized data stream is parsed and its features are extracted, aligned, and encoded to obtain multi-source feature vectors; the multi-source feature vectors are weighted and fused to obtain a preliminary indicator set; based on the preliminary indicator set, the predefined key areas are gridded, and the dynamic correlation characteristics of each spatial grid unit are analyzed to obtain a spatial dynamic weight matrix; the spatial dynamic weight matrix is weighted and fused with the preliminary indicator set, and optimized comprehensive state indicators are obtained through spatial context calibration; based on the optimized comprehensive state indicators, a global environmental situation awareness map is generated.
[0029] Step 300: Input the global environmental situational awareness map into the pre-trained multi-objective collaborative decision-making model; the multi-objective collaborative decision-making model forms a multi-objective decision feature set by analyzing the resource entities and their relationships in the map; based on the multi-objective decision feature set, perform decision optimization to obtain a comprehensive collaborative scheduling scheme.
[0030] Step 400 involves parsing and encapsulating the comprehensive collaborative scheduling scheme to obtain an executable instruction sequence; then, through a distributed communication architecture, the executable instruction sequence is distributed in parallel to the corresponding decision nodes and control terminals to complete the real-time scheduling of resources and the collaborative release of control instructions.
[0031] In this embodiment of the invention, data silos are broken down by collecting multimodal heterogeneous data; preprocessing the data to form a standardized data stream unifies data format standards and improves data adaptability; parsing and extracting features from the standardized data stream and completing alignment encoding strengthens the expressive power of multi-source features; a preliminary indicator set is obtained through weighted fusion, and a spatial dynamic weight matrix is constructed by combining gridded processing of key areas and analysis of spatial dynamic correlation characteristics, which can improve the pertinence and rationality of feature fusion; comprehensive state indicators are optimized through spatial context calibration to generate a global environmental situational awareness map, constructing a structured situational representation, which helps to deeply understand complex environments; the global environmental situational awareness map is input into a pre-trained model to analyze resource entities and their relationships to enrich the dimensions of decision features; a multi-agent collaborative algorithm is used to optimize decision-making and achieve efficient collaborative reasoning, which can improve the integrity and collaboration of decision-making; the collaborative scheduling scheme is parsed and encapsulated to obtain an executable instruction sequence, and the standardized execution instructions ensure the accuracy of AI decision implementation; parallel instruction issuance through a distributed communication architecture improves the efficiency and real-time performance of AI decision execution and strengthens the closed-loop capability from decision-making to execution.
[0032] In some embodiments, step 100 above involves collecting multimodal heterogeneous data, including images, text, time-series data, and graph structure data; preprocessing the multimodal heterogeneous data to obtain a standardized data stream, including a1-a5.
[0033] a1, Parallel acquisition of multimodal heterogeneous data from multiple heterogeneous data sources, including at least one of the following: monitoring images, sensor nodes, mobile units, and environmental information data sources.
[0034] In some embodiments, multimodal heterogeneous data includes images, text, time-series data, and graph-structured data.
[0035] In some embodiments, a1 can be specifically implemented as follows: Distributed data acquisition nodes are deployed to capture scene image data in real time using high-definition cameras for monitoring image data sources. Simultaneously, an attitude calibration algorithm is used to detect attitude parameters in the captured images. Pitch, yaw, and roll deviations of the images are corrected according to a preset calibration model to ensure the consistency of the image's spatial attitude. For sensor nodes, time-series sensing data such as temperature, humidity, and air pressure are collected. For mobile units, dynamic data such as location trajectories and state parameters are collected. For environmental information data sources, textual data such as meteorological reports and geographical descriptions, as well as graph structure data such as regional association topology, are obtained. During the acquisition process, a k-nearest neighbor spatial indexing algorithm is integrated to construct a k-nearest neighbor index structure based on spatial location, quickly matching adjacent and related data of various heterogeneous data sources in the spatial dimension to achieve synchronous association acquisition of spatially related data. A multi-threaded concurrent acquisition mechanism is adopted to synchronously receive multimodal data from various heterogeneous data sources. Acquisition scheduling is optimized through spatial indexing, and image data quality is ensured through attitude calibration, ensuring the comprehensiveness, parallelism, and spatial consistency of data acquisition.
[0036] a2 performs format parsing and data cleaning on multimodal heterogeneous data, removing invalid and redundant data to obtain structured data.
[0037] In some embodiments, a2 above can be specifically implemented as follows: for the collected multimodal heterogeneous data, a format parsing tool adapted to different data types is used to convert image data into a unified pixel matrix format, text data into a standard character encoding format, time series data into a timestamp-associated numerical sequence format, and graph structure data into a standardized storage format with node-edge association; redundant information such as duplicate records and null values is removed by rule matching and filtering, and invalid data with format errors or exceeding reasonable range is removed by data validity verification, and the processed data is organized into structured data with a unified structure.
[0038] a3 performs outlier detection and missing data repair on structured data to obtain reduced data.
[0039] In some embodiments, a3 above can be specifically implemented as follows: First, statistical analysis methods combined with domain threshold standards are used to identify outliers in numerical time-series data in structured data, detect node-edge association anomalies in graph structure data, and determine semantic contradictions in text data; then, the detected outliers are processed using reasonable replacement or removal strategies, and missing values in the dataset are processed according to data type classification. For non-spatial related data, mean imputation or similarity-based association imputation is selected for repair; for numerical time-series data containing spatial location information (such as sensor data distributed across different spatial nodes)... The spatial geometric interpolation algorithm is used for repair. The specific calculation process is as follows: First, the spatial coordinate information of the missing data point is extracted. Then, the effective adjacent data points within a preset range around the coordinates are selected. The spatial distance and geometric position relationship between each adjacent data point and the missing point are determined. The weight coefficients are assigned according to the inverse proportion of the spatial distance. The adjacent data points that are closer to the missing point have a higher weight ratio. The initial interpolation result of the missing point is calculated by weighted summation. Then, the interpolation value is finely adjusted in combination with the temporal change law of the data to ensure that the repaired data not only conforms to the spatial distribution characteristics but also maintains the temporal continuity. Finally, the reduced data with unified data format, complete content, and consistent spatiotemporal correlation is formed.
[0040] a4 performs time series alignment and spatial location matching on the standardized data to establish a unified spatiotemporal benchmark and obtain spatiotemporally aligned data.
[0041] In some embodiments, a4 above can be specifically implemented as follows: extracting timestamp information from each protocol data, adjusting time-series data from different data sources to the same time granularity based on a high-precision time synchronization protocol, and achieving alignment of each modality data in the time dimension; for image data, mobile unit location data, and graph structure topology data containing spatial information, a unified geographic coordinate system is first used for preliminary spatial location calibration, associating each data with the corresponding spatial coordinate points or regions, and then the iterative nearest point algorithm is fused to optimize the spatial alignment accuracy. The specific calculation process is as follows: selecting a data source with clear spatial features and uniform distribution as a reference dataset, and using other datasets to be calibrated as target datasets, extracting key spatial feature points from the two datasets respectively, including the pixel coordinates of the images. Feature points, the coordinates of the mobile unit, and the coordinates of the topological nodes in the graph structure are analyzed. A spatial transformation matrix is initialized, including translation, rotation, and scaling parameters. The nearest neighbor pairs of feature points in the target dataset and the reference dataset are iteratively searched, and the spatial distance error between each pair of corresponding points is calculated to construct an error function. The optimal transformation matrix that minimizes the error function is solved using the least squares method, and the spatial coordinates of the target dataset are updated. The nearest neighbor pair matching, error calculation, and transformation matrix optimization steps are repeated until the error change between two adjacent iterations is less than a preset threshold, at which point the iteration stops. The positional deviations of spatial data from different data sources are corrected using an iterative nearest-point algorithm, establishing a more accurate spatial correlation mapping between data, resulting in spatiotemporally aligned data with both a unified time reference and a high-precision spatial reference.
[0042] a5 normalizes the spatiotemporally aligned data to eliminate dimensional differences and obtain a standardized data stream.
[0043] In some embodiments, the above-mentioned a5 can be specifically implemented as follows: For numerical data (including time-series sensing data, location coordinate data, etc.) in spatiotemporal aligned data, a min-max normalization or z-score standardization method is adopted to map data with different dimensions and different numerical ranges to a unified numerical interval, eliminating the problem of data incomparability caused by differences in dimensions, while maintaining the original distribution characteristics and correlation patterns of the data, and finally forming a standardized data stream with unified format, spatiotemporal consistency, and unified dimensions, providing a highly adaptable data foundation for subsequent feature extraction and fusion.
[0044] In this embodiment of the invention, by acquiring multimodal data from multiple heterogeneous data sources in parallel, data silos can be broken down, enriching the dimensions of data input; the data coverage can be broadened, providing a diverse data foundation for multi-source information fusion; by parsing and cleaning the format of the multimodal heterogeneous data, invalid and redundant information can be removed, and the data form can be standardized; data noise interference can be reduced, and data quality can be improved; by carrying out outlier detection and missing data repair, data defects can be remedied, ensuring data integrity and consistency, and strengthening data reliability; by performing time series alignment and spatial location matching, a unified spatiotemporal benchmark can be established; the spatiotemporal misalignment problem of multi-source data can be eliminated, giving data spatiotemporal correlation; finally, by performing normalization processing on spatiotemporally aligned data, dimensional differences can be eliminated; data scale standards can be unified, and data comparability can be improved.
[0045] In some embodiments, step 200 above involves parsing and extracting features from the standardized data stream, aligning and encoding it to obtain multi-source feature vectors; weighting and fusing the multi-source feature vectors to obtain a preliminary indicator set; based on the preliminary indicator set, performing gridding on predefined key areas and analyzing the dynamic correlation characteristics of each spatial grid unit to obtain a spatial dynamic weight matrix; weighting and fusing the spatial dynamic weight matrix with the preliminary indicator set and calibrating it through spatial context to obtain an optimized comprehensive state indicator; and generating a global environmental situation awareness map based on the optimized comprehensive state indicator, including b1-b16.
[0046] b1. Perform time-dimensional analysis on the standardized data stream, dividing the continuous standardized data stream into equally spaced time slice sequences. Specifically, this includes: First, analyzing the acquisition frequency of the standardized data stream and the decision latency requirements of the application scenario. For example, the acquisition frequency is 10Hz, and the decision latency requirement of the application scenario is 500ms. The time interval parameter is determined to be 100ms by calculating the matching degree between data transmission bandwidth and processing computing power. A sliding window partitioning method is adopted with a step size of 50ms. At the same time, the sliding window geometric constraint algorithm is integrated to optimize the window partitioning. First, extract the spatial coordinate boundaries of each modality data in the standardized data stream, determine the total number of preset spatial grids in the key area, such as 10,000, and set the window spatial coverage integrity threshold to 90%, that is, each sliding window must cover no less than 90% of the preset spatial grids. For each potential time slice, the number of effective spatial grids covered by each modality is calculated based on the spatial coordinates of each modality. The spatial coverage integrity is obtained by the ratio of the number of effective grids to the total number of grids. If the ratio is less than 90%, the window start time is dynamically adjusted in 1ms increments, with the adjustment range not exceeding 20ms. The spatial coverage integrity of the adjusted window is recalculated until the threshold requirement is met. The window that meets the geometric constraints is taken as the final time slice. Each slice contains complete multimodal heterogeneous data within the 100ms time period. The timestamp continuity of each slice is checked, with an allowable error range of ±1ms. Data segments with duplicate or missing timestamps are removed to ensure the orderliness, integrity, and consistency of spatial coverage of the time series data. This provides regular, data-free, and spatiotemporally adapted time units for subsequent time-segmented processing.
[0047] b2. Based on equally spaced time slice sequences, spatial registration processing is performed on the multimodal data within each time slice to establish spatial reference data under a unified coordinate system. The spatial reference data includes moving object trajectory data, fixed sensor data, and moving sensor data. Specifically, this includes: selecting the World Geodetic System 1984 (WGS-84) geographic coordinate system as the unified spatial reference based on equally spaced time slice sequences; and optimizing spatial registration accuracy by fusing projection transformation algorithms. First, the original projection type of each modal data is identified through data source metadata parsing or preset configuration, including but not limited to Universal Transverse Mercator (UTM). Mercator (UTM) projection, Gauss-Kruger projection, and Mercator projection were used to extract key parameters of the corresponding original projections, including central meridian longitude, datum ellipsoid parameters, projection scale factor, and offset. The specific calculation process involved: for the original projection coordinate data, geodetic coordinates were derived backwards based on the forward projection model; the geodetic coordinates (latitude and longitude) under the original datum were converted to geodetic coordinates under the WGS-84 datum through ellipsoid parameter transformation; then, an error correction model was used to compensate for projection distortion errors, ensuring that the coordinate error after projection transformation was less than or equal to 0.3m. Spatial registration was performed on the multimodal data within each time slice; the trajectory data of moving objects was first converted to WGS-84 datum through projection transformation. S-84 latitude and longitude coordinates are used, and then the GPS positioning deviation is corrected through a preset coordinate transformation matrix, with a correction accuracy of ±0.5m. Fixed sensor data first transforms the projected coordinates of the original installation location into WGS-84 latitude and longitude coordinates through the above process, accurate to 6 decimal places, and then binds them one by one with the sensor data and stores them in GeoJSON format. Mobile sensor data first performs projection transformation on the original projected coordinates output by the positioning module, and then synchronizes with the WGS-84 coordinates output by the Beidou positioning module in real time, updating the coordinate association relationship every 10ms. Finally, the data is integrated to form a unified coordinate system spatial reference data containing mobile object trajectory data, fixed sensor data and mobile sensor data.
[0048] b3. Based on the trajectory data of moving objects in the spatial reference data, determine the object distribution density within each grid cell to obtain the spatiotemporal feature matrix; based on fixed sensor data, construct a flow feature vector by statistically analyzing the correlation between event intervals and the number of events per unit time; using moving sensor data, perform outlier removal and mean calculation on the dynamic parameter sampling sequence to obtain the regional average dynamic characteristics. Specifically, this includes: based on the spatial reference data, first setting a grid size of 10m×10m according to the accuracy requirements of the application scenario, determining the grid division rules according to the key area range, determining the latitude and longitude coordinates of the grid cell boundaries, quickly locating the grid cell to which the moving object trajectory data belongs using the R-tree spatial index, counting the real-time number of moving objects in each grid cell, and calculating the object distribution density (unit: objects / m). 2 A spatiotemporal feature matrix (dimension: number of grids × number of time slices) containing time slice identifiers, grid numbers, and corresponding distribution densities is constructed. Based on fixed sensor data, an event trigger record for each sensor is extracted using a 1-second sliding time window. The event occurrence interval and the number of events per unit time within the window are statistically analyzed (number of events within the window / window duration). The positive or negative correlation between the two is analyzed, and a flow feature vector reflecting the frequency and pattern of event occurrence is constructed, with dimensions including sensor number, event frequency, and mean interval. Using mobile sensor data, a dynamic parameter sampling sequence is divided into 100m × 100m spatial regions. Abnormal sampling values in the sequence are removed using the 3σ criterion. The arithmetic mean of the effective dynamic parameters in each region is calculated to obtain the regional average dynamic characteristics reflecting the overall dynamic state of the region.
[0049] b4 aligns the spatiotemporal feature matrix, flow feature vector, and regional average dynamic feature by feature dimension alignment, and then concatenates them using vectorization to obtain a multi-source feature vector. Specifically, this involves first parsing the dimensional structure and data type of the spatiotemporal feature matrix, flow feature vector, and regional average dynamic feature (converting them to 32-bit floating-point numbers (float32)). The spatiotemporal feature matrix has a dimension of grid number × time slice number, the flow feature vector has a dimension of sensor number × 2, and the regional average dynamic feature has a dimension of region number × 1. The spatiotemporal feature matrix is then flattened into a one-dimensional vector. The flow feature vector is spatially expanded according to the grid to which the sensor belongs, with a dimension of grid number × time slice number × 2. The regional average dynamic feature is spatially expanded according to the grid contained in the region, with a dimension of grid number × time slice number × 1. This ensures that the three dimensions are completely consistent and the data types are matched, thus completing the feature dimension alignment. Then, according to the time slice order, the three types of aligned features are vectorized and concatenated along the feature dimensions to form a multi-source feature vector corresponding to each time slice, with a dimension of grid number × time slice number × 4, thereby achieving effective integration of multi-type feature information.
[0050] b5, based on multi-source feature vectors, determines the attention weight matrix of each feature vector through a pre-defined fully connected layer and a normalization exponent (softmax) function. Specifically, it includes: a pre-defined fully connected layer structure based on the multi-source feature vector dimension, feature transformation requirements, and gradient stabilization objectives. This fully connected layer contains an input layer, two hidden layers, and an output layer. The number of neurons in the input layer is consistent with the dimension of the multi-source feature vectors, where the multi-source feature vector dimension is the number of grids × the number of time slices × 4. The number of neurons in the first hidden layer is set to half that of the input layer, and a Rectified Linear Unit (ReLU) activation function is used to achieve non-linear feature transformation. The number of neurons in the second hidden layer is set to one-quarter that of the input layer, and a Leaky Rectified Linear Unit (ReLU) activation function is used. The LeakyReLU activation function is used to avoid gradient vanishing. The number of neurons in the output layer is the same as that in the input layer. An L2 regularization mechanism is preset to suppress overfitting. The multi-source feature vectors are reshaped into one-dimensional vectors and then input into a fully connected layer. After passing through the input layer, they are passed to two hidden layers for feature transformation. The output layer outputs the original scores for each feature and performs preset L2 regularization. The regularized original scores are then input into a softmax function, which normalizes the original scores of each feature, mapping them to the range of 0 to 1, with the sum of all feature scores being 1. The resulting value is the attention weight for each feature, used to quantify the contribution of different features to the overall feature representation. Finally, the attention weight matrix corresponding to each feature vector is obtained, with the same dimension as the multi-source feature vectors. Gradient descent optimization ensures that the weight of key features accounts for no less than 30%, intuitively reflecting the importance of each feature.
[0051] b6. The attention weight matrix and multi-source feature vectors are weighted and summed to obtain the attention-weighted fusion features. Specifically, this involves: first, confirming the dimensionality consistency between the attention weight matrix and the multi-source feature vectors, both being three-dimensional structures of grid number × time slice number × 4, ensuring a one-to-one correspondence between each weight value and its corresponding feature element in terms of grid number, time slice sequence, and feature dimension; performing element-wise multiplication on each weight value in the attention weight matrix and its corresponding feature element in the multi-source feature vector, retaining six decimal places during the operation, to obtain a three-dimensional weighted feature component set, whose dimension is consistent with the weight matrix and multi-source feature vectors; and then performing time slice dimension calculations. The process iterates through each grid and each feature dimension, summing the weighted feature components corresponding to all time slices within the same grid and feature dimension to obtain the time aggregated value for each grid and each feature dimension. Then, it integrates all the time aggregated values of each grid according to the feature dimension to form a four-dimensional feature vector for each grid. Through the above two-step summation operation, the weighted contributions of each feature in different time and spatial dimensions are aggregated, ultimately generating an attention-weighted fusion feature with a dimension of grid number × 4. This process, through the differentiated allocation of weight values, ensures that the weighted components of key features dominate after summation, thereby highlighting the supporting role of key features in the overall feature expression, while weakening the influence of the weighted components corresponding to secondary features.
[0052] b7. The attention-weighted fusion features are subjected to dimensionality reduction and standardization to obtain a preliminary index set. Specifically, this includes: first, data normalization of the attention-weighted fusion features with a dimension of grid number × 4, organizing them into a two-dimensional data matrix with four feature dimensions as columns and the feature data corresponding to each grid as rows; based on this two-dimensional data matrix, the covariance between each feature dimension is calculated, constructing a 4×4 covariance matrix. Covariance calculation uses the feature data of all grids as a complete sample set, solving for the covariance values of all sample data under any two feature dimensions one by one, ensuring that each element in the covariance matrix corresponds to the linear correlation between the two feature dimensions; eigenvalue decomposition is performed on the constructed covariance matrix to obtain four eigenvalues and each... The eigenvectors that uniquely correspond to each eigenvalue are used to sort the four eigenvalues in descending order of value. The variance contribution rate of each eigenvalue to the sum of all eigenvalues is calculated one by one, and then the variance contribution rates are accumulated in sorting order to obtain the cumulative variance contribution rate. When the cumulative variance contribution rate first reaches the preset screening condition of greater than or equal to 95%, the accumulation operation is stopped, and all eigenvectors corresponding to the eigenvalues involved in the accumulation are recorded. These eigenvectors are then arranged into a projection matrix in the sorting order of their corresponding eigenvalues. The number of columns in the projection matrix is the number of principal components retained, which is also the number of features after dimensionality reduction. The original two-dimensional data matrix is multiplied by the projection matrix to obtain the dimensionality-reduced feature matrix, which retains the number of rows as the number of grids and the number of columns as the number of principal components retained.
[0053] Subsequently, z-score standardization is performed on the dimensionality-reduced feature matrix. First, the feature mean and sample standard deviation of all elements in the dimensionality-reduced feature matrix are calculated. The feature mean is the arithmetic mean of all elements, and the sample standard deviation is the unbiased standard deviation calculated based on all elements, retaining six significant decimal places in the calculation. For each element in the dimensionality-reduced feature matrix, the difference between the original value and the feature mean is calculated, and then the difference is divided by the sample standard deviation to complete the standardization transformation of a single element. This transformation process is repeated until all elements in the dimensionality-reduced feature matrix have been processed, mapping the standardized feature matrix to a unified numerical range, completely eliminating the differences in dimensions and numerical ranges between different feature dimensions. Finally, a preliminary index set with dimensions of grid number × number of dimensionality-reduced features is obtained.
[0054] b8. Based on a preliminary indicator set and combined with predefined key area boundaries, construct a regular grid topology covering the area. Specifically, this includes: predefining key area boundaries using a preset method, considering the functional requirements of the specific application scenario, the actual management scope, and the effective coverage of heterogeneous data sources. This boundary is defined as a set of latitude and longitude coordinates of several consecutive polygon vertices, forming a standardized boundary definition file. The boundary definition file is read through a GIS interface to obtain the complete latitude and longitude coordinate information of the predefined key area, including the longitude and latitude values of the boundary vertices. The actual area of the key area, such as 10 km², is calculated based on the obtained boundary coordinates. 2 Based on the scenario's precision requirements for data processing, such as 1m, the grid size parameter is determined to be 1m×1m. The north-south span (in meters) of the region is calculated by converting the latitude difference between the northernmost and southernmost vertices of the boundary, and this north-south span value is used as the number of rows in the grid. The east-west span (in meters) of the region is calculated by converting the longitude difference between the easternmost and westernmost vertices of the boundary, and this east-west span value is used as the number of columns in the grid, ensuring that the grid structure can completely cover the key areas without any spatial omissions. A rectangular grid division rule is adopted, and a regular grid topology structure is constructed according to the set number of rows, columns, and size parameters. The grid is then divided into rows from left to right. Starting from the top corner grid, each grid is assigned a unique number in the format of row number and column number. Based on the number of rows, columns, and size parameters of each grid, and combined with the starting latitude and longitude coordinates of the key area boundary, the latitude and longitude coordinates of the four corners of each grid are calculated one by one, establishing a one-to-one mapping relationship between the unique grid number and the latitude and longitude coordinates of the four corners of the grid. An R-tree index structure is constructed to store the mapping relationship between the grid number and the corresponding coordinate in the index, enabling the function of quickly querying the corresponding coordinate by grid number or quickly locating the grid number by coordinate. Finally, a clear spatial grid system is formed, providing a structured carrier for spatial feature mapping.
[0055] b9, based on a regular grid topology, maps the preliminary indicator set to each spatial grid cell, and calculates the statistical distribution characteristics of environmental state indicators within each grid cell. Specifically, this includes: first, extracting the latitude and longitude coordinates associated with each indicator data in the preliminary indicator set; comparing these coordinates with the preset latitude and longitude coordinate ranges of each spatial grid cell; determining whether the coordinates of each indicator data fall within the upper and lower limits of longitude and latitude of the target grid cell, thus completing the preliminary matching between the indicator data and its corresponding spatial grid cell; verifying the matching results through the mapping relationship between grid number and spatial coordinates, eliminating invalid indicator data whose coordinates exceed all grid ranges, ensuring that each valid indicator data uniquely corresponds to one spatial grid cell; and grouping the environmental state indicator data within each grid cell according to the unique identifier of the time slice, such as the time slice sequence number, so that indicator data of the same grid cell and the same time slice are grouped together, forming a three-dimensional grouping structure of grid cell, time slice, and indicator data.
[0056] For each environmental status indicator data in each data set, first count the total number of valid data, i.e., the amount of data remaining after removing null and outlier values. When calculating the arithmetic mean, sum all valid indicator data in the set, divide the sum by the total number of valid data, and retain six significant decimal places. When calculating the unbiased variance, first subtract the arithmetic mean of the set from each valid data, square each difference, sum all squared differences to obtain the total sum of squares, divide the total sum of squares by the total number of valid data minus one (i.e., degrees of freedom n-1), and also retain six significant decimal places.
[0057] When calculating the median, all valid data in the group are sorted in ascending order. If the total number of valid data is odd, the median is the value in the middle position after sorting (position is the total number of valid data plus one divided by two). If the total number of valid data is even, the median is the value in the two middle positions after sorting (positions are the total number of valid data divided by two and the total number of valid data divided by two plus one, respectively). The arithmetic mean of these two values is then calculated as the median.
[0058] When calculating the mode, the frequency of each different indicator value within the group is counted, and the value with the highest frequency is recorded. If two or more values have the same frequency and are both the highest, the arithmetic mean of these values is calculated as the mode of the group of indicator data. For all environmental state indicator data within each grid cell, the above calculation process of arithmetic mean, unbiased variance, median, and mode is repeated to comprehensively characterize the distribution pattern and dispersion of environmental state indicators under different time slices within each grid cell. The four statistical results corresponding to each grid cell and each time slice are integrated according to indicator type to form a grid cell-specific statistical distribution feature with a dimension of grid number × time slice number × 4. An indexed storage method is used to associate grid number, time slice identifier, and statistical results, supporting quick retrieval and calling of corresponding statistical data by time dimension or spatial dimension.
[0059] b10, based on the statistical distribution characteristics of environmental state indicators within each grid cell, calculates the dynamic correlation strength between adjacent grid cells through spatial autocorrelation analysis. Specifically, it employs spatial autocorrelation analysis, selecting the Moran index as the global correlation calculation indicator and the Local Indicators of Spatial Association (LISA) statistic as the local correlation calculation indicator. Using the Queen adjacency rule (grids sharing edges or vertices are considered adjacent), it performs correlation calculations between the mean statistical distribution characteristics of the environmental state indicators of each grid cell and the mean statistical distribution characteristics of surrounding adjacent grid cells. The calculation results quantify the degree of correlation between adjacent grid cells, mapping the correlation strength range to a 0-1 interval. A correlation strength greater than or equal to 0.7 is considered strong correlation, 0.3 to 0.7 is considered moderate correlation, and less than or equal to 0.3 is considered weak correlation. This yields the dynamic correlation strength between each grid cell and each adjacent grid cell, capturing the linkage characteristics in the spatial dimension.
[0060] b11, combining the statistical distribution characteristics of environmental state indicators within each grid cell and the dynamic correlation strength between adjacent grid cells, a spatial dynamic weight matrix is obtained through weighted fusion calculation. Specifically, this includes: an adaptive learning method based on a gradient boosting tree model, setting initial weight coefficients for statistical distribution characteristics and dynamic correlation strength (statistical distribution characteristics 0.6, dynamic correlation strength 0.4), optimizing the weight coefficients through 5-fold cross-validation to ensure minimal feature fusion error in the validation set, and constraining the total weight coefficient sum to 1. The mean of the statistical distribution characteristics of the environmental state indicators of each grid cell is multiplied by the corresponding weight coefficient to obtain the statistical feature weight value. The dynamic correlation strength value of the grid cell with adjacent grid cells is multiplied by the corresponding weight coefficient, and then the correlation strength weight values of all adjacent grid cells are summed to obtain the total correlation strength weight value. The statistical feature weight value is added to the total correlation strength weight value to obtain the dynamic weight value of each grid cell. The dynamic weight values of all grid cells are arranged according to grid number to form a spatial dynamic weight matrix (dimension: number of grids × number of grids), with rows corresponding to the target grid and columns corresponding to adjacent grids.
[0061] b12 performs element-wise multiplication of the spatial dynamic weight matrix and the preliminary index set to achieve weighted feature fusion in the spatial domain, resulting in a weighted feature map. Specifically, this involves: first, extracting the dimensional parameters of the spatial dynamic weight matrix and the preliminary index set, determining that the spatial dynamic weight matrix is a two-dimensional structure, with the row dimension corresponding to the total number of target grids and the column dimension corresponding to the total number of adjacent grids; and the preliminary index set is a two-dimensional structure, with the row dimension corresponding to the total number of target grids and the column dimension corresponding to the total number of features. A tensor dimension compatibility check is then performed to verify whether the row dimension of the spatial dynamic weight matrix and the row dimension of the preliminary index set are completely equal, ensuring that the index identifier of each target grid in the two types of data corresponds one-to-one, thus meeting the core requirement of dimension matching for broadcast operations.
[0062] The broadcast expansion rules are determined, and the spatial dynamic weight matrix is expanded from a two-dimensional structure to a three-dimensional tensor. The expansion dimension is the feature dimension. The expansion method is to copy the associated weight values between each grid in the matrix equally along the feature dimension, so that each associated weight forms a one-to-one mapping relationship with all feature dimensions of the corresponding grid in the initial index set. The expanded three-dimensional tensor dimension is the number of target grids × the number of adjacent grids × the number of features. At the same time, the initial index set is expanded from a two-dimensional structure to a three-dimensional tensor. The expansion dimension is the adjacent grid dimension. The expansion method is to copy the feature vector of each grid equally along the adjacent grid dimension, so that each feature value forms a one-to-one mapping relationship with the associated weights between the corresponding grids in the spatial dynamic weight matrix. The expanded three-dimensional tensor dimension is consistent with the expanded dimension of the weight matrix.
[0063] The target grid dimensions of the 3D weight tensor are traversed in row-major order. The 3D subtensor corresponding to each target grid is processed sequentially. The dimension of the subtensor is the number of adjacent grids × the number of features. For each target grid's 3D subtensor, an element-wise multiplication operation is performed with the corresponding target grid's 3D subtensor in the 3D index tensor. That is, the weight value and feature value corresponding to the same adjacent grid index and the same feature index in the subtensor are directly multiplied. During the operation, 6 significant decimal places are retained, and the rounding rules are used to handle the last digits to ensure controllable calculation accuracy.
[0064] After element-wise multiplication, the three-dimensional product subtensor of each target grid is summed according to the dimensions of adjacent grids. The weighted contributions of all adjacent grids to each feature dimension of the target grid are summarized, maintaining a precision of 6 decimal places during the summation process to avoid precision loss. After traversing all target grids and completing the above operations, the summation result of each target grid is organized into a one-dimensional feature vector, with the vector dimension being the number of features. All one-dimensional feature vectors are arranged sequentially according to the target grid index order to form a weighted feature mapping with a dimension of target grid number × feature number. This mapping achieves deep integration of spatial dynamic correlation into feature expression by aggregating the element-wise weighted sum of the initial index set with the contributions of adjacent grids through a spatial dynamic weight matrix, thereby achieving precise adjustment of the initial index and enhancing the expression effect of spatial correlation features.
[0065] b13, based on weighted feature mapping, performs context calibration on the feature propagation of adjacent network nodes to obtain context-aware features; the context-aware features are then compressed and standardized to obtain optimized comprehensive state indicators. Specifically, this includes: first, based on the dynamic association strength between adjacent grids obtained in b10, constructing a feature propagation network for adjacent grid nodes based on a graph neural network, with each spatial grid cell as a network node, each node carrying the weighted feature vector output by b12, and using the existence of adjacency relationships (including direct adjacency and indirect adjacency within 3 orders) between grids as network edges, with the initial weight of the edge assigned to the corresponding dynamic association strength between the grids; defining the 3rd order neighborhood: the 1st order neighborhood is a directly adjacent grid that shares an edge or vertex with the target grid, the 2nd order neighborhood is a grid that is directly adjacent to the 1st order neighborhood grid but not included in the 1st order neighborhood, and the 3rd order neighborhood is a grid that is directly adjacent to the 2nd order neighborhood grid but not included in the 1st to 2nd order neighborhood, setting the propagation distance threshold to 3 grids, i.e., only allowing features to propagate within the 3rd order neighborhood.
[0066] The calculation rules for the propagation weight coefficients are defined. Centered on the target grid, the propagation weight coefficient of the first-order neighborhood is the corresponding dynamic association strength multiplied by 0.8 to the power of 1, the propagation weight coefficient of the second-order neighborhood is the corresponding dynamic association strength multiplied by 0.8 to the power of 2, and the propagation weight coefficient of the third-order neighborhood is the corresponding dynamic association strength multiplied by 0.8 to the power of 3. After calculation, the validity of the weight coefficient of each neighborhood is verified, and the propagation paths corresponding to the weight coefficients less than 0.01 are eliminated to avoid interference of weak associations on feature propagation.
[0067] The feature propagation process is initiated. Each grid cell splits its weighted feature vector into multiple feature components according to the propagation weight coefficients of each neighborhood, and transmits them to the corresponding neighboring nodes in the 3rd order neighborhood. At the same time, each grid node receives the feature components transmitted by all neighboring nodes in its 3rd order neighborhood, classifies and statistically analyzes all received feature components, and records the source grid, propagation order and propagation weight coefficient of each feature component.
[0068] Calculate the weighted normalization coefficients of the received features by summing the propagation weight coefficients of all feature components received by the node to obtain the total weight. For each received feature component, divide its propagation weight coefficient by the total weight to obtain the normalized weight of that feature component, ensuring that the sum of the normalized weights of all received feature components is 1. Perform a weighted average calculation on the received feature components according to the feature dimension. For each feature dimension, multiply the value of all received feature components in that dimension by the corresponding normalized weight, and then sum all the products to obtain the calibrated value for that dimension. Repeat this process for all features. The context-aware feature vector of the grid node is formed by the weighted average calculation of the dimensions, realizing the interactive calibration of features of adjacent nodes. The context-aware feature vectors of all grid nodes are compressed using a 2×2 max pooling kernel with a pooling step size of 1. Each context-aware feature vector is divided into four groups according to four consecutive feature dimensions, resulting in a total of four groups. The four feature values in each group are compared, and the feature value with the largest value is selected as the pooling result of the group. The pooling operation is completed for all groups in sequence to obtain compressed features with the feature dimensions reduced to one-quarter of the original feature number.
[0069] The min-max standardization method is used to unify the scale of the compressed features. First, the minimum and maximum values of each dimension of the compressed features of all grid nodes are calculated. For each dimension value of the compressed features of each grid node, the minimum value of the corresponding dimension is subtracted from the value. Then, the difference is divided by the difference between the maximum and minimum values of the corresponding dimension to map the feature value to the interval between 0 and 1. Six decimal places are retained in the calculation process. Finally, the feature vectors of all grid nodes after standardization are arranged in order of grid number to obtain an optimized comprehensive state index with a dimension of one-quarter of the number of grids and features. This ensures that the index has contextual relevance, dimensional simplification and scale uniformity.
[0070] b14, based on the optimized comprehensive state index, performs feature upsampling and spatial resolution enhancement to obtain a high-resolution feature map. Specifically, this includes: first, extracting the dimensional parameters of the optimized comprehensive state index to determine its corresponding original feature map resolution, which is directly related to the number of grids. For example, when the number of grids is 1000×1000, the original feature map resolution is 1000×1000. The target resolution is determined based on the spatial detail requirements of the application scenario, such as 2000×2000. The upsampling factor is calculated by the ratio of the target resolution to the original resolution. When the target resolution is twice the original resolution, the upsampling factor is determined to be 2. Simultaneously, the optimized comprehensive state index is organized into a two-dimensional feature map format according to the grid number order. The row dimension of the feature map corresponds to the north-south distribution of the grid, and the column dimension corresponds to the east-west distribution of the grid. Each pixel carries the feature value of the corresponding grid, ensuring that the data structure is compatible with the subsequent upsampling algorithm.
[0071] If bilinear interpolation is used for upsampling, the specific calculation process is as follows: Create a blank interpolation feature map according to the target resolution, with the number of rows and columns of pixels matching the target resolution. For each pixel to be interpolated in the interpolation feature map, calculate its mapped coordinates in the original feature map using the upsampling factor, i.e., original coordinates = interpolation point coordinates / upsampling factor. Determine the four nearest original feature map pixels around the interpolation point based on the mapped coordinates (i.e., the four adjacent original pixels, top, bottom, left, and right), and record the row and column coordinates and corresponding feature values of these four original pixels in the original feature map. Calculate the difference between the row and column coordinates of the interpolation point and each adjacent original pixel. The horizontal and vertical distances are obtained, with the distance in pixels. Weights are assigned according to the inverse square law rule, meaning the smaller the squared distance between the interpolation point and the original pixel, the larger the corresponding weight. After weight calculation, normalization is performed to ensure that the sum of the weights of the four original pixels is 1. The feature value of each original pixel is multiplied by its corresponding weight to obtain four weighted feature values. These four weighted feature values are summed to obtain the feature value of the interpolation point, retaining six decimal places during the calculation. All interpolation points in the interpolation feature map are processed one by one in row priority order to fill the pixel gaps in the low-resolution feature map, finally generating a high-resolution feature map consistent with the target resolution.
[0072] If a transposed convolutional layer is used to perform deconvolution for upsampling, the specific calculation process is as follows: Preset the core parameters of the transposed convolutional layer: the kernel size is set to 3×3, the stride is set to 2, the padding method is edge padding, and the number of padding pixels is set to 1, ensuring that the output feature map resolution is exactly twice that of the input feature map; initialize the 3×3 kernel weights using the Xavier initialization method, adaptively allocating initial values based on the input and output feature dimensions to avoid training instability caused by excessively large or small initial weights; use the two-dimensional feature map corresponding to the optimized comprehensive state index as the input feature map, and pad the edges of the input feature map by adding one pixel to each of the four edges (top, bottom, left, and right). The padding value is taken as the average of the edge pixels of the input feature map to avoid loss of edge information. Deconvolution operation is performed on the padded input feature map with a stride of 2. The 3×3 convolution kernel is slid to cover each pixel region of the input feature map in turn. Each element of the convolution kernel is multiplied by the pixel value of the corresponding region of the input feature map. All the product results are summed to obtain the intermediate result of the convolution operation. According to the sliding rule of stride of 2, the intermediate results are arranged in the output feature map with a 1-pixel interval to automatically fill the pixel gaps. At the same time, the weight distribution of the convolution kernel is balanced to avoid the generation of checkerboard artifacts after upsampling. After completing all convolution sliding operations, the output feature map with the same resolution as the target resolution (e.g., 2000×2000) is obtained.
[0073] After completing the upsampling process using either of the two methods described above, the output feature map is reorganized into a tensor format according to the grid number and feature dimension to obtain a high-resolution feature map. Its dimension is the number of target grids (the row and column product corresponding to the target resolution) × the number of target features (consistent with the number of features of the optimized comprehensive state index). The spatial resolution of this feature map is significantly improved compared to the original feature map, and the spatial detail information is richer, which can accurately characterize the feature differences of different grid units.
[0074] b15, based on high-resolution feature mapping, obtains a continuously spatially distributed feature field through a spatial interpolation algorithm. Specifically, it includes: firstly, extracting the grid parameters corresponding to the high-resolution feature mapping, determining the spatial coordinate boundaries, size, and feature values of each grid cell. This feature mapping has covered the main spatial range of the predefined key area, with only blank spatial regions between grid cells; according to the spatial detail expression requirements, the blank spatial regions and the original grid coverage areas are uniformly divided into 0.5m×0.5m uniform interpolation grids. Each interpolation grid corresponds to a spatial point with a feature value to be estimated. The latitude and longitude coordinates of all interpolation grid points are recorded to ensure that the interpolated feature field can completely cover the entire key area.
[0075] If the ordinary Kriging interpolation algorithm is selected for eigenvalue estimation, the specific calculation process is as follows: First, preprocess all grid cell eigenvalues in the high-resolution feature map to remove outliers. Then, using the 3σ criterion, retain the valid eigenvalues and their corresponding grid coordinates as interpolation sample points. Next, calculate the spatial distance between any two sample points and the difference in their corresponding eigenvalues. Based on these statistical data, fit a Gaussian variogram to determine the sill value, nugget value, and range parameters of the variogram, ensuring a goodness-of-fit R². 2 The value is greater than or equal to 0.95. For each interpolation grid point, 8 valid sample points (i.e., the 8 nearest grid cells) are selected from the nearest to the farthest spatial distance, with that point as the center. If there are fewer than 8 sample points within the range, all sample points within the range are used as the interpolation basis. The spatial correlation between the interpolation point and each selected sample point is calculated based on the fitted Gaussian variogram. A Kriging equation system is constructed by combining the spatial distribution law of the sample point feature values. The equation system is solved to obtain the weight coefficient corresponding to each sample point, and the sum of the weight coefficients is 1. The feature value of each sample point is multiplied by the corresponding weight coefficient, and all products are accumulated to obtain the feature estimate of the interpolation grid point. The average absolute error between the estimated value and the feature values of the 3 nearest surrounding sample points is calculated. If the error is less than or equal to 5%, the estimated value is retained. If the error exceeds 5%, 10 nearest sample points are selected again and the above estimation process is repeated until the error requirement is met.
[0076] If the radial basis function interpolation algorithm is selected for eigenvalue estimation, the specific calculation process is as follows: Determine the smoothing parameter of the Gaussian radial basis function, which is set to 1 / 2 of the grid size in the high-resolution feature map. For example, if the original grid size is 1m, the smoothing parameter is 0.5m. For each interpolation grid point, sort and select the 8 nearest valid sample points (grid cells in the high-resolution feature map) by spatial distance, and record the eigenvalues of the sample points and their spatial distances from the interpolation points. Substitute the spatial distance between each sample point and the interpolation point into the Gaussian radial basis function to calculate the correlation degree between the sample point and the interpolation point (the correlation degree decreases as the distance increases). Normalize the correlation degree of all sample points. The process involves processing the data to obtain the weight coefficients for each sample point. The feature value of each sample point is multiplied by its corresponding weight coefficient, and all products are summed to obtain the estimated feature value for that interpolated grid point. Similarly, the difference between the estimated value and the feature values of surrounding sample points is checked to ensure that the difference is less than or equal to 5%. If this is not met, the smoothing parameter is adjusted (within 0.8 to 1.2 times the original parameter) and recalculated until the continuity and consistency requirements are met. Feature value estimation for all interpolated grid points is performed sequentially in row-major order. Finally, the grid feature values of the original high-resolution feature map are integrated with the estimated feature values of all interpolated grid points to form a continuous spatial distribution feature field covering the entire key area without spatial discontinuities.
[0077] b16, based on a continuously spatially distributed feature field, maps the feature field values to corresponding color space values to obtain a raster image; based on the raster image, calculates the gradient distribution of the feature field to obtain contour vector data; and performs layer overlay and semantic annotation processing on the raster image and contour vector data to obtain a global environmental situation awareness map, specifically including: firstly, establishing the feature field values and red-green-blue (Red-Green-Blue) coordinates... The one-to-one mapping rule for RGB color space values is as follows: First, the eigenvalue range of all spatial points in the continuous spatial distribution feature field is statistically analyzed and normalized to the interval of 0 to 1. Then, five consecutive numerical intervals are divided: 0 to 0.2, 0.2 to 0.4, 0.4 to 0.6, 0.6 to 0.8, and 0.8 to 1.0. A corresponding RGB color value is assigned to each numerical interval: 0 to 0.2 corresponds to RGB255, 255, 255; 0.2 to 0.4 corresponds to RGB200, 200, 255; 0.4 to 0.6 corresponds to RGB150, 150, 255; 0.6 to 0.8 corresponds to RGB100, 100, 255; and 0.8 to 1.0 corresponds to RGB50, 50, 255. This ensures that higher values correspond to higher color saturation and that the color transition between adjacent intervals is natural.
[0078] Traverse each spatial point in the continuous spatial distribution feature field, read the feature value of that point, determine its numerical range, and determine the corresponding RGB color value according to the mapping rules. The color value is an integer (ranging from 0 to 255). Create a blank raster image at a resolution of 0.5m / pixel. The number of rows and columns of the raster image is calculated based on the latitude and longitude range of the key area and the pixel resolution (number of rows = north-south span of the area / 0.5m, number of columns = east-west span of the area / 0.5m). Write the RGB color value of each spatial point into the corresponding pixel position of the raster image according to its latitude and longitude coordinates. After completing the color value mapping of all spatial points, generate a raster image in Portable Network Graphics (PNG) format. This image intuitively reflects the spatial distribution of the feature field.
[0079] The Sobel operator is used to calculate the gradient values of each spatial point in the feature field. The specific process is as follows: a 3×3 Sobel horizontal and vertical operator is constructed. A 3×3 neighborhood spatial point is selected with each spatial point as the center, and the feature values of each point in the neighborhood are extracted. The neighborhood feature values are convolved with the horizontal operator to obtain the horizontal gradient value of the point. The neighborhood feature values are convolved with the vertical operator to obtain the vertical gradient value of the point. The gradient intensity is calculated by taking the square root of the sum of the squares of the horizontal and vertical gradient values, and retaining three significant decimal places. The contour interval is set according to the gradient intensity: the contour interval is set to 0.1 when the gradient intensity is less than or equal to 0.1, and the contour interval is set to 0.05 when the gradient intensity is greater than 0.1. All spatial points in the feature field are traversed, and spatial points with equal feature values are extracted. These points are connected in spatial order to form closed or continuous contour lines. All contour lines constitute contour line vector data, which contains the coordinate string of the contour lines, the corresponding feature values, and the interval information.
[0080] Using the layer blending function of Geographic Information System (GIS) software, the generated raster image was imported as the base map layer, and then the contour vector data was imported as the overlay layer into the GIS software. The transparency of the overlay layer was adjusted to 50% to ensure that the base map colors and the overlay contour lines were clearly visible. Semantic annotation information was added, including grid number (marked at the center of the corresponding grid according to the original grid number), indicator name (marked in the upper left corner of the image), numerical range (marked in the upper right corner of the image), and legend (marked in the lower right corner of the image, including the numerical range and corresponding color block). The annotation font was Arial, and the font size was set to 10pt. The annotation positions avoided areas of abrupt change in feature values. The gradient strength was used to determine that areas with a gradient strength greater than 0.3 were abrupt change areas to avoid the annotations obscuring key feature distribution information. After completing the layer overlay and semantic annotation, the results were exported as a global environmental situational awareness map in GeographicTagged Image File Format (GeoTIFF). This format supports latitude and longitude coordinate positioning and feature value attribute query functions to ensure the practicality and ease of use of the situational awareness map.
[0081] In this embodiment of the invention, by dividing a continuous data stream into a sequence of equally spaced time slices, an ordered time frame can be established, improving the controllability of data temporal sequence; spatial registration is performed on multimodal data within each time slice to establish spatial reference data under a unified coordinate system; spatial misalignment of multi-source data is eliminated, and spatial correlation of data is strengthened, providing a unified foundation for artificial intelligence to integrate spatially heterogeneous data; spatiotemporal feature matrices, flow feature vectors, and regional average dynamic features are extracted for different types of spatial reference data; the dimensions of feature expression are enriched, and the unique characteristics of various types of data are explored; multi-source feature vectors are obtained by dimensional alignment and vectorization of multiple features; the problem of inconsistent dimensions of multi-source features is solved, and a structured feature set is constructed to reduce complexity.
[0082] Attention weight matrix is calculated using a fully connected layer and softmax function; key feature components are automatically identified to increase the weight ratio of core information; the attention weight matrix is weighted and summed with multi-source feature vectors to obtain attention-weighted fusion features; key feature expression is strengthened and irrelevant feature interference is suppressed; the attention-weighted fusion features are dimensionality reduced and standardized to obtain a preliminary index set, simplifying feature dimensions and unifying feature scale; a regular grid topology structure is constructed by combining key region boundaries to map abstract indicators to specific spatial units, establishing a spatial structured framework to help understand the spatial distribution patterns of data; the preliminary index set is mapped to each grid unit to calculate the statistical distribution characteristics of environmental state indicators, explore the state distribution patterns within the grid, and provide data support; the dynamic correlation strength of adjacent grid units is calculated through spatial autocorrelation analysis to capture the dynamic dependency relationship between grids, construct a spatial correlation network, and enhance the perception of spatial linkage features; the spatial dynamic weight matrix is obtained by combining statistical distribution characteristics and dynamic correlation strength to achieve spatial adaptive adjustment of weights, optimize spatial feature weighting logic, and improve the spatial adaptability of feature fusion.
[0083] By multiplying the spatial dynamic weight matrix element-wise with the preliminary index set, a weighted feature map is obtained, integrating spatial dynamic correlation into the feature expression and enhancing the spatial differentiation representation capability of the features. Context calibration is performed on the feature propagation of adjacent nodes, and after compression and standardization, an optimized comprehensive state index is obtained, improving the contextual consistency of the features. Feature upsampling and spatial resolution enhancement are performed on the comprehensive state index to obtain a high-resolution feature map, strengthening the spatial detail expression of the features and laying the foundation for generating a fine-grained situational awareness map. The high-resolution feature map is transformed into a continuously distributed feature field through a spatial interpolation algorithm, and discrete grid features are transformed into continuous spatial representations. The feature field is mapped to a raster image, contour vector data is calculated, and layers are overlaid and semantically labeled to obtain a global environmental situational awareness map, constructing an intuitive and semantically rich situational representation and improving the interpretability of the decision-making process.
[0084] In some embodiments, step 300 above involves inputting a global environmental situation awareness map into a pre-trained multi-objective collaborative decision-making model; the multi-objective collaborative decision-making model forms a multi-objective decision feature set by analyzing the resource entities and relationships in the global environmental situation awareness map; and based on the multi-objective decision feature set, it performs decision optimization to obtain a comprehensive collaborative scheduling scheme, including: c1-c5.
[0085] c1. Analyze the overall environmental situational awareness map, extract key state indicators to characterize task processing efficiency, resource utilization balance, and service demand satisfaction, and construct a multi-objective decision feature set. Specifically, this includes: first, loading the overall environmental situational awareness map, which contains spatially aligned and feature-fused multimodal data association information; and then using a Geographic Information System (GIS)... The System, GIS) parsing interface extracts resource entity information from the map, including the identifiers, spatial coordinates, and operational status data of entities such as fixed sensors, mobile devices, and task execution units. It also parses the relationships between entities, covering spatial proximity, task dependency, and resource supply and demand relationships. Based on the core objectives of the application scenario, extraction rules for three types of key status indicators are determined: task processing efficiency is calculated by statistically analyzing the number of tasks completed per unit time, the average time from task receipt to completion, and the task queue length; resource utilization balance is obtained through statistical analysis of the load rate, resource occupancy time, and idle time percentage of each resource unit; and service demand satisfaction is determined by parameters such as the ratio of the number of demand responses meeting the target to the total demand, and whether the demand response delay is within a preset threshold. Data cleaning is performed on the extracted indicators to remove outliers and missing values, unify the data type and dimensions of the indicators, and construct a structured multi-objective decision feature set according to entity identifier, timestamp, and indicator category. The feature set is stored in tensor format for easy reading and processing by subsequent models.
[0086] c2, based on a multi-objective decision feature set, constructs a global multi-objective optimization function that simultaneously optimizes operational efficiency, resource cost, and service quality, and sets operational constraints for each functional unit. Specifically, based on the multi-objective decision feature set, it determines the core dimensions of the global multi-objective optimization: the operational efficiency dimension aims to maximize task throughput and minimize average processing latency; the resource cost dimension aims to minimize total energy consumption and equipment loss coefficient; and the service quality dimension aims to maximize demand satisfaction and optimize user satisfaction. A weighted summation method is used to construct the global multi-objective optimization function. The weight coefficients are determined using the analytic hierarchy process (AHP) combined with scenario requirements. A judgment matrix is constructed using a 1-9 scale. After normalization and consistency testing (consistency ratio (CR) less than or equal to 0.1), the final weights are determined to ensure that the importance of each optimization objective matches the actual application requirements.
[0087] The constraint system construction is optimized using a fusion constraint convex set projection algorithm. The specific calculation process is as follows: First, each type of constraint is mapped to a corresponding convex set. The hardware constraint convex set is bounded by the maximum load threshold of the device and the upper limit of the computing power; the time constraint convex set is bounded by the longest allowable delay for task response and the shortest interval for resource scheduling; and the resource constraint convex set is bounded by the total energy supply limit and the quota for the use of key consumables. For each convex set, mathematical boundary conditions are defined to ensure that the constraints satisfy the closure and convexity requirements of the convex set. The intersection of each type of constraint convex set is calculated to obtain the global constraint convex set. By randomly generating 100 to 200 sets of candidate solutions, covering combinations of parameters such as device load, task delay, and resource consumption, the candidate solutions are substituted into the global constraint convex set for projection calculation. The shortest distance from the candidate solution to the boundary of the convex set is calculated. If the distance from a candidate solution to the boundary of the convex set is 0... If the distance from the candidate solution to the convex set is greater than 0, it indicates that the constraint convex set is empty. In this case, the rigid constraint threshold needs to be reduced by 5% to 10%, such as adjusting the maximum load threshold of the equipment from 80% to 85% of the rated load. The constraint convex set is then reconstructed and the projection verification is repeated until a candidate solution falls within the convex set. Finally, the boundary thresholds of each constraint condition are determined. The hardware constraint threshold is set based on the equipment performance parameters with a 20% redundancy. The time constraint threshold is determined based on the service level agreement to determine the upper limit of the response delay of 95% of the tasks. The resource constraint threshold is set by deducting a 15% fluctuation reserve from the resource supply capacity. A complete constraint system that is both feasible and stringent is formed. The constraint convex set projection algorithm ensures that the optimization process is carried out within the feasible region of the convex set, guaranteeing that the global multi-objective optimization function has an optimal solution.
[0088] c3 decomposes the global multi-objective optimization function into local value functions associated with each agent; each agent generates a local policy based on its local value function and the environment state; all local policies are integrated by a centralized coordinator, which uses an attention mechanism to evaluate their global cooperative utility and obtains a non-dominated solution set as the initial cooperative policy set through iterative search within the constrained solution space. Specifically, this includes: adopting a multi-agent reinforcement learning paradigm based on value function decomposition, and selecting a decomposition-based multi-objective evolutionary algorithm (Multi-Objective Evolutionary Algorithm based on...). MOEAD (Multi-Objective Decomposition) divides the global multi-objective optimization function according to the agent's responsibility. Agents are divided according to functional unit type or spatial region. Each agent corresponds to a local sub-objective. During the decomposition process, it is ensured that the objective direction of each local value function is consistent with the global optimization function, and the cumulative result of all local value functions is equivalent to the global optimization function. An independent reinforcement learning policy network is configured for each agent. The network input is the local features of the agent's responsibility area corresponding to the multi-objective decision feature set and the real-time environmental state. The environmental state includes the current resource occupancy, task queue length, service demand change trend, etc. Based on the optimization objective of its own local value function, each agent generates a local policy through the policy network. The local policy includes specific operation instructions such as resource allocation ratio, task execution priority, and device start / stop scheme.
[0089] This paper integrates distributed consensus algorithms to optimize the consistency and effectiveness of local policies. A practical Byzantine fault-tolerant consensus algorithm is selected to construct a distributed consensus mechanism. The consensus participation nodes are set to all agents, the consensus rounds are capped at 3, and the fault tolerance threshold is one-third of the total number of agents, meaning that at most one-third of the total number of agents are allowed to output abnormal policies. Specifically, each agent broadcasts its generated local policy, along with its own identifier and a policy generation timestamp, to all other agent nodes. After receiving the local policies broadcast by all nodes, each agent first verifies the policy's format and the validity of its own identifier, eliminating policies with incorrect formats or invalid identifiers. Then, based on the constraints of the local value function, it checks whether the received policy meets the local sub-objective requirements of its agent and counts the number of valid policies that meet the requirements.
[0090] In the consensus preparation phase, each agent performs a hash operation on the valid policies to generate a policy digest, and then broadcasts the digest along with its own vote (support or opposition) to all nodes. After all nodes have completed the voting broadcast, each agent counts the number of support votes for each local policy and calculates the proportion of support votes to the total number of valid agents. When the support ratio of a local policy reaches two-thirds or more, the policy is deemed to have passed consensus verification. If there are policies whose support ratio does not reach the threshold, feedback is given to the corresponding generating agent, requiring it to adjust its policy parameters based on the reasons for failure in the consensus feedback (such as violation of local constraints or conflict with other policies), regenerate the local policy, and repeat the above broadcasting and verification process until all agents' local policies have passed consensus verification.
[0091] A centralized coordinator is constructed, with a built-in attention mechanism module. This module assigns attention weights by calculating the correlation between each consensus-reaching local policy and the global optimization objective, focusing on local policies that significantly impact global performance. The coordinator imports all consensus-reaching local policies into the constrained solution space, sets an upper limit for the number of iterations and a convergence error threshold, and iteratively optimizes the solution space using a greedy search algorithm. In each iteration, the global collaborative utility of the policy combination is evaluated, and policy combinations that violate the constraints or have low utility are eliminated until the upper limit for the number of iterations is reached or the utility improvement is less than the convergence threshold. The output is a non-dominated solution set that satisfies the Pareto optimality condition, and this solution set is used as the initial collaborative policy set.
[0092] c4. Perform multi-dimensional performance evaluation and comprehensive trade-off analysis on the initial collaborative strategy set, and calculate the comprehensive utility value of each candidate strategy. Sort the comprehensive utility values of each candidate strategy to obtain the comprehensive utility ranking. Specifically, this includes: constructing a hierarchical multi-dimensional performance evaluation system, determining a three-level hierarchical structure, with the target layer being the comprehensive performance evaluation result, the dimension layer corresponding to the three core dimensions of operational efficiency, resource cost, and service quality, and the indicator layer being the specific quantitative indicators under each dimension. The operational efficiency dimension includes task completion rate and average processing latency, the resource cost dimension includes unit task energy consumption and equipment depreciation and loss, and the service quality dimension includes demand satisfaction rate and response compliance rate. All indicators are positively processed to ensure that the higher the value, the better the performance.
[0093] The weights of each level are determined based on the analytic hierarchy process (AHP). A hierarchical model of target layer, dimension layer, and indicator layer is constructed. The relative importance between dimensions and between indicators of the same dimension is compared and scored pairwise using a scale of 1 to 9 to form a corresponding judgment matrix. The maximum eigenvalue and eigenvector of the matrix are calculated, and the initial weights are obtained after normalization. The rationality is verified by a consistency test (CR ≤ 0.1). If it fails, the scores are adjusted until it passes. Finally, the weights of the dimension layer and indicator layer are determined, and the sum of the weights of the same level is 1.
[0094] The comprehensive utility value is calculated using a hierarchical geometric weighted algorithm. The specific process is as follows: For each candidate strategy in the initial collaborative strategy set, it is substituted into the evaluation system one by one. Based on the strategy simulation execution data, the actual scores of each indicator are calculated and normalized to the range of 0 to 1. The calculation is performed step by step according to the hierarchical rules. First, for each dimension, the scores of the two indicators under that dimension are multiplied by the corresponding power of the indicator weights to obtain the dimension geometric weighted score, which is retained to four decimal places. Then, the scores of the three dimensions are multiplied by the corresponding power of the dimension weights to obtain the comprehensive utility value of the candidate strategy, which is retained to four decimal places. The comprehensive utility values of all candidate strategies are sorted in descending order to obtain the comprehensive utility ranking result. At the same time, the original indicator scores, dimension weighted scores, and comprehensive utility values of each strategy are recorded and organized into a structured evaluation report according to the strategy number.
[0095] c5, based on comprehensive utility ranking, selects a dominant strategy from the initial set of collaborative strategies to obtain a comprehensive collaborative scheduling scheme. Specifically, this includes: Firstly, based on the comprehensive utility ranking results, prioritizing the top-ranked candidate strategies and performing secondary verification based on the real-time needs of the application scenario. For emergency response scenarios, the service quality dimension score is the primary focus; for cost-sensitive scenarios, the resource cost dimension score is the primary focus. From the verified candidate strategies, the strategy with the highest comprehensive utility value and relatively balanced scores across dimensions is selected as the dominant strategy. If multiple strategies have the same comprehensive utility value, random sampling combined with scenario adaptability analysis is used to determine the final dominant strategy. The dominant strategy is then broken down into specific executable operational instructions, including resource allocation schemes for each functional unit, task scheduling order, and runtime parameter settings, clarifying the execution subject, execution time, and execution standards. The feasibility of the dominant strategy is verified by simulating resource consumption, task completion, and service quality performance during strategy execution to ensure the strategy meets all constraints and has no logical conflicts. Finally, a structured comprehensive collaborative scheduling scheme is formed, including strategy details, execution steps, expected performance, and emergency adjustment mechanisms, supporting direct import into the system for execution.
[0096] In this embodiment of the invention, a global environmental situational awareness map is analyzed and key state indicators are extracted to construct a multi-objective decision feature set. The visualized spatial situation is transformed into structured decision features, strengthening the correlation between features and task processing efficiency, resource utilization balance, and service demand satisfaction. A global multi-objective optimization function is constructed based on the multi-objective decision feature set, and operational constraints for functional units are set. The optimization direction and boundary limits of the decisions are determined, establishing a standardized multi-objective optimization framework. This makes the decision logic of the artificial intelligence system clearer and improves the systematicness and controllability of the multi-objective optimization task. The global optimization function is transformed into a local value function, generating local strategies which are then integrated by a centralized coordinator to achieve synergistic unity between global objectives and local behaviors, leveraging the advantages of distributed processing and the global control capabilities of centralized coordination. A multi-dimensional performance evaluation is performed on the initial collaborative strategy set, calculating and ranking the comprehensive utility value to establish a quantitative strategy evaluation system, providing an objective basis for artificial intelligence strategy selection. The dominant strategy is selected based on the comprehensive utility ranking, forming a comprehensive collaborative scheduling scheme. This achieves precise selection of the optimal strategy, constructing a complete closed loop from feature extraction to decision output, and strengthening the comprehensive adaptability of the decision scheme.
[0097] In some embodiments, step 400 above involves parsing and encapsulating the comprehensive collaborative scheduling scheme to obtain an executable instruction sequence; the executable instruction sequence is then distributed in parallel to the corresponding decision nodes and control terminals through a distributed communication architecture to complete the collaborative release of real-time resource scheduling and control instructions, including d1-d4.
[0098] d1 performs semantic parsing and instruction conversion on the comprehensive collaborative scheduling scheme, generating an executable instruction sequence containing device control instruction sets, resource scheduling instruction sets, and service publishing instruction sets. Specifically, this includes: first, loading the comprehensive collaborative scheduling scheme, which contains structured information such as global resource allocation strategies, device operating parameter configurations, service execution standards, and constraints; then, using a semantic parsing engine to decompose the scheme layer by layer, first extracting the decision objectives, execution subject identifiers, operation object types, and core parameters; then, decomposing the logical relationships and execution timing requirements of each decision item; and finally, based on a pre-defined instruction conversion rule base, mapping the decomposed decision information into specific executable instructions, generating three types of instruction sets. The device control instruction set contains various execution... The system includes a unique identifier for each device, an operation type, target parameter values, and execution time limits. Operation types include start, stop, and parameter adjustment. The resource scheduling instruction set includes resource type, allocation object identifier, allocation ratio, occupation duration, and release conditions. Resource types include computing resources, energy resources, and storage resources. The service publishing instruction set includes service identifier, service provider, target service object, service start threshold, and service quality monitoring indicators. The generated instruction set undergoes syntax validation and logical conflict detection, eliminating instructions with format errors, parameters exceeding device thresholds, and logical contradictions. Missing necessary fields are added, and the instructions are sorted according to execution sequence to form an ordered sequence of executable instructions. The sequence is stored in JavaScript Object Notation (JSON) format to ensure readability and parsability.
[0099] d2, based on a preset communication protocol, encapsulates the executable instruction sequence into standardized instruction data packets recognizable by each control terminal. Specifically, this includes: pre-configuring the communication protocol; researching the compatibility, transmission efficiency, and terminal adaptation range of mainstream industrial-grade communication protocols; considering the terminal type characteristics of target application scenarios such as government services, AI-driven aquaculture, smart homes, and digital agriculture; selecting stable and widely adaptable protocol types, including Message Queuing Telemetry Transport (MQTT), Hypertext Transfer Protocol version 2 (HTTP / 2), and Modicon Bus Transmission Control Protocol (ModbusTCP); collecting parameters such as hardware models, operating system versions, communication interface specifications, and data transmission requirements of various control terminals; establishing an adaptation rule base for terminal parameters and protocol characteristics; and determining the optimal communication protocol for different parameter combinations. Finally, the selected protocol types and adaptation rule base are solidified into the system configuration module, forming a preset communication protocol set that supports automatic matching and invocation during subsequent terminal access.
[0100] Based on the hardware model, operating system, and communication interface capabilities of the control terminal, the system's preset adaptation rule base is invoked to match the corresponding preset communication protocol for different types of control terminals, ensuring communication compatibility between the protocol and the terminal. A unified structure for standardized instruction data packets is defined, consisting of a header, body, and checksum. The header contains metadata such as the terminal's unique identifier, instruction set type, data packet length, and timestamp. The body is a binary data stream of the executable instruction sequence adapted to the protocol. The checksum is calculated using a 32-bit Cyclic Redundancy Check (CRC32) algorithm to verify the integrity of data transmission. The executable instruction sequence is serialized according to the matched communication protocol, converting the JSON-formatted instruction sequence into binary data supported by the corresponding protocol. The header information is filled according to the defined data packet structure, the serialized binary instruction data is written into the body, and the checksum is calculated and filled. The encapsulated standardized instruction data packets are format-verified, checking the integrity of the header fields, the consistency between the body data length and the identifier length, and the validity of the checksum. After verification, the data packets are stored according to terminal type to prepare for subsequent parallel distribution.
[0101] d3. Standardized instruction data packets are distributed in parallel to the corresponding decision nodes and control terminals through a distributed message middleware. Specifically, this includes: deploying a distributed message middleware cluster, selecting Kafka as the core middleware, configuring multiple broker nodes to achieve load balancing and fault redundancy, creating dedicated message queues based on terminal identifiers and decision node numbers, and establishing a one-to-one mapping between instruction data packets and message queues; starting the data packet distribution service, reading the categorized and stored standardized instruction data packets, obtaining the target terminal and decision node identifiers corresponding to each data packet, and delivering the data packets to the corresponding message queues according to the mapping relationship; and enabling a parallel distribution mechanism through the middleware... The system features multi-threaded processing capabilities, simultaneously pushing instruction data packets to multiple message queues. During distribution, a flow control strategy dynamically adjusts the data packet sending rate based on real-time network bandwidth to prevent network congestion. The middleware cluster monitors the distribution process in real-time, recording the distribution time, target queue, and sending status of each data packet (success, failure, or pending retry). For failed distribution packets, a preset retry mechanism (maximum of 3 retries, with a 1-second interval between each retry) is used to re-push the packet until successful distribution or the retry limit is reached. Decision nodes and control terminals listen to the corresponding message queues via long connections, receiving instruction data packets in real-time to ensure rapid instruction transmission and synchronous reception.
[0102] d4. Receive status feedback data from each control terminal and dynamically adjust the distribution strategy of subsequent instructions based on the status feedback data through a feedback control mechanism to complete the real-time scheduling of global resources and the coordinated release of control instructions. Specifically, this includes: establishing a status feedback receiving interface based on Transmission Control Protocol / Internet Protocol (TCP / IP), which supports high-concurrency data reception; allocating an independent feedback data receiving port to each control terminal to ensure orderly reception of feedback data; receiving status feedback data returned by each control terminal, including terminal identifier, instruction reception status, instruction execution progress, current device operating parameters, resource occupancy, service execution results, and abnormal alarm information; unpacking and parsing the received feedback data, extracting key status indicators, and removing invalid and duplicate data; and constructing a feedback control mechanism, which includes a status evaluation module and a strategy adjustment module. The status evaluation module compares the parsed status indicators with the preset targets of the coordinated scheduling scheme to evaluate the compliance of instruction execution, including device response timeliness, resource allocation rationality, and service quality compliance rate, and identifies execution deviations and abnormal problems.
[0103] The optimization logic is based on the integration of a dynamic adjustment algorithm using Bézier curves. The specific calculation process is as follows: First, key state indicator data, historical adjustment strategies, and corresponding execution effects from the most recent 5 to 10 feedback cycles are collected to establish a mapping dataset between adjustment parameters and state deviations, providing data support for Bézier curve parameter configuration. Curve parameter rules are set for different execution deviation scenarios, with a uniform third-order Bézier curve. The curve shape is determined by four control points, with the selection of control points based on historical optimal operating state data, current state data, target state data, and stable threshold data. If the terminal feedback execution delay exceeds the threshold, the delay deviation value, i.e., the actual delay, is calculated first. The difference from the preset threshold is used to select four control points based on the mapping dataset. The first control point is the distribution priority corresponding to the lowest historical latency, the second control point is the current distribution priority, the third control point is the target priority calculated according to the latency deviation ratio, and the fourth control point is the priority safety threshold when the system is running stably. A third-order Bézier curve is plotted based on the coordinate values of the four control points. The adjustment period is divided into 10 to 15 equal time steps. The priority discrete value corresponding to each time step on the curve is calculated. At the same time, the network bandwidth increment is allocated according to the slope of the curve. The larger the slope, the more obvious the latency deviation, and the larger the bandwidth allocation increment, so as to achieve smooth linkage adjustment between priority and bandwidth.
[0104] If resource occupancy is too high, calculate the occupancy deviation value, which is the difference between the actual occupancy rate and the preset safety threshold. Select four control points: the resource allocation ratio corresponding to the historical optimal occupancy rate, the current resource allocation ratio, the target allocation ratio calculated based on the occupancy deviation, and the minimum allocation ratio threshold corresponding to safe resource operation. Plot a third-order Bézier curve, divide the time step according to the adjustment period, calculate the discrete value of the resource allocation ratio corresponding to each step, and gradually reduce the resource occupancy ratio of non-critical tasks according to the curve trend to ensure stable resource supply for critical tasks and avoid task interruption caused by sudden changes in allocation ratio. If an equipment abnormality alarm occurs, select four control points: the normal instruction sequence position before the abnormality occurs, the current instruction sequence position, the target position after the emergency instruction is inserted, and the sequence connection position after normalization. Plot a second-order Bézier curve to simplify the number of control points and improve response speed. Calculate the smooth transition path for emergency instruction insertion, determine the instruction execution interval and insertion timing for each time step, insert the emergency adjustment instruction at the beginning of the instruction sequence according to the curve planning path, and adjust the execution sequence of subsequent instructions to ensure a smooth connection between emergency handling and regular tasks.
[0105] After each adjustment cycle, the adjustment effect is verified based on the new status feedback data, the control point coordinates of the Bézier curve are corrected, and the adjustment curve for the next round is recalculated to form a dynamic calibration mechanism. The adjusted distribution strategy calculated according to the Bézier curve is updated to the scheduling and configuration center of the distributed message middleware. The middleware executes subsequent instruction distribution according to the new strategy, continuously receiving feedback data, evaluating the execution status, and optimizing the distribution strategy through the Bézier curve to achieve real-time dynamic scheduling of global resources and collaborative release of control instructions.
[0106] In this embodiment of the invention, semantic parsing and instruction conversion are performed on the collaborative scheduling scheme to generate multiple types of executable instruction sequences; abstract decision schemes are transformed into structured and concrete execution instructions, enhancing the adaptability of instructions to devices, resources, and services; standardized instruction data packets are encapsulated based on a preset communication protocol; unified instruction transmission format and specifications are implemented to break down protocol barriers between different control terminals; instruction data packets are distributed in parallel through a distributed message middleware; the parallel processing advantages of the distributed architecture are leveraged to shorten the instruction issuance time across multiple nodes; status feedback data is received and the instruction distribution strategy is dynamically adjusted; and a closed-loop control mechanism of decision-making, execution, and feedback is constructed to ensure that instruction distribution aligns with the real-time operating status, improving the real-time adaptive capability of the artificial intelligence system and enhancing the dynamic optimization level of global resource scheduling.
[0107] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the aforementioned functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0108] Embodiments of the present invention also provide a perception-cooperative decision-making system based on multimodal heterogeneous data fusion, such as... Figure 2 The diagram shown illustrates a perception-cooperative decision-making system based on multimodal heterogeneous data fusion, as provided in an embodiment of the present invention. The system includes:
[0109] The multi-source acquisition module is used to acquire multimodal heterogeneous data, including images, text, time-series data, and graph structure data; it preprocesses the multimodal heterogeneous data to obtain a standardized data stream.
[0110] The fusion analysis module is used to parse and extract features from the standardized data stream, and then align and encode it to obtain multi-source feature vectors. These multi-source feature vectors are then weighted and fused to obtain a preliminary indicator set. Based on this preliminary indicator set, predefined key areas are gridded, and the dynamic correlation characteristics of each spatial grid cell are analyzed to obtain a spatial dynamic weight matrix. This spatial dynamic weight matrix is then weighted and fused with the preliminary indicator set, and optimized using spatial context calibration to obtain a comprehensive state indicator. Finally, based on the optimized comprehensive state indicator, a global environmental situational awareness map is generated.
[0111] The collaborative decision-making module is used to input the global environmental situational awareness map into the pre-trained multi-objective collaborative decision-making model. The multi-objective collaborative decision-making model forms a multi-objective decision feature set by analyzing the resource entities and their relationships in the global environmental situational awareness map. Based on the multi-objective decision feature set, decision optimization is performed to obtain a comprehensive collaborative scheduling scheme.
[0112] The execution module is used to parse and encapsulate the comprehensive collaborative scheduling scheme into an executable instruction sequence; through a distributed communication architecture, the executable instruction sequence is distributed in parallel to the corresponding decision nodes and control terminals to complete the real-time scheduling of resources and the collaborative release of control instructions.
[0113] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the perception collaborative decision-making system based on multimodal heterogeneous data fusion can be divided into different functional modules to complete all or part of the functions described above.
[0114] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0115] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be executed by computer instructions instructing related hardware. The program can be stored in the aforementioned computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be any of the foregoing embodiments or memory. The aforementioned computer-readable storage medium can also be an external storage device of the authentication device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the authentication device. Further, the aforementioned computer-readable storage medium can include both internal storage units of the authentication device and external storage devices. The aforementioned computer-readable storage medium is used to store the aforementioned computer program and other programs and data required by the authentication device. The aforementioned computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0116] This application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to execute any of the perceptual collaborative decision-making methods based on multimodal heterogeneous data fusion provided in the above embodiments.
[0117] The above are specific embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multi-modal heterogeneous data fusion based perception collaborative decision method, characterized in that, The method comprises: Step 100, collecting multi-modal heterogeneous data, the multi-modal heterogeneous data comprising images, texts, time series data and graph structure data; preprocessing the multi-modal heterogeneous data to obtain standardized data streams; Step 200, analyzing and extracting features of the standardized data streams, and performing alignment and encoding to obtain multi-source feature vectors; performing weighted fusion on the multi-source feature vectors to obtain a preliminary index set; based on the preliminary index set, performing grid processing on a predefined key area, and analyzing the dynamic correlation characteristics of each spatial grid unit to obtain a spatial dynamic weight matrix; performing weighted fusion on the spatial dynamic weight matrix and the preliminary index set, and performing spatial context calibration to obtain an optimized comprehensive state index; based on the optimized comprehensive state index, generating a global environment situation awareness map; Step 300, inputting the global environment situation awareness map into a pre-trained multi-target collaborative decision model; the multi-target collaborative decision model forms a multi-target decision feature set by analyzing resource entities and associated relationships in the global environment situation awareness map; based on the multi-target decision feature set, decision optimization is performed to obtain a comprehensive collaborative scheduling scheme; Step 400, performing instruction analysis and packaging on the comprehensive collaborative scheduling scheme to obtain an executable instruction sequence; the executable instruction sequence is parallelly issued to corresponding decision nodes and control terminals through a distributed communication architecture to complete real-time scheduling of resources and collaborative publishing of control instructions.
2. The method of claim 1, wherein, The step 100 comprises: parallelly collecting the multi-modal heterogeneous data from a plurality of heterogeneous data sources, the plurality of heterogeneous data sources comprising at least one of the following: monitoring images, sensor nodes, mobile units and environmental information data sources; performing format analysis and data cleaning on the multi-modal heterogeneous data to remove invalid and redundant data and obtain structured data; performing outlier detection and missing data repair processing on the structured data to obtain reduced data; performing time series alignment and spatial position matching on the reduced data to establish a unified space-time reference to obtain space-time aligned data; performing normalization processing on the space-time aligned data to eliminate dimension differences to obtain the standardized data streams.
3. The method of claim 1, wherein, The step 200 comprises: performing time dimension analysis on the standardized data streams to divide the continuous standardized data streams into an equal-interval time slice sequence; based on the equal-interval time slice sequence, performing spatial registration processing on the multi-modal heterogeneous data in each time slice to establish spatial reference data under a unified coordinate system, the spatial reference data comprising mobile object trajectory data, fixed sensor data and mobile sensor data; based on the mobile object trajectory data in the spatial reference data, determining the object distribution density in each grid unit to obtain a space-time feature matrix; based on the fixed sensor data, constructing a flow feature vector by statistically analyzing the correlation between event intervals and the number of event occurrences per unit time; using the mobile sensor data, performing outlier rejection and mean value calculation on a dynamic parameter sampling sequence to obtain a regional average dynamic feature; The spatiotemporal feature matrix, the traffic feature vector and the area average dynamic feature are aligned in feature dimension, and a multi-source feature vector is obtained through vectorization splicing.
4. The method of claim 3, wherein, The step 200 further includes: Based on the multi-source feature vector, an attention weight matrix of each feature vector is determined through a preset fully connected layer and a normalized exponential softmax function; The attention weight matrix and the multi-source feature vector are weighted and summed to obtain an attention weighted fusion feature; The attention weighted fusion feature is processed by feature dimension reduction and standardization to obtain a preliminary index set; Based on the preliminary index set, a regular grid topology structure covering the area is constructed in combination with a pre-defined key area boundary; Based on the regular grid topology structure, the preliminary index set is mapped to the spatial grid cells to determine the statistical distribution characteristics of the environmental state indicators in each grid cell; Based on the statistical distribution characteristics of the environmental state indicators in each grid cell, the dynamic correlation strength between adjacent grid cells is calculated through spatial autocorrelation analysis; The statistical distribution characteristics of the environmental state indicators in each grid cell and the dynamic correlation strength between adjacent grid cells are integrated to obtain the spatial dynamic weight matrix through weighted fusion calculation.
5. The method of claim 4, wherein, The step 200 further includes: The spatial dynamic weight matrix and the preliminary index set are multiplied element by element to realize feature weighted fusion in the spatial domain and obtain a weighted feature mapping; Based on the weighted feature mapping, the feature propagation of adjacent network nodes is spatially context calibrated to obtain context-aware features; the context-aware features are processed by feature compression and standardization to obtain the optimized comprehensive state indicators; Based on the optimized comprehensive state indicators, feature upsampling and spatial resolution enhancement are performed to obtain a high-resolution feature mapping; Based on the high-resolution feature mapping, a continuous spatial distribution feature field is obtained through a spatial interpolation algorithm; Based on the continuous spatial distribution feature field, the feature field values are mapped to corresponding color space values to obtain a raster image; based on the raster image, the gradient distribution of the feature field is calculated to obtain contour line vector data; the raster image and the contour line vector data are subjected to layer superposition and semantic labeling processing to obtain the global environmental situation awareness map.
6. The method of claim 1, wherein, The step 300 includes: The global environmental situation awareness map is parsed to extract key state indicators representing task processing efficiency, resource utilization balance and service demand satisfaction, and a multi-objective decision feature set is constructed; Based on the multi-objective decision feature set, a global multi-objective optimization function that simultaneously optimizes operation efficiency, resource cost and service quality is constructed, and operation constraints of each functional unit are set. The global multi-objective optimization function is decomposed into a local value function associated with each agent; each agent generates a local strategy based on its local value function and the environment state; all local strategies are integrated by a centralized coordinator, which evaluates its global coordination utility using an attention mechanism, and obtains a non-dominated solution set as an initial coordination strategy set by iterative search within the constraint solution space; The initial coordination strategy set is evaluated and analyzed in multiple dimensions, and the comprehensive utility value of each candidate strategy is calculated; the candidate strategies are sorted according to the comprehensive utility value, and a comprehensive utility ranking is obtained; Based on the comprehensive utility ranking, a dominant strategy is selected from the initial coordination strategy set to obtain the comprehensive coordination scheduling scheme.
7. The method of claim 6, wherein, The step 400 includes: The semantic analysis and instruction conversion of the comprehensive coordination scheduling scheme are performed to generate the executable instruction sequence including the device control instruction set, the resource scheduling instruction set and the service publishing instruction set; Based on the preset communication protocol, the executable instruction sequence is packaged into a standardized instruction data packet recognizable by each control terminal; The standardized instruction data packet is distributed in parallel to the corresponding decision node and control terminal through the distributed message middleware; Receive the state feedback data returned by each control terminal, and dynamically adjust the distribution strategy of subsequent instructions based on the state feedback data through the feedback control mechanism to complete the real-time scheduling of global resources and the coordinated publishing of control instructions.
8. A multi-modal heterogeneous data fusion based perception collaborative decision system, the system implements the method of any one of claims 1 to 7, characterized in that, It includes: A multi-source acquisition module is used to acquire multi-modal heterogeneous data, and the multi-modal heterogeneous data includes images, texts, time series data and graph structure data; The multi-modal heterogeneous data is preprocessed to obtain a standardized data stream; A fusion analysis module is used to analyze and extract features of the standardized data stream, align and encode the standardized data stream, obtain a multi-source feature vector, weight fuse the multi-source feature vector, obtain a preliminary index set, grid process a pre-defined key area based on the preliminary index set, analyze dynamic correlation characteristics of each spatial grid unit, obtain a spatial dynamic weight matrix, weight fuse the spatial dynamic weight matrix and the preliminary index set, and calibrate through a spatial context to obtain an optimized comprehensive state index, and generate a global environment situation awareness map based on the optimized comprehensive state index; A collaborative decision module is used to input the global environment situation awareness map into a pre-trained multi-objective collaborative decision model; the multi-objective collaborative decision model forms a multi-objective decision feature set by analyzing resource entities and correlation relationships in the global environment situation awareness map; Based on the multi-objective decision feature set, decision optimization is performed to obtain a comprehensive coordination scheduling scheme; An execution module is used to perform instruction analysis and packaging on the comprehensive coordination scheduling scheme to obtain an executable instruction sequence; the executable instruction sequence is distributed in parallel to the corresponding decision node and control terminal through a distributed communication architecture to complete real-time scheduling of resources and coordinated publishing of control instructions.
9. A computing device, comprising: It includes: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program, which when executed by a processor, implements the method of any one of claims 1-7.
Citation Information
Patent Citations
Unmanned intelligent multi-mode information fusion and target perception system and operation method
CN116310689A
Intelligent agent autonomous decision control method based on multi-modal data fusion
CN120469238A