A multi-modal data transmission optimization method for smart city
By improving the CorrNet network and using differentiated coding scheduling technology, the problems of dynamic adjustment and real-time optimization in multimodal data transmission in smart cities have been solved, achieving efficient and reliable data transmission and improving network resource utilization and data processing accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 仁和善德(天津)城市建设有限公司
- Filing Date
- 2025-12-24
- Publication Date
- 2026-06-19
AI Technical Summary
Existing technologies struggle to dynamically adjust encoding methods and resource allocation based on the city's operational status during multimodal data transmission in smart cities. This leads to increased latency of critical data, exacerbated link congestion, accumulated packet loss, and decreased transmission efficiency. Furthermore, the lack of real-time path optimization capabilities and feedback mechanisms results in low network resource utilization and insufficient reliability.
An improved CorrNet network is used for multimodal feature extraction. Combined with differentiated coding and intelligent scheduling techniques, coding control instructions and scheduling control instructions are generated through region partitioning and redundant channel enhancement to achieve differentiated data compression and coding processing. Real-time performance evaluation is also performed during transmission.
It achieves high accuracy, controllable latency, and high resource utilization in multimodal data transmission, significantly reduces the probability of packet loss in highly dynamic network environments, improves the reliability and scheduling efficiency of data processing, and enhances the efficiency of network resource utilization.
Smart Images

Figure CN121728113B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart city data communication technology, and in particular to a method for optimizing multimodal data transmission for smart cities. Background Technology
[0002] In recent years, the construction of smart cities has advanced rapidly. Cities have deployed a large number of multi-source sensing terminals, including cameras, road and traffic infrastructure sensors, and vehicle-to-the-road communication equipment. The resulting video, sensor data, and collaborative sensing data exhibit multimodal, continuous, and high-density characteristics. Existing technologies typically employ fixed encoding parameters, uniform compression strategies, and static transmission paths, processing multimodal data with the same priority. This makes it difficult to dynamically adjust encoding methods and resource allocation based on the city's operational status. Furthermore, due to the complexity of urban network topologies and the significant changes in link status over time, existing solutions generally rely on static or heuristic scheduling, lacking the ability to optimize real-time paths across links and multiple nodes.
[0003] In complex traffic environments and highly dynamic network conditions, these methods can easily lead to increased latency of critical data, exacerbated link congestion, accumulated packet loss, and decreased cross-node transmission efficiency. Furthermore, traditional feature extraction models struggle to simultaneously account for the impact of regional differences and link packet loss on data quality, failing to output high-quality features suitable for generating encoding and scheduling strategies. Moreover, existing technologies lack feedback mechanisms for the data transmission process, making it impossible to evaluate encoding and scheduling strategies based on actual performance, resulting in low network resource utilization and insufficient reliability.
[0004] Therefore, how to provide a multimodal data transmission optimization method for smart cities is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a multimodal data transmission optimization method for smart cities. This invention integrates multimodal feature extraction, differential coding, and intelligent scheduling technologies to achieve efficient and reliable transmission of urban data, with advantages such as high accuracy, controllable latency, and high resource utilization.
[0006] A method for optimizing multimodal data transmission for smart cities according to an embodiment of the present invention includes the following steps:
[0007] Collect raw multimodal data generated in the city, preprocess it, and generate a standardized multimodal dataset;
[0008] An improved CorrNet network is used to perform joint feature extraction on a standardized multimodal dataset, outputting multimodal features and corresponding importance labels according to preset scene rules;
[0009] Operational data is collected from edge nodes, roadside units, and base stations to organize and analyze the urban network and generate network status information.
[0010] Multimodal features, importance labels, and network state information are correlated and analyzed, and corresponding encoding methods, compression ratios, and redundancy levels are determined for various types of data to generate encoding control instructions.
[0011] Based on the encoding control instructions, the multimodal data is subjected to differentiated compression and encoding processing to generate a priority-ordered queue of data to be transmitted, with encoded data units as elements.
[0012] Based on network status information and historical scheduling results, path calculation and resource allocation analysis are performed on the data to be transmitted. Scheduling control instructions are generated through tabu search and implemented in urban network nodes.
[0013] During data transmission, feedback indicators are continuously collected, and the effectiveness of the current encoding and scheduling strategies is evaluated in conjunction with network status information. Transmission performance evaluation results are generated and uploaded to the system monitoring terminal.
[0014] Optionally, the multimodal raw data includes video data from cameras, environmental and traffic sensor data from roads and traffic infrastructure, and collaborative sensing data from vehicle-to-road communication devices. The preprocessing includes performing time alignment, format parsing, structured processing, noise removal, and outlier removal on the video data based on a unified time reference; performing time synchronization, format parsing, noise suppression, and outlier filtering on the environmental and traffic sensor data; performing time calibration, field parsing, and structured processing on the collaborative sensing data; and performing slicing and resampling on all multimodal data according to time windows.
[0015] Optionally, the output of the multimodal features and importance labels specifically includes:
[0016] Each data segment in the standardized multimodal dataset is indexed according to its collection location and the corresponding urban network area, and each data segment is associated with a preset regional identifier.
[0017] The urban network consists of the connection relationships of edge nodes, roadside units and base stations deployed within the city. Its topology is obtained by collecting network configuration data, link connection information and node identification information of each node, and performing regularization parsing processing on each connection relationship.
[0018] The multimodal data segments that have completed the region indexing process are input into the improved CorrNet network. Based on the region identifier to which each data segment belongs, region-related feature representations are generated in the corresponding region-shared hidden layer blocks.
[0019] The improved CorrNet network replaces the single shared hidden layer of the original CorrNet network with a multi-region shared hidden layer structure based on the city topology, introduces regional blocks corresponding to city network nodes and adds topological regularization constraints between regions, and adds supplementary channels in the shared hidden layer as redundant feature processing paths.
[0020] In the regional shared hidden layer block, the regional related feature representation is processed by regional association. Based on the topological connectivity of adjacent regions in the urban network, the regional related features belonging to adjacent regions are continuously weighted according to the preset regional association coefficient to obtain the regional continuous feature set.
[0021] Redundant feature extraction is performed on multimodal data segments by supplementary channels in the improved CorrNet network. Redundant features corresponding to each data segment are generated based on the hidden layer random deactivation rule and redundant channel reconstruction constraints, thus generating a set of redundant features.
[0022] The set of regional continuity features and the set of redundant features are fused together to form a set of shared spatial features that includes regional continuity attributes and redundancy recovery attributes, based on the feature fusion rules.
[0023] A calibration and reconstruction constraint is applied to the shared spatial feature set. A reconstruction calibration value is generated by weighted summation of regional reconstruction error and redundant reconstruction error. The shared spatial feature set is then calibrated for consistency based on the reconstruction calibration value to generate multimodal features.
[0024] The calibrated shared spatial feature representation is processed to generate importance. Based on the preset scene rules, the regional attributes, semantic information and task sensitivity of each data segment are comprehensively calculated to obtain the corresponding importance label and output it together with the multimodal features.
[0025] Optionally, the generation of the network status information specifically includes:
[0026] Operational data is collected from edge nodes, roadside units, and base stations, and timestamps and node identifiers are recorded during collection to generate an operational dataset. The operational data includes network bandwidth, latency, packet loss rate, and node load.
[0027] The data in the running dataset is processed by discarding abnormal time points to form a running data sequence on a unified time axis. Then, normalization processing is performed on network bandwidth, latency, packet loss rate and node load respectively, and each running data is mapped to the same numerical range to obtain a standardized running data sequence.
[0028] The standardized operational data sequence is grouped according to the links and nodes in the urban network topology, and network bandwidth, latency, packet loss rate and node load are combined into operational data vectors in a fixed order.
[0029] For each running data vector, a comprehensive running score is calculated by multiplying network bandwidth, latency, packet loss rate and node load by their respective weight coefficients and summing them to generate a comprehensive running score value for the corresponding link and node.
[0030] The comprehensive operational scores of each link and node are organized according to the time dimension and topological location to generate network status information with nodes and links as indexes and comprehensive operational scores and corresponding operational data as records.
[0031] Optionally, the generation of the encoded control instructions specifically includes:
[0032] Multimodal features, importance labels, and network state information are associated according to the correspondence of data segments to generate data description records for each data segment;
[0033] Read the importance level identifier and importance score value corresponding to each data segment in the data description record, divide the data segments into priority level categories with different numbers according to the importance level identifier, use the importance score value as the priority quantification score used for coding priority calculation, and generate priority parsing results;
[0034] The network status information in the data description record is parsed, the comprehensive operation score value corresponding to the data segment transmission path and the operation data on the corresponding path are read, comparison processing is performed according to the preset threshold set, the network status level corresponding to the data segment is determined, and the network status level is mapped to the network status level quantization value to generate the network status parsing result.
[0035] Based on the priority parsing results and the network state parsing results, a joint coding strategy is calculated to determine the coding method, compression ratio and redundancy level of each data segment;
[0036] The encoding method, compression ratio, and redundancy level are combined with the corresponding data segment identifier to generate encoding control instructions.
[0037] Optionally, the generation of the data queue to be transmitted specifically includes:
[0038] Based on the encoding control instructions, the data segments in the multimodal data are associated with the corresponding encoding methods, compression ratios, redundancy levels, and priority parsing results to generate a set of encodings to be processed.
[0039] Select data segments sequentially from the set to be encoded, read the corresponding encoding method, compression ratio, redundancy level and priority quantization score, and generate corresponding encoding strategy records based on preset parameters;
[0040] Based on the encoding method, target compression ratio and target redundancy level recorded in the encoding strategy record, differential compression and encoding processing is performed on the corresponding data segments. Encoded data units corresponding to the data segment identifier are generated through data fragmentation, transform coding, entropy coding and redundancy information appending, and priority level category and priority quantization score are written into the encoded data units.
[0041] All encoded data units are sorted according to priority level category, priority quantization score, and time order of data segments to generate a priority-sorted queue of data to be transmitted, with encoded data units as elements.
[0042] Optionally, the generation of the scheduling control instructions specifically includes:
[0043] Based on network status information and the queue of data to be transmitted, the city network topology is analyzed to resolve links and nodes, generating a scheduling input dataset;
[0044] Select the start node and end node corresponding to the encoded data unit in the data queue to be transmitted from the scheduling input dataset. Construct several initial transmission paths according to the urban network topology. Allocate bandwidth usage ratio on each initial transmission path according to the link available bandwidth and the priority quantization score of the encoded data unit. Generate an initial scheduling scheme including path sequence and bandwidth usage ratio allocation results.
[0045] Based on network status information and historical scheduling results, a path cost evaluation function is constructed for candidate transmission paths to calculate the path cost and generate a global scheduling evaluation value.
[0046] The initial scheduling scheme is iteratively optimized using the tabu search method to obtain the target scheduling scheme.
[0047] Based on the target scheduling plan, scheduling control instructions are generated and sent to the corresponding city network nodes for scheduling implementation.
[0048] Optionally, the generation of the transmission performance evaluation results specifically includes:
[0049] Feedback indicators are collected from urban network nodes during data transmission and associated with data segments to generate a set of feedback indicator records.
[0050] The feedback indicator record set is collected and structured in conjunction with network status information to generate a set of feedback evaluation inputs;
[0051] Perform transmission performance calculations on the feedback evaluation input set to generate path-level transmission performance scores;
[0052] All path-level transmission performance scores are aggregated according to the data segment to which the encoded data unit belongs, priority level category, and transmission time window. Transmission performance evaluation results oriented towards data segments are generated and uploaded to the system monitoring terminal.
[0053] The beneficial effects of this invention are:
[0054] This invention constructs a multimodal data transmission optimization system for smart city environments, achieving end-to-end linked processing from feature extraction and encoding control to scheduling optimization, resulting in a significant comprehensive improvement compared to existing technologies. Firstly, the improved CorrNet network proposed in this invention, based on regional partitioning and redundant channel enhancement, can extract shared spatial features with regional continuity, semantic consistency, and redundancy recovery capabilities in multi-regional urban network topologies. Furthermore, through an importance label generation mechanism, it comprehensively characterizes the task sensitivity, semantic attributes, and urban regional dynamics of data fragments. This allows encoding and scheduling strategies to be dynamically adjusted according to scenario requirements and regional attributes, rather than being based on fixed parameters or a single priority, thereby improving the accuracy and reliability of data processing from the source.
[0055] Secondly, the coding strategy generation mechanism proposed in this invention introduces the joint calculation of priority parsing results and network state parsing results, realizing differentiated allocation of coding methods, compression ratios, and redundancy levels. Through refined coding path design of fragmentation, transform coding, entropy coding, and redundant information appending, this invention can flexibly adjust data compression and redundancy intensity under different bandwidth, latency, and packet loss conditions, thereby ensuring that high-priority data has higher fidelity and transmission reachability. Compared with uniform compression or simple forced redundancy in existing technologies, this invention can significantly reduce the probability of packet loss of critical data in highly dynamic network environments, shorten signaling response time, and improve overall link utilization. In addition, the structured design of the coded data unit enables the scheduling optimization stage to accurately identify priorities, network states, and data attributes, thereby achieving cross-module information collaboration.
[0056] Finally, this invention introduces a cross-node scheduling optimization mechanism based on tabu search, combining real-time network status information and historical scheduling results to achieve dynamic optimization of path calculation and resource allocation for urban networks. This mechanism can continuously correct the transmission path and bandwidth usage ratio under complex multi-node, multi-link topologies, significantly improving the scheduling efficiency of multimodal data in urban networks and avoiding local optima and resource conflicts that occur in traditional path planning. Simultaneously, this invention continuously collects feedback indicators during data transmission and generates transmission performance evaluation results for data segments based on network status information, enabling the system to achieve real-time monitoring of the execution effects of coding and scheduling strategies, improving the overall system's adaptability and network resource utilization efficiency. In summary, this invention achieves substantial innovations in multimodal feature extraction, coding optimization, and intelligent scheduling, effectively improving the reliability, timeliness, and resource utilization efficiency of multimodal data transmission in smart city scenarios, and has significant engineering application value and promotional significance. Attached Figure Description
[0057] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0058] Figure 1 This is a flowchart of a multimodal data transmission optimization method for smart cities proposed in this invention;
[0059] Figure 2 This is a schematic diagram of the improved CorrNet network feature processing path for a multimodal data transmission optimization method for smart cities proposed in this invention. Detailed Implementation
[0060] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0061] refer to Figure 1-2 A method for optimizing multimodal data transmission for smart cities includes the following steps:
[0062] Collect raw multimodal data generated in the city, preprocess it, and generate a standardized multimodal dataset;
[0063] An improved CorrNet network is used to perform joint feature extraction on a standardized multimodal dataset, outputting multimodal features and corresponding importance labels according to preset scene rules;
[0064] Operational data is collected from edge nodes, roadside units, and base stations to organize and analyze the urban network and generate network status information.
[0065] Multimodal features, importance labels, and network state information are correlated and analyzed, and corresponding encoding methods, compression ratios, and redundancy levels are determined for various types of data to generate encoding control instructions.
[0066] Based on the encoding control instructions, the multimodal data is subjected to differentiated compression and encoding processing to generate a priority-ordered queue of data to be transmitted, with encoded data units as elements.
[0067] Based on network status information and historical scheduling results, path calculation and resource allocation analysis are performed on the data to be transmitted. Scheduling control instructions are generated through tabu search and implemented in urban network nodes.
[0068] During data transmission, feedback indicators are continuously collected, and the effectiveness of the current encoding and scheduling strategies is evaluated in conjunction with network status information. Transmission performance evaluation results are generated and uploaded to the system monitoring terminal.
[0069] In this embodiment, the multimodal raw data includes video data from cameras, environmental and traffic sensor data from roads and traffic infrastructure, and collaborative sensing data from vehicle-to-road communication devices. The preprocessing includes performing time alignment, format parsing, structured processing, noise removal, and outlier removal on the video data based on a unified time reference; performing time synchronization, format parsing, noise suppression, and outlier filtering on the environmental and traffic sensor data; performing time calibration, field parsing, and structured processing on the collaborative sensing data; and performing slicing and resampling on all multimodal data according to time windows.
[0070] In this embodiment, the output of the multimodal features and importance labels specifically includes:
[0071] Each data segment in the standardized multimodal dataset is indexed according to its collection location and the corresponding urban network area, and each data segment is associated with a preset regional identifier.
[0072] The urban network consists of the connection relationships of edge nodes, roadside units and base stations deployed within the city. Its topology is obtained by collecting network configuration data, link connection information and node identification information of each node, and performing regularization parsing processing on each connection relationship.
[0073] The region identifier includes region number information for indicating the urban network region to which the data fragment belongs, region topology attribute information for characterizing the region’s positional relationship in the urban network topology, region boundary information for indicating the spatial extent of the region, and region attribute information for describing the infrastructure configuration within the region.
[0074] The multimodal data segments that have completed the region indexing process are input into the improved CorrNet network. Based on the region identifier to which each data segment belongs, region-related feature representations are generated in the corresponding region-shared hidden layer blocks.
[0075] The improved CorrNet network replaces the single shared hidden layer of the original CorrNet network with a multi-region shared hidden layer structure based on the city topology, introduces regional blocks corresponding to city network nodes and adds topological regularization constraints between regions, and adds supplementary channels in the shared hidden layer as redundant feature processing paths.
[0076] The formation of the regional shared hidden layer block includes: parsing the single shared hidden layer in the original related network according to the preset urban network regions in the urban topology; mapping the region identifier information of each region in the urban network to the hidden layer unit index of the original shared hidden layer according to the region partitioning rules; and grouping the hidden layer units of the original shared hidden layer according to the region identifier based on the mapping result, so that each group forms a regional shared hidden layer block; further, performing sequential arrangement processing on multiple regional shared hidden layer blocks according to the regional adjacency relationship of the urban topology, so that the regional shared hidden layer blocks form a multi-region shared hidden layer structure in the hidden layer structure corresponding to the urban topology, which replaces the single shared hidden layer of the original related network.
[0077] The generation of the region-related feature representation includes: determining the corresponding region-shared hidden layer block for multimodal data segments carrying region identifiers according to the region identifiers; sending the input data of each data segment into the encoding path of the corresponding region-shared hidden layer block for feature extraction; forming an initial region feature vector by performing linear mapping processing, nonlinear transformation processing and multidimensional feature compression processing on the input data; and further performing feature aggregation processing on the initial region feature vector according to the feature aggregation rules of the region-shared hidden layer block, so that data segments corresponding to the same region identifier form region-related feature representations in the shared hidden layer space.
[0078] In the regional shared hidden layer block, the regional related feature representation is processed by regional association. Based on the topological connectivity of adjacent regions in the urban network, the regional related features belonging to adjacent regions are continuously weighted according to the preset regional association coefficient to obtain the regional continuous feature set.
[0079] Redundant feature extraction is performed on multimodal data segments by supplementary channels in the improved CorrNet network. Redundant features corresponding to each data segment are generated based on the hidden layer random deactivation rule and redundant channel reconstruction constraints, thus generating a set of redundant features.
[0080] The generation of the redundant feature set includes: converting multimodal data fragments processed by region indexing into input representations of the supplementary channel through the input mapping unit of the supplementary channel, and sending the input representations into the hidden layer processing unit of the supplementary channel; in the hidden layer processing unit, the hidden layer units are zeroed according to the hidden layer random deactivation rule and the execution unit is set to zero according to the set deactivation probability to form a redundant input sequence including deactivation modes; inputting the redundant input sequence into the redundant reconstruction unit of the supplementary channel, performing a reconstruction operation on the redundant input sequence, obtaining the reconstruction error value of the non-deactivation unit by calculating the corresponding difference between the reconstruction operation output and the original input, and weighting all reconstruction error values according to the preset redundant error weight to generate a redundant channel reconstruction constraint value for characterizing the redundancy recovery capability of the supplementary channel; applying the redundant channel reconstruction constraint value to the redundant input sequence, performing error calibration processing on the sequence to form a redundant feature vector corresponding to each data fragment; further, according to the one-to-one correspondence between the data fragments and the redundant feature vectors, all redundant feature vectors are aggregated to generate a redundant feature set for characterizing the redundant expression capability of multimodal data in the supplementary channel;
[0081] The set of regional continuity features and the set of redundant features are fused together to form a set of shared spatial features that includes regional continuity attributes and redundancy recovery attributes, based on the feature fusion rules.
[0082] The generation of the shared spatial feature set includes: pairing each feature vector in the regional continuity feature set and the redundant feature set according to the correspondence of data segments; using the regional continuity feature vector and the redundant feature vector belonging to the same data segment as fusion input; performing feature splicing, feature weighting, and feature alignment processing on each pair of fusion inputs according to preset feature fusion rules; generating the fused shared spatial feature vector by jointly mapping the regional continuity attribute and the redundancy recovery attribute; further organizing and collecting all shared spatial feature vectors according to the order of data segments to generate a shared spatial feature set that simultaneously represents the regional continuity attribute and the redundancy recovery attribute.
[0083] A calibration and reconstruction constraint is applied to the shared spatial feature set. A reconstruction calibration value is generated by weighted summation of regional reconstruction error and redundant reconstruction error. The shared spatial feature set is then calibrated for consistency based on the reconstruction calibration value to generate multimodal features.
[0084] The generation of multimodal features includes: performing reconstruction processing on each shared spatial feature vector in the shared spatial feature set; obtaining a regional reconstruction error value by calculating the corresponding difference between the reconstruction output of the shared spatial feature vector in the region-related feature dimension and the region continuity feature set; and obtaining a redundant reconstruction error value by calculating the corresponding difference between the reconstruction output of the shared spatial feature vector in the redundant feature dimension and the redundant feature set; generating a calibration reconstruction constraint value for global calibration by weighting and summing the regional reconstruction error value and the redundant reconstruction error value according to a preset error weighting rule; applying the calibration reconstruction constraint value to the shared spatial feature set and performing consistency calibration processing on each shared spatial feature vector to achieve feature consistency between the region continuity attribute and the redundant recovery attribute in a unified feature space; and further, collecting the consistent calibrated shared spatial feature vectors according to the order of data segments to generate multimodal features for characterizing the region attributes, semantic attributes, and redundant recovery attributes of multimodal data.
[0085] The importance generation process is performed on the calibrated shared spatial feature representation. Based on the preset scene rules, the regional attributes, semantic information and task sensitivity of each data segment are comprehensively calculated to obtain the corresponding importance label and output it together with the multimodal features.
[0086] The scenario rules include: region attribute rules, semantic information rules, task sensitivity rules, importance weighted calculation rules, and importance division rules;
[0087] The regional attributes include regional identification information, regional topological association information, regional feature density information, and regional dynamic status information.
[0088] The semantic information includes object category semantics, behavioral semantics, scene semantic tags, and semantic association information;
[0089] The task sensitivity includes task relevance coefficient, time sensitivity index, risk sensitivity index, and priority reference information;
[0090] The importance labels include importance level identifiers, importance score values, decision priority parameters, and label association information indicating the correspondence between importance labels and multimodal features;
[0091] The output of the importance label includes: performing feature parsing processing on the calibrated shared space feature representation according to the correspondence of data segments; extracting regional attributes, semantic information and task sensitivity information from each shared space feature representation according to preset scene rules; performing weighted calculation on regional attributes, semantic information and task sensitivity information to generate an importance calculation value corresponding to each data segment; and performing interval mapping processing on the importance calculation value according to preset importance division rules to determine the final importance label of each data segment.
[0092] In this embodiment, the generation of the network status information specifically includes:
[0093] Operational data is collected from edge nodes, roadside units, and base stations, and timestamps and node identifiers are recorded during collection to generate an operational dataset. The operational data includes network bandwidth, latency, packet loss rate, and node load.
[0094] The data in the running dataset is processed by discarding abnormal time points to form a running data sequence on a unified time axis. Then, normalization processing is performed on network bandwidth, latency, packet loss rate and node load respectively, and each running data is mapped to the same numerical range to obtain a standardized running data sequence.
[0095] The standardized operational data sequence is grouped according to the links and nodes in the urban network topology, and network bandwidth, latency, packet loss rate and node load are combined into operational data vectors in a fixed order.
[0096] For each running data vector, a comprehensive running score is calculated by multiplying network bandwidth, latency, packet loss rate and node load by their respective weight coefficients and summing them to generate a comprehensive running score value for the corresponding link and node. The comprehensive running score value is associated with the time scale and node identifier.
[0097] The comprehensive operational scores of each link and node are organized according to the time dimension and topological location to generate network status information with nodes and links as indexes and comprehensive operational scores and corresponding operational data as records.
[0098] In this embodiment, the generation of the encoding control instructions specifically includes:
[0099] Multimodal features, importance labels, and network status information are associated according to the correspondence of data segments to generate a data description record for each data segment. The data description record includes the multimodal features, importance labels, and network status information of the data segment in the urban network.
[0100] Read the importance level identifier and importance score value corresponding to each data segment in the data description record, divide the data segments into priority level categories with different numbers according to the importance level identifier, use the importance score value as the priority quantification score used for coding priority calculation, and generate priority parsing results;
[0101] The network status information in the data description record is parsed, the comprehensive operation score value corresponding to the data segment transmission path and the operation data on the corresponding path are read, comparison processing is performed according to the preset threshold set, the network status level corresponding to the data segment is determined, and the network status level is mapped to the network status level quantization value to generate the network status parsing result.
[0102] Based on the priority parsing results and network status parsing results, a joint coding strategy calculation is performed to determine the coding method, compression ratio and redundancy level of each data segment. The joint coding strategy calculation includes: multiplying the priority quantization score and the network status level quantization value by the corresponding weight coefficients and summing them to generate a coding strategy score value, and retrieving the corresponding coding method, compression ratio and redundancy level in a preset strategy mapping table according to the coding strategy score value.
[0103] The encoding method, compression ratio, and redundancy level are combined with the corresponding data segment identifier to generate an encoding control instruction, which includes a data segment identifier field, an encoding method field, a compression ratio field, and a redundancy level field.
[0104] In this embodiment, the generation of the data queue to be transmitted specifically includes:
[0105] Based on the encoding control instructions, the data segments in the multimodal data are associated with the corresponding encoding methods, compression ratios, redundancy levels, and priority parsing results to generate a set of encodings to be processed.
[0106] Data segments are selected sequentially from the set to be encoded and processed. The corresponding encoding method, compression ratio, redundancy level, and priority quantization score are read. The corresponding encoding strategy record is generated by weighting according to preset parameters. The generation of the encoding strategy record includes weighting the compression ratio according to preset compression adjustment parameters to generate a target compression ratio, and weighting the redundancy level according to preset redundancy adjustment parameters to generate a target redundancy level. The encoding method, target compression ratio, and target redundancy level are combined according to the data segment identifier.
[0107] Based on the encoding method, target compression ratio and target redundancy level recorded in the encoding strategy record, differential compression and encoding processing is performed on the corresponding data segments. Encoded data units corresponding to the data segment identifier are generated through data fragmentation, transform coding, entropy coding and redundancy information appending, and priority level category and priority quantization score are written into the encoded data units.
[0108] The generation of the encoded data unit includes: performing data segmentation processing on the data segments determined according to the encoding strategy and according to the data segment identifier; sequentially dividing the original data sequence of the data segments according to the preset segment length and performing time-aligned segmentation according to the preset time window to obtain the segmented sequence for encoding processing; performing transform coding processing on each segmented sequence according to the encoding path corresponding to the encoding method; generating transform coefficient sequences by performing block-level decomposition, coefficient transformation, and multi-dimensional feature compression on the segmented sequences; performing entropy coding processing on the transform coefficient sequences; generating a basic encoded bitstream by performing statistical modeling on the symbol information of the transform coefficient sequences, probability estimation on the amplitude information of the transform coefficient sequences, and codeword generation on both; and processing the basic encoded codestream according to the target compression ratio. The stream undergoes compression ratio adjustment processing, which involves truncation, codeword reallocation, and quantization level adjustment of the base encoded bitstream to generate a compressed bitstream. Based on the target redundancy level, redundant information is appended to the compressed bitstream by generating a cyclic parity sequence, a block parity sequence, and a redundancy parity segment of a specified length. These sequences are then appended to the compressed bitstream in a predetermined order to form a redundant appended bitstream. Priority level category and priority quantization score corresponding to the data segment identifier are written into the redundant appended bitstream. Encoded data units containing encoded result data, redundant appended data, and priority record information are constructed according to the data segment identifiers. All encoded data units are then processed and aggregated.
[0109] All encoded data units are sorted according to priority level category, priority quantization score, and time order of data segments to generate a priority-sorted queue of data to be transmitted, with encoded data units as elements.
[0110] In this embodiment, the generation of the scheduling control command specifically includes:
[0111] Based on network status information and the queue of data to be transmitted, the city network topology is analyzed to resolve links and nodes, generating a scheduling input dataset;
[0112] The generation of the scheduling input dataset includes: parsing the comprehensive operation score, network bandwidth, latency, packet loss rate, and node load recorded in the network status information according to links and nodes; reading the link identifier field and node identifier field of each record item, filtering the time scale field of the record item, and splitting the indicator field of the record item item by item, assigning the indicator value corresponding to the link identifier to the link record, and assigning the indicator value corresponding to the node identifier to the node record; further combining the priority level category and priority quantization score of the encoded data units in the data queue to be transmitted, performing availability judgment processing on the link records to form an available link set, performing availability judgment processing on the node records to form an available node set, and performing path combination processing based on the link connection relationship in the available link set to form a candidate transmission path set; and writing the available link set, available node set, candidate transmission path set, and available bandwidth information of the links into the scheduling input dataset in a preset data structure order.
[0113] Select the start node and end node corresponding to the encoded data unit in the data queue to be transmitted from the scheduling input dataset. Construct several initial transmission paths according to the urban network topology. Allocate bandwidth usage ratio on each initial transmission path according to the link available bandwidth and the priority quantization score of the encoded data unit. Generate an initial scheduling scheme including path sequence and bandwidth usage ratio allocation results.
[0114] Based on network status information and historical scheduling results, a path cost evaluation function is constructed for candidate transmission paths to calculate the path cost and generate a global scheduling evaluation value. The path cost evaluation function is obtained by multiplying the inverse of the available bandwidth of the link, the link delay value, the link packet loss rate value, and the link node load value on the candidate transmission path by the corresponding bandwidth cost weight, delay cost weight, packet loss cost weight, and load cost weight, respectively, and then summing them up. The path cost values of each candidate transmission path are then aggregated according to the preset path cost aggregation rules to generate a global scheduling evaluation value.
[0115] The initial scheduling scheme is iteratively optimized using the tabu search method to obtain the target scheduling scheme.
[0116] The target scheduling scheme is obtained by adjusting the transmission paths of some coded data units according to the current scheduling scheme in each iteration, redistributing the bandwidth usage ratio on the relevant paths, and reordering the scheduling of coded data units among different paths to generate a neighborhood scheduling scheme set. For each candidate scheduling scheme in the neighborhood scheduling scheme set, the corresponding global scheduling evaluation value is calculated according to the path cost evaluation function. The characteristics of recently visited scheduling schemes are recorded according to a preset taboo list to prevent the scheduling schemes recorded in the taboo list from being selected as the current solution again. Under the premise of satisfying the taboo criterion and the boundary criterion, the neighborhood scheduling scheme with a global scheduling evaluation value better than the current scheduling scheme is selected as the new current scheduling scheme. During the iteration process, the scheduling scheme with the best global scheduling evaluation value is retained as the target scheduling scheme according to the preset iteration number threshold and the scheduling evaluation value convergence condition.
[0117] Based on the target scheduling plan, scheduling control instructions are generated and sent to the corresponding city network nodes for scheduling implementation.
[0118] The generation of the scheduling control instructions includes: reading the cross-node transmission path and bandwidth usage ratio of each coded data unit from the target scheduling scheme, writing the path forwarding relationship, link bandwidth occupancy ratio and queue forwarding priority into the scheduling control instruction field, organizing them into a scheduling control instruction set according to nodes, and executing the scheduling implementation in the city network nodes.
[0119] In this embodiment, the generation of the transmission performance evaluation result specifically includes:
[0120] Feedback indicators are collected from urban network nodes during data transmission and associated with data fragments to generate a set of feedback indicator records. The feedback indicators include latency, packet loss rate, retransmission rate, and content decoding effect.
[0121] The feedback indicator record set is collected and structured in conjunction with network status information to generate a feedback evaluation input set. The generation of the feedback evaluation input set includes: collecting the feedback indicator records according to time windows and transmission paths; performing path-level organization on the latency, packet loss rate, retransmission rate and content decoding effect on the same path within the same time window; and combining the organized indicators with the comprehensive operation score value and link operation data in the network status information according to the path dimension.
[0122] The transmission performance is calculated on the feedback evaluation input set to generate a path-level transmission performance score. The transmission performance calculation includes multiplying the latency index, packet loss rate index, retransmission rate index and content decoding effect quality score of each path by the corresponding evaluation weight coefficient, and then weighting and summing them with the comprehensive operation score in the network status information according to the comprehensive operation evaluation weight coefficient to obtain the path-level transmission performance score.
[0123] All path-level transmission performance scores are aggregated according to the data segment, priority level category, and transmission time window to which the encoded data unit belongs, generating a data segment-oriented transmission performance evaluation result and uploading it to the system monitoring terminal. The generation of the transmission performance evaluation result includes: aggregating the transmission performance scores of each data segment on different paths according to a preset weighting rule to form a data segment-level transmission performance evaluation value, and outputting the evaluation value together with the corresponding feedback indicators, network status information, and coding strategy and scheduling strategy parameters to generate the transmission performance evaluation result.
[0124] Example 1:
[0125] To verify the feasibility of this invention in practice, it was applied to the urban road integrated perception and transmission environment of a city's intelligent traffic management center. In this scenario, urban arterial roads and expressways are interwoven, traffic flow fluctuates with morning and evening peak hours, and a large number of video surveillance devices, environmental and traffic sensors, and vehicle-to-road communication devices simultaneously generate high-density multimodal data, including high-definition video streams, temperature and humidity data, visibility information, and vehicle operation status broadcasts. Traditional methods struggle to guarantee critical data during peak hours, easily leading to significantly increased latency, exacerbated loss of perceived information, and slow dispatch response, affecting traffic situation analysis and emergency command. This invention addresses the high-concurrency, multimodal, and dynamically fluctuating link scenarios in cities by improving the CorrNet network through regional segmentation, differentiated coding strategies, and a tabu search scheduling framework to enhance the transmission stability and priority guarantee capability of critical data in complex environments.
[0126] During application, the system runs continuously from early morning to evening, synchronously collecting multimodal data along the roads. All data, after entering the system, undergoes time alignment, structured parsing, noise reduction, and time windowing to form a standardized format. Based on the urban road and communication topology, the system divides the roads into multiple regions and maps video clips, sensor clips, and collaborative information to the corresponding regions. Subsequently, an improved CorrNet shared hidden layer structure is used to extract regional features, generating regional continuity and redundancy features, enabling the final multimodal features to distinguish between high-risk congestion areas and ordinary areas. According to scene rules, the system assigns higher importance labels to video clips containing accident risks and lower importance labels to low-impact data, allowing for precise implementation of differentiated compression and redundancy strategies, thereby enhancing the resilience of important data and reducing overall bandwidth pressure.
[0127] When network bandwidth decreases, packet loss increases, or link load rises during peak evening hours, the system dynamically adjusts transmission paths based on real-time network conditions. Encoded data units are prioritized, the scheduling module calculates the cost of candidate paths, and iteratively optimizes using tabu search to dynamically allocate paths and bandwidth. During transmission, the system continuously collects feedback metrics such as latency, packet loss, retransmission, and decoding quality, and combines these with network status information to generate a transmission performance evaluation result, which is then uploaded to the monitoring terminal to ensure the reliability of policy execution.
[0128] After continuous operation during morning peak, off-peak, and evening peak hours, monitoring results show that, under complex network conditions, this invention, relying on regional shared hidden layer feature extraction, importance identification, differentiated coding, and intelligent scheduling, enables high-value video data to maintain stable, low-latency transmission even during congested periods, avoiding the performance drop caused by uniform coding or fixed paths in traditional solutions. Simultaneously, moderate compression and lightweight redundancy of low- and medium-value data effectively save bandwidth resources, significantly improving overall network utilization and verifying the applicability and stability of this invention in multimodal data transmission for smart cities.
[0129] Table 1. Performance Comparison of the Invention and Traditional Methods in Smart City Multimodal Data Transmission Scenarios.
[0130]
[0131] As can be clearly seen from Table 1, the method of the present invention is superior to the traditional method in many indicators.
[0132] Firstly, comparing the average latency of critical video data, the method of this invention achieves 145ms, significantly lower than the 310ms of traditional methods. The core reason for the reduced latency lies in the fact that this invention extracts importance labels through an improved CorrNet network and applies a lower compression ratio and stronger redundancy strategy to high-importance data during the encoding stage, automatically giving critical data higher priority during queue scheduling. Simultaneously, the tabu search scheduling framework adjusts paths in real time based on network conditions, ensuring that critical data is promptly diverted from congested links, thereby guaranteeing stable, low-latency transmission.
[0133] Regarding the packet loss rate of critical video data, the method of this invention achieves only 2.1%, far lower than the 7.8% of traditional methods. This is because traditional methods employ a uniform encoding strategy for all data, resulting in weak recovery capabilities after packet loss. In contrast, this invention generates target redundancy levels based on importance and network status. High-importance video segments are appended with cyclic check sequences, group check sequences, and redundant check segments, enabling high recovery capabilities even under weak link conditions, thereby significantly reducing the packet loss rate.
[0134] In the overall retransmission rate of multimodal data, the method of this invention achieves 4.3%, significantly better than the 12.4% of the traditional method. This is because the present invention implements differentiated compression during the encoding stage, reducing the load on low-to-medium importance sensing data through lightweight encoding, thereby reducing link pressure from the source. Meanwhile, high-importance data, with the addition of sufficient redundancy information, can be correctly recovered during the first transmission, reducing the need for retransmission and thus lowering the overall system retransmission rate.
[0135] In terms of high-value video decoding quality score, this invention achieved 89 points, significantly higher than the 68 points of the traditional method. This is because the invention employs higher-quality parameter combinations for high-value data in the transform coding and redundancy addition stages, ensuring high recoverability even with transmission link fluctuations. Simultaneously, by assigning greater coding resources to key videos through importance tags, it retains more feature information during compression ratio adjustments.
[0136] Regarding overall network bandwidth usage, the method of this invention achieves only 605 Mb / s, a significant reduction compared to the traditional method's 860 Mb / s. This improvement stems from a differentiated coding strategy: high compression ratios are applied to insensitive data, while redundancy is precisely allocated to sensitive data, thus avoiding indiscriminate bandwidth consumption caused by a one-size-fits-all approach to coding.
[0137] In terms of on-time delivery rate of critical data, this invention achieves 93%, far exceeding the 71% of traditional methods. This is because this invention uses a priority sorting queue, which prioritizes the scheduling of highly important data. At the same time, the tabu search algorithm continuously optimizes the path and bandwidth allocation, ensuring that critical data can still be delivered on time even in complex network conditions.
[0138] Regarding the average load fluctuation of nodes, this invention achieves 11%, significantly better than the traditional method's 26%. This is mainly due to the load balancing feature in the path scheduling of this invention: the path cost evaluation function includes the node load dimension, which prevents scheduling from continuously placing critical business load on a few nodes, resulting in a more balanced overall load.
[0139] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A smart city oriented multi-modal data transmission optimization method, characterized in that, Includes the following steps: Collect raw multimodal data generated in the city, preprocess it, and generate a standardized multimodal dataset; An improved CorrNet network is used to perform joint feature extraction on a standardized multimodal dataset, outputting multimodal features and corresponding importance labels according to preset scene rules; The output of the multimodal features and importance labels specifically includes: Each data segment in the standardized multimodal dataset is indexed according to its collection location and the corresponding urban network area, and each data segment is associated with a preset regional identifier. The urban network consists of the connection relationships of edge nodes, roadside units and base stations deployed within the city. Its topology is obtained by collecting network configuration data, link connection information and node identification information of each node, and performing regularization parsing processing on each connection relationship. The multimodal data segments that have completed the region indexing process are input into the improved CorrNet network. Based on the region identifier to which each data segment belongs, region-related feature representations are generated in the corresponding region-shared hidden layer blocks. The improved CorrNet network replaces the single shared hidden layer of the original CorrNet network with a multi-region shared hidden layer structure based on the city topology, introduces regional blocks corresponding to city network nodes and adds topological regularization constraints between regions, and adds supplementary channels in the shared hidden layer as redundant feature processing paths. In the regional shared hidden layer block, the regional related feature representation is processed by regional association. Based on the topological connectivity of adjacent regions in the urban network, the regional related features belonging to adjacent regions are continuously weighted according to the preset regional association coefficient to obtain the regional continuous feature set. Redundant feature extraction is performed on multimodal data segments by supplementary channels in the improved CorrNet network. Redundant features corresponding to each data segment are generated based on the hidden layer random deactivation rule and redundant channel reconstruction constraints, thus generating a set of redundant features. The set of regional continuity features and the set of redundant features are fused together to form a set of shared spatial features that includes regional continuity attributes and redundancy recovery attributes, based on the feature fusion rules. A calibration and reconstruction constraint is applied to the shared spatial feature set. A reconstruction calibration value is generated by weighted summation of regional reconstruction error and redundant reconstruction error. The shared spatial feature set is then calibrated for consistency based on the reconstruction calibration value to generate multimodal features. The importance generation process is performed on the calibrated shared spatial feature representation. Based on the preset scene rules, the regional attributes, semantic information and task sensitivity of each data segment are comprehensively calculated to obtain the corresponding importance label and output it together with the multimodal features. Operational data is collected from edge nodes, roadside units, and base stations to organize and analyze the urban network and generate network status information. Multimodal features, importance labels, and network state information are correlated and analyzed, and corresponding encoding methods, compression ratios, and redundancy levels are determined for various types of data to generate encoding control instructions. The generation of the encoding control instructions specifically includes: Multimodal features, importance labels, and network state information are associated according to the correspondence of data segments to generate data description records for each data segment; Read the importance level identifier and importance score value corresponding to each data segment in the data description record, divide the data segments into priority level categories with different numbers according to the importance level identifier, use the importance score value as the priority quantification score used for coding priority calculation, and generate priority parsing results; The network status information in the data description record is parsed, the comprehensive operation score value corresponding to the data segment transmission path and the operation data on the corresponding path are read, comparison processing is performed according to the preset threshold set, the network status level corresponding to the data segment is determined, and the network status level is mapped to the network status level quantization value to generate the network status parsing result. Based on the priority parsing results and the network state parsing results, a joint coding strategy is calculated to determine the coding method, compression ratio and redundancy level of each data segment; The encoding method, compression ratio, and redundancy level are combined with the corresponding data segment identifier to generate encoding control instructions; Based on the encoding control instructions, the multimodal data is subjected to differentiated compression and encoding processing to generate a priority-ordered queue of data to be transmitted, with encoded data units as elements. Based on network status information and historical scheduling results, path calculation and resource allocation analysis are performed on the data to be transmitted. Scheduling control instructions are generated through tabu search and implemented in urban network nodes. During data transmission, feedback indicators are continuously collected, and the effectiveness of the current encoding and scheduling strategies is evaluated in conjunction with network status information. Transmission performance evaluation results are generated and uploaded to the system monitoring terminal.
2. The multimodal data transmission optimization method for smart cities according to claim 1, characterized in that, The multimodal raw data includes video data from cameras, environmental and traffic sensor data from roads and traffic infrastructure, and collaborative sensing data from vehicle-to-road communication devices. The preprocessing includes performing time alignment, format parsing, structured processing, noise removal, and outlier removal on the video data based on a unified time reference; performing time synchronization, format parsing, noise suppression, and outlier filtering on the environmental and traffic sensor data; performing time calibration, field parsing, and structured processing on the collaborative sensing data; and performing slicing and resampling on all multimodal data according to time windows.
3. The multimodal data transmission optimization method for smart cities according to claim 1, characterized in that, The generation of the network status information specifically includes: Operational data is collected from edge nodes, roadside units, and base stations, and timestamps and node identifiers are recorded during collection to generate an operational dataset. The operational data includes network bandwidth, latency, packet loss rate, and node load. The data in the running dataset is processed by discarding abnormal time points to form a running data sequence on a unified time axis. Then, normalization processing is performed on network bandwidth, latency, packet loss rate and node load respectively, and each running data is mapped to the same numerical range to obtain a standardized running data sequence. The standardized operational data sequence is grouped according to the links and nodes in the urban network topology, and network bandwidth, latency, packet loss rate and node load are combined into operational data vectors in a fixed order. For each running data vector, a comprehensive running score is calculated by multiplying network bandwidth, latency, packet loss rate and node load by their respective weight coefficients and summing them to generate a comprehensive running score value for the corresponding link and node. The comprehensive operational scores of each link and node are organized according to the time dimension and topological location to generate network status information with nodes and links as indexes and comprehensive operational scores and corresponding operational data as records.
4. The multimodal data transmission optimization method for smart cities according to claim 1, characterized in that, The generation of the data queue to be transmitted specifically includes: Based on the encoding control instructions, the data segments in the multimodal data are associated with the corresponding encoding methods, compression ratios, redundancy levels, and priority parsing results to generate a set of encodings to be processed. Select data segments sequentially from the set to be encoded, read the corresponding encoding method, compression ratio, redundancy level and priority quantization score, and generate corresponding encoding strategy records based on preset parameters; Based on the encoding method, target compression ratio and target redundancy level recorded in the encoding strategy record, differential compression and encoding processing is performed on the corresponding data segments. Encoded data units corresponding to the data segment identifier are generated through data fragmentation, transform coding, entropy coding and redundancy information appending, and priority level category and priority quantization score are written into the encoded data units. All encoded data units are sorted according to priority level category, priority quantization score, and time order of data segments to generate a priority-sorted queue of data to be transmitted, with encoded data units as elements.
5. The multimodal data transmission optimization method for smart cities according to claim 1, characterized in that, The generation of the scheduling control instructions specifically includes: Based on network status information and the queue of data to be transmitted, the city network topology is analyzed to resolve links and nodes, generating a scheduling input dataset; Select the start node and end node corresponding to the encoded data unit in the data queue to be transmitted from the scheduling input dataset. Construct several initial transmission paths according to the urban network topology. Allocate bandwidth usage ratio on each initial transmission path according to the link available bandwidth and the priority quantization score of the encoded data unit. Generate an initial scheduling scheme including path sequence and bandwidth usage ratio allocation results. Based on network status information and historical scheduling results, a path cost evaluation function is constructed for candidate transmission paths to calculate the path cost and generate a global scheduling evaluation value. The initial scheduling scheme is iteratively optimized using the tabu search method to obtain the target scheduling scheme. Based on the target scheduling plan, scheduling control instructions are generated and sent to the corresponding city network nodes for scheduling implementation.
6. The multimodal data transmission optimization method for smart cities according to claim 1, characterized in that, The generation of the transmission performance evaluation results specifically includes: Feedback indicators are collected from urban network nodes during data transmission and associated with data segments to generate a set of feedback indicator records. The feedback indicator record set is collected and structured in conjunction with network status information to generate a set of feedback evaluation inputs; Perform transmission performance calculations on the feedback evaluation input set to generate path-level transmission performance scores; All path-level transmission performance scores are aggregated according to the data segment to which the encoded data unit belongs, priority level category, and transmission time window. Transmission performance evaluation results oriented towards data segments are generated and uploaded to the system monitoring terminal.