A pedestrian paving stone environment perception method based on multimodal fusion

By employing a multimodal fusion-based pedestrian paving environment perception method, which utilizes pressure distribution, surface texture images, and temperature and humidity data, combined with dynamic thresholds and spatial features, the perception area and time window are dynamically adjusted. This overcomes the limitations of single-modal perception and enables efficient and accurate environmental condition monitoring and resource optimization.

CN121147639BActive Publication Date: 2026-03-13HANGZHOU LIHUAN ENVIRONMENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing methods for sensing pedestrian pavement environments rely on single-modal data, which makes it difficult to comprehensively and accurately reflect the environmental status. Furthermore, the unreasonable allocation of sensing resources leads to misjudgments and increased operation and maintenance costs.

Method used

A multimodal fusion method for permeable paving stone environment perception is adopted to acquire pressure distribution data, surface texture images and environmental temperature and humidity data. Combined with preset environmental parameter dynamic thresholds and spatial characteristics of the target perception area, the perception target area and time window are dynamically adjusted, and a progressive multi-source data fusion method is used for perception operations.

Benefits of technology

It achieves comprehensive and accurate perception of the environmental status of pedestrian paving stones, improves perception efficiency and accuracy, rationally allocates perception resources, reduces operation and maintenance costs, and adapts to dynamic changes in the environment and pedestrian flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121147639B_ABST
    Figure CN121147639B_ABST
Patent Text Reader

Abstract

This invention relates to the field of pedestrian environment sensing technology and discloses a method for pedestrian paving stone environment sensing based on multimodal fusion. The method includes acquiring a multimodal sensing data sequence of pedestrian paving stones within a preset monitoring area, including pressure distribution, surface texture images, and environmental temperature and humidity data sequences; then, combining preset dynamic threshold environmental parameters with spatial characteristics of the target sensing area, dynamically sensing target areas are selected through positive equilibrium matching. The former is pre-configured based on historical sensing data statistical distribution and preset environmental anomaly judgment rules, while the latter is calculated based on paving stone paving density, peak pedestrian flow density, and the distribution range of anomaly area heat maps; next, multiple continuous sensing time windows are divided according to the target area boundary location and the distribution topology of multimodal sensing nodes; finally, sensing operations for each window are completed sequentially using a progressive multi-source data fusion method, achieving comprehensive sensing of the pedestrian paving stone environmental status.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pedestrian environment perception technology, specifically a pedestrian paving tile environment perception method based on multimodal fusion. Background Technology

[0002] In the field of urban pedestrian system operation and maintenance and pedestrian safety assurance, pedestrian paving stones, as infrastructure that directly interacts with pedestrians, play a crucial role in improving the safety of pedestrian spaces and optimizing operation and maintenance efficiency through environmental condition perception. Currently, environmental perception methods for pedestrian paving stones mostly rely on single-type sensing data for monitoring operations. For example, they may only use pressure sensors to obtain information about the stress on the paving stone surface, or only rely on temperature and humidity sensors to collect surrounding environmental parameters. Such single-modal sensing methods have significant limitations.

[0003] From a practical application perspective, the environmental condition of pedestrian paving stones is influenced by a combination of factors, and relying solely on a single type of data is insufficient to comprehensively and accurately reflect the true situation. For example, when the surface of a paving stone is damaged, it may be accompanied by abnormal pressure distribution and changes in surface texture. If monitoring is only conducted through pressure sensors, misjudgments may occur due to differences in pedestrian traffic conditions. Furthermore, under extreme weather conditions, changes in temperature and humidity may alter the material properties of the paving stones, thereby affecting their load-bearing capacity. Without the perception of visual information such as surface texture, it is difficult to detect potential safety hazards in a timely manner. In addition, most existing sensing methods adopt a monitoring mode with fixed areas and fixed time windows, failing to consider factors such as differences in pedestrian flow density in different areas and dynamic changes in environmental parameters. This leads to unreasonable allocation of sensing resources, resulting in untimely data collection in areas with high pedestrian traffic, while wasting sensing resources in areas with sparse pedestrian traffic.

[0004] Traditional sensor data processing lacks an effective multi-source data fusion mechanism. Different types of sensor data operate independently, failing to fully extract the correlations between them, resulting in low accuracy in environmental condition assessments. For example, pressure distribution data might indicate an anomaly in a certain area, but combined with temperature and humidity data, it might suggest that the area's paving stones are frozen due to low temperatures. Failure to fuse these two types of data could lead to a misjudgment of paving stone structural damage, resulting in unnecessary repairs and increased maintenance costs. These problems make existing pedestrian paving stone environmental sensing methods insufficient to meet the practical needs of refined management, efficient operation and maintenance, and pedestrian safety in urban pedestrian systems. A new method is needed that can integrate multimodal sensor data, dynamically adjust sensing strategies, and achieve accurate sensing. Summary of the Invention

[0005] The purpose of this invention is to provide a pedestrian paving stone environment perception method based on multimodal fusion to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides a pedestrian paving tile environment perception method based on multimodal fusion, the method comprising:

[0007] Acquire a multimodal sensing data sequence of pedestrian pavers within a preset monitoring area. The multimodal sensing data sequence includes a pressure distribution data sequence, a surface texture image sequence, and an environmental temperature and humidity data sequence.

[0008] Based on the positive equilibrium relationship between preset dynamic threshold environmental parameters and spatial characteristics of the target sensing area, the preset dynamic threshold environmental parameters are matched with the spatial characteristics of the target sensing area to filter out dynamic sensing target areas that meet preset conditions; the preset dynamic threshold environmental parameters are obtained by pre-configuration based on the statistical distribution of historical sensing data and preset environmental anomaly judgment rules; the spatial characteristics of the target sensing area are calculated based on the pedestrian paving density, peak pedestrian flow density, and the distribution range of the heat map of abnormal areas;

[0009] Based on the boundary location of the dynamically sensing target area and the distribution topology of the multimodal sensing nodes obtained through screening, the sensing task of the dynamically sensing target area is divided into multiple continuous sensing time windows.

[0010] Based on the sequence of continuous sensing time windows, a progressive multi-source data fusion method is used to perform the sensing operation for each continuous sensing time window in turn, until the sensing operation corresponding to the last continuous sensing time window is completed.

[0011] Preferably, the matching of the preset environmental parameter dynamic threshold with the target perception area spatial features based on the positive equilibrium relationship between the preset environmental parameter dynamic threshold and the target perception area spatial features includes:

[0012] Identify environmental state control factors used for cross-level transmission and environmental feature feedback factors used for reverse feedback within the dynamic sensing target area;

[0013] The environmental state control factor is input into the next-level sensing subsystem, and the environmental feature feedback factor is transmitted back to the previous-level sensing subsystem to optimize the sensing parameter configuration.

[0014] Preferably, based on the boundary positions of the dynamically perceived target region and the distribution topology of the multimodal sensing nodes obtained through screening, the sensing task of the dynamically perceived target region is divided into multiple continuous sensing time windows, including:

[0015] Obtain the boundary location and multimodal sensing node distribution topology of the dynamically sensing target area; the multimodal sensing node distribution topology is a spatial distribution map formed by pressure sensor nodes, image acquisition nodes and environmental sensor nodes within the sensing area;

[0016] Based on the boundary location of the dynamically sensing target area and the topology of the multimodal sensing node distribution, the multimodal data acquisition time axis of the dynamically sensing target area is obtained, and the multimodal data acquisition time axis is divided into multiple continuous sensing time windows.

[0017] Preferably, before dividing the sensing task of the dynamically sensing target area into multiple consecutive sensing time windows based on the boundary position of the dynamically sensing target area and the distribution topology of the multimodal sensing nodes obtained through screening, the method further includes:

[0018] Multimodal feature classification is performed on the set of abnormal event records in the historical perception database to determine multiple sets of abnormal event features for different categories.

[0019] By traversing multiple sets of categorized abnormal event features, environmental features are extracted in a concentrated manner to obtain multiple concentrated values ​​of environmental features and the associated duration of multiple environmental features.

[0020] Preferably, the step of traversing multiple sets of categorized abnormal event features to extract environmental features and obtain multiple sets of environmental feature values ​​and multiple environmental feature association durations includes:

[0021] By associating the duration of multiple environmental features as multiple temporal receptive fields of multiple feature extraction levels, and using the configured multiple feature extraction levels to perform multi-scale environmental feature analysis on multimodal sensing data sequences, a set of environmental features at multiple scales is obtained.

[0022] Preferably, before performing the sensing operation for each continuous sensing time window sequentially using a progressive multi-source data fusion method according to the order of the continuous sensing time windows, the method further includes:

[0023] Based on the predicted start time values ​​of multiple consecutive sensing time windows and the preset time association standard, the sensing coordination matching degree between adjacent consecutive sensing time windows is calculated.

[0024] When the perceived coordination matching degree does not reach the preset matching threshold, the re-optimization parameters are determined based on the time delay coefficient and the distribution of non-matching features.

[0025] The sensing parameters of the current continuous sensing time window are dynamically adjusted based on the re-optimization parameters.

[0026] Preferably, the step of performing the sensing operation for each continuous sensing time window sequentially using a progressive multi-source data fusion method includes:

[0027] Cross-modal interactive fusion of environmental feature sets at multiple scales is performed to obtain a fused environmental feature set;

[0028] The fused environmental feature set is used as an environmental state control factor and input into the next-level sensing subsystem for state coordination analysis.

[0029] Preferably, the step of performing cross-modal interactive fusion of multiple scale environmental feature sets to obtain a fused environmental feature set includes:

[0030] Extract the normalized value of feature similarity between feature sets of different scales;

[0031] Dynamically weight the feature similarity normalization value and the set of environmental features at adjacent scales;

[0032] Cross-modal interaction fusion of all scale environmental feature sets is achieved through an iterative weight allocation process.

[0033] Preferably, after completing the sensing task corresponding to the last continuous sensing time window, the process includes:

[0034] Based on the total amount of multimodal sensing data actually collected and the spatial characteristics of the target sensing area, the environmental state sensing coverage rate is calculated.

[0035] Preferably, before calculating the environmental state perception coverage rate based on the total amount of actually collected multimodal sensing data and the spatial characteristics of the target sensing area, the method further includes:

[0036] The threshold for identifying abnormal events is iteratively adjusted based on the concentrated values ​​of environmental features until the rate of change of the environmental feature aggregation is less than the preset rate of change threshold.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] By acquiring multimodal sensing data sequences including pressure distribution data sequences, surface texture image sequences, and environmental temperature and humidity data sequences, the limitations of traditional single-modal sensing methods are broken. This allows for comprehensive capture of environmental status information of pedestrian paving stones from multiple dimensions, including mechanics, vision, and environment. The collaborative acquisition of multiple types of sensing data can cover multiple key evaluation dimensions such as the structural performance of the paving stones themselves, surface morphological characteristics, and surrounding environmental conditions. This avoids misjudgments or omissions of environmental status due to insufficient data from a single type, making the perception of the environmental status of pedestrian paving stones more comprehensive and objective.

[0039] In the dynamic sensing target area selection stage, matching and selection are performed based on the positive equilibrium relationship between preset dynamic threshold environmental parameters and the spatial characteristics of the target sensing area. The preset dynamic threshold environmental parameters are pre-configured by combining historical sensing data statistical distribution and environmental anomaly judgment rules, fully referencing past sensing experience to ensure the rationality and adaptability of the threshold setting, effectively addressing dynamic changes in environmental parameters under different seasons and weather conditions. The spatial characteristics of the target sensing area are calculated based on the density of pedestrian paving stones, peak pedestrian flow density, and the distribution range of anomaly area heat maps, fully considering the actual usage and potential risk level differences in different areas. This selection method can accurately locate areas requiring key monitoring, making sensing operations more targeted and avoiding the waste of sensing resources under the traditional fixed-area monitoring model. It also ensures that sensing operations are prioritized in areas with abnormal environmental parameters, dense pedestrian flow, or high risk, improving sensing efficiency and effectiveness.

[0040] The sensing task for the dynamically sensing target area is divided into multiple continuous sensing time windows. This division is based on the boundary location of the target area and the topological distribution of multimodal sensing nodes. This ensures precise matching between the sensing task and the node layout, guaranteeing efficient sensing operations within each time window and avoiding incomplete or duplicate data collection due to excessively large sensing areas or uneven node distribution. The continuous sensing time windows allow for phased and orderly sensing operations based on the dynamic changes in the pedestrian pavement environment, adapting to changes in environmental parameters and pedestrian flow across different time periods. This makes the sensing process more flexible and enables timely capture of dynamic evolution trends in the environment.

[0041] A progressive multi-source data fusion approach is adopted, sequentially performing sensing operations within each consecutive sensing time window. This allows for the gradual integration of multimodal sensing data from each time window during the sensing process, fully exploring the correlations and potential information between different data types. Compared to the traditional method of processing each type of data independently, progressive fusion continuously optimizes data processing results as the sensing operation progresses. As the sensing time windows advance sequentially, the amount of fused data gradually increases, and the synergistic effects between data become increasingly apparent, gradually improving the accuracy of the assessment of the pedestrian pavement environment. Furthermore, this fusion method does not require waiting for all sensing data to be collected before unified processing; it enables timely data fusion and analysis after the sensing operation in each time window. This facilitates the rapid detection of environmental anomalies, shortens anomaly response time, and provides timely information support for operational and maintenance decisions or safety early warnings.

[0042] This method dynamically adjusts the target sensing area and sensing time window to achieve rational allocation of sensing resources. Sensing resources are concentrated in areas and time periods with high pedestrian traffic and high environmental risk to ensure effective sensing in key areas. Sensing resources are appropriately reduced in areas with low pedestrian traffic and stable environmental conditions to avoid resource idleness, thereby improving the overall operational efficiency of the sensing system and reducing sensing costs. In summary, this method effectively improves the comprehensiveness, accuracy, and efficiency of pedestrian pavement environmental sensing, better meeting the actual needs of refined management, operation, and pedestrian safety assurance in urban pedestrian systems. Attached Figure Description

[0043] Figure 1 This is a schematic diagram illustrating the working principle of the pedestrian paving stone environment perception method based on multimodal fusion described in this invention.

[0044] Figure 2 A flowchart for classifying and extracting features of abnormal events;

[0045] Figure 3 A flowchart for calculating the coordination matching degree and optimizing parameters;

[0046] Figure 4 This is a flowchart for cross-modal interaction fusion. Detailed Implementation

[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] Please see Figure 1 This invention provides a method for perceiving pedestrian paving stones based on multimodal fusion, the method comprising:

[0049] Dynamic monitoring of the pedestrian paving environment is achieved through the collaborative acquisition and analysis of pressure distribution data sequences, surface texture image sequences, and environmental temperature and humidity data sequences. The system first acquires multimodal sensing data sequences within a pre-defined monitoring area. These sequences include spatiotemporal distribution data collected by a pressure sensor array, surface texture change images captured by a high-resolution camera, and microenvironmental parameters recorded by a distributed temperature and humidity sensor network. By establishing a mapping relationship between the statistical distribution of historical sensing data and anomaly detection rules, dynamic environmental parameter thresholds with adaptive adjustment capabilities are pre-configured. A spatial grid model is constructed based on the pedestrian paving density, and a probability distribution map of anomaly areas is generated by combining the thermodynamic simulation results of pedestrian flow density peaks, ultimately forming a spatial characteristic description of the target sensing area. Based on the matching results of dynamic thresholds and spatial characteristics, the system automatically divides the dynamic sensing target areas that require key monitoring and decomposes the sensing task into temporally continuous sensing windows according to the multimodal node topology, employing a progressive fusion strategy to complete the hierarchical processing of multi-source data. For example, a standardized paving tile unit integrating three types of sensing modules was developed: a 4×4 array of thin-film pressure sensors (accuracy up to 0.1N) is embedded inside the tile, connected to an edge computing chip via a flexible circuit board; the surface is covered with a wear-resistant transparent ceramic layer, beneath which is a built-in 1-megapixel miniature camera (equipped with infrared illumination to adapt to low-light environments), with the lens angle calibrated to ensure full coverage of the tile surface; a sensor interface is reserved on the side for connecting temperature and humidity sensing probes (measurement range: temperature -20℃~60℃, humidity 0~100%RH, sampling frequency 1Hz). The paving tile units are connected via a PoE power supply and data transmission integrated interface to form a grid-like sensing network. During installation, they are embedded into the gaps between regular paving tiles at a standard 50cm×50cm interval, without damaging the original pavement structure.

[0050] The regional management terminal deploys a data aggregation gateway, which periodically wakes up local brick units via a self-developed communication protocol (compatible with NB-IoT) to simultaneously collect pressure distribution matrices (refreshed every 10ms), surface texture images (captured one frame every 2s), and temperature and humidity data (recorded in sets every 30s). The raw data is preprocessed (e.g., pressure data noise reduction, image distortion correction) before being packaged and uploaded, forming a structured multimodal sensing data sequence. Simultaneously, the gateway's built-in storage module caches nearly 24 hours of data to prevent data loss due to network interruptions. Data is automatically retransmitted after network recovery, ensuring the continuity and integrity of the data sequence and providing a reliable data foundation for subsequent dynamic sensing target area selection.

[0051] Example 1: During the matching process of dynamically perceived target areas, the system performs multi-dimensional analysis of environmental state parameters through a feature decomposition module. This module adopts a hierarchical feature extraction architecture. The first layer processes the raw sensor data stream, converting the voltage signal of the pressure sensor into a pressure distribution matrix, the RGB data of the image acquisition node into a grayscale texture map, and quantizing the analog signal of the environmental sensor into a digital reading. The second layer of feature processing focuses on the spatiotemporal characteristics of various types of data, performing gradient field calculation on the pressure distribution matrix to identify the propagation direction and intensity of pressure changes; performing optical flow analysis on the texture image sequence to capture the motion trajectory of surface microstructures; and performing time series decomposition on environmental parameters to separate periodic variation components and sudden fluctuation components. The feature vector set formed after the two layers of processing is fed into a factor classifier for deep analysis.

[0052] The factor classifier consists of parallel neural network branches, each performing a nonlinear transformation on the feature vectors of a specific modality. The pressure feature branch uses a 3D convolutional kernel to extract the temporal evolution pattern of spatial pressure distribution, the image feature branch uses an attention mechanism to focus on texture anomaly regions, and the environment feature branch uses a long short-term memory network to capture long-term dependencies in parameter changes. The higher-order features output by each branch undergo cross-correlation analysis in the fusion layer, and their correlation strength is determined by calculating the mutual information entropy between features. Feature combinations with correlations exceeding a threshold are labeled as environmental state control factors, which reflect the cross-level transmission patterns of environmental states. For example, when a sudden change in pressure gradient accompanied by accelerated texture changes is detected in a certain region, the system will pass the joint pattern of these two features as a control factor to the next-level subsystem.

[0053] The identification process of environmental feature feedback factors focuses on the reverse transmission mechanism of abnormal events. The system establishes a feature backtracking channel for abnormal events. When a certain area is determined to be in an abnormal state, feature source analysis is triggered. This analysis locates the spatial center point of the abnormal event and expands outward to search for the change trajectory of related features. For example, for crack damage on the surface of floor tiles, the system will backtrack the history of texture changes, pressure concentration, and temperature and humidity fluctuations in the area before the crack appeared. After time inversion processing, these historical feature data generate a set of feature vectors describing the anomaly formation process, which are transmitted to the upper-level system as feedback factors. The feedback channel adopts a bidirectional transmission architecture, transmitting real-time monitoring data in the forward direction and historical analysis results in the reverse direction. The transmission process of control factors and feedback factors adopts different encoding strategies. The transmission of control factors uses a lightweight encoding scheme, converting feature vectors into low-dimensional latent space representations, reducing data transmission volume while retaining key feature information. The encoding of feedback factors uses differential compression technology, transmitting only the deviation between the current feature and the historical benchmark, improving the transmission efficiency of the reverse channel. The encoders and decoders of the two types of factors share some network layer parameters, maintaining the consistency of feature processing while taking into account the specific needs of different transmission directions. In the transmission path, control factors are directly sent to each sensing node through a bus network, while feedback factors are sent back to the central processing unit through point-to-point connections.

[0054] After receiving feedback factors, the upper-level system initiates a parameter optimization process. This process assesses the reliability of the feedback factors and eliminates false feedback caused by noise interference. Effective feedback factors are input into a parameter adjustment model, which generates parameter adjustment suggestions by analyzing the correlation between historical optimization records and current feedback. For example, if multiple feedbacks consistently show that the sensitivity of a pressure sensor in a certain area decreases under high-temperature conditions, the system will establish a temperature compensation coefficient and automatically adjust the calibration parameters of the sensor in that area. Parameter adjustment employs a gradual strategy, making only small corrections to a subset of parameters each time to avoid drastic changes that could lead to system instability. The corrected parameters are then sent to each sensing node via the control channel, completing one full feedback optimization cycle.

[0055] During system operation, the system continuously maintains state tables for control factors and feedback factors. The control factor state table records the transmission path, timeliness, and update frequency of each factor, ensuring that lower-level systems obtain the latest environmental status information. The feedback factor state table tracks the source tracing depth, analysis progress, and optimization results of each factor, forming a closed-loop management mechanism. The two state tables are synchronized using timestamps. When a time mismatch is detected between a control factor and its corresponding feedback factor, a data retransmission mechanism is triggered. This bidirectional factor management approach enables the system to adapt to dynamic changes in environmental conditions and maintain consistent perception accuracy.

[0056] In areas with a high incidence of abnormal events, the system activates an enhanced monitoring mode. In this mode, the update frequency of control factors is increased, and the analysis depth of feedback factors is enhanced, resulting in high-density two-way data exchange. Simultaneously, the scale of factor transmission in non-critical areas is reduced, optimizing system resource allocation. The scope of the enhanced monitoring area is dynamically adjusted based on the spatial distribution density of feedback factors; when the number of feedback factors in a certain area exceeds a threshold, the monitoring range is automatically expanded. This flexible resource allocation mechanism ensures monitoring quality in key areas while maintaining the overall system's operational efficiency.

[0057] The interaction between control factors and feedback factors generates rich system log data. These logs record the time, content, and processing results of each factor transmission, forming a traceable operational chain. This log data is used for offline analysis of system performance and to identify potential optimization opportunities. For example, analyzing factor response latency in the logs can reveal network transmission bottlenecks; statistically analyzing the optimization effects of feedback factors can evaluate the effectiveness of different parameter adjustment strategies. The log analysis results guide long-term system improvement, gradually enhancing the intelligence level of environmental perception.

[0058] The system establishes differentiated factor management strategies for different types of sensing environments. In densely populated areas, the focus is on monitoring dynamic changes in pressure distribution, with control factors primarily based on high-frequency pressure fluctuation characteristics. In harsh environments, the emphasis is on detecting anomalies in temperature and humidity parameters, with feedback factors incorporating more calibration information from environmental sensors. This targeted strategy enables the system to adapt to diverse application scenarios and maintain stable performance. The management strategies for each area are dynamically adjusted based on real-time monitoring data to ensure the system is always in optimal operating condition.

[0059] Security measures during factor transmission include data encryption and integrity verification. Control factors are digitally signed before transmission to prevent information tampering; feedback factors employ end-to-end encryption to protect sensitive monitoring data. Each factor transmission includes a verification code, which is only processed after successful verification by the receiver. These security mechanisms ensure reliable system operation in complex network environments and prevent malicious interference from affecting perception results. Simultaneously, the system retains multiple historical versions of factors, allowing for rapid rollback to a stable state in the event of anomalies, enhancing the system's fault tolerance.

[0060] Through this two-way factor transmission mechanism, the system achieves closed-loop control of environmental perception. Control factors transmit key characteristics of the environmental state downwards, guiding sensing nodes to collect targeted data; feedback factors reflect the actual effect of system operation upwards, driving continuous optimization of parameter configuration. This two-way information flow architecture enables the system to continuously adapt to environmental changes, gradually improving perception accuracy and response speed, forming a self-improving intelligent monitoring system. Throughout the entire process, the generation, transmission, and processing of various factors follow strict standards and specifications, ensuring the predictability and stability of system behavior.

[0061] Example 2: See Figure 2 The construction of the multimodal sensing node distribution topology begins with the spatial calibration of all sensing devices within the area. The installation location of each pressure sensor node is precisely positioned at the centimeter level using a laser rangefinder, recording its row and column coordinates within the pedestrian paving matrix. The spatial parameters of the image acquisition nodes include not only the physical location of the camera itself but also optical characteristics such as lens focal length and field of view. This data is used to construct a 3D model of the visible field of view for each node. The location information of the environmental sensor nodes, combined with their detection radius parameters, forms a spatial description of the spherical coverage area. This basic spatial data is input into the topology generation system, which calculates the signal transmission delay and coverage overlap between nodes to establish a network graph reflecting the actual physical connections.

[0062] The Delaunay triangulation algorithm is applied to model the spatial relationships between nodes. This algorithm connects three adjacent sensor nodes to form a triangular mesh. Each triangular unit represents a basic sensing area. The system evaluates the uniformity of network coverage by calculating the geometric properties of these units, such as side length ratios and angular distribution. In network edge regions or areas with weak coverage, the system automatically marks the locations where additional nodes are needed. The spatiotemporal coupling model takes into account the physical limitations of signal transmission, such as the response delay of pressure sensors, the computation time of image processing, and the transmission time of environmental data. These time parameters are converted into virtual displacements in a spatial coordinate system, forming a four-dimensional spatiotemporal topology.

[0063] The boundary determination process for dynamically perceived target areas employs adaptive contour recognition technology. This involves a preliminary scan of the pre-defined monitoring area to identify potential areas of interest characterized by abnormal pressure distribution, significant texture changes, or fluctuations in environmental parameters. The initial boundaries of these areas are delineated using edge detection algorithms and then overlaid with the node topology. When there is a significant deviation between the boundary line and the node coverage area, the system adjusts the boundary orientation based on the nearest node location to ensure sufficient sensing nodes support each dynamically perceived target area. The boundary adjustment algorithm prioritizes the collaborative coverage of multimodal nodes; that is, defining an ideal boundary requires an effective combination of at least three types of sensing nodes—pressure, image, and environment—within the area.

[0064] The generation of the multimodal data acquisition timeline relies on the clock synchronization mechanism of each sensing node. A high-precision time protocol is used to calibrate the time of all nodes and establish a unified time reference. During data acquisition, the sampling time of each node is recorded as an offset relative to the reference time. The division of the timeline takes into account the differences in the acquisition cycles of different modal data; for example, a pressure sensor may sample at a frequency of 100Hz, while an environmental sensor operates at a frequency of 1Hz. The system determines the basic time unit that can cover all data streams by calculating the least common multiple of the sampling times of each modality, serving as the basis for constructing the continuous sensing time window.

[0065] The historical perception database stores abnormal event records categorized according to multimodal features. Each category of abnormal events has a corresponding feature template. For example, the feature template for a floor tile crack event includes pressure distribution patterns, crack texture features, and temperature and humidity correlation changes. The classification process employs a hierarchical clustering algorithm to reduce the dimensionality of the original event data and perform density clustering in the feature space. The system automatically determines the optimal number of clusters, maximizing the feature similarity of similar events and highlighting the feature differences between different events. Each cluster center forms a typical feature vector for that type of event, used for classifying new events.

[0066] The extraction of environmental feature set values ​​employs a sliding time window analysis technique, performing time alignment processing on the historical records of each type of abnormal event to identify the patterns of feature changes before and after the event. For example, when analyzing events such as loose floor tiles, the system will find that the pressure distribution in the area exhibits an asymmetrical pattern before the event, surface texture shows slight displacement, and ambient humidity increases slightly. These cross-modal feature changes are quantified into numerical indicators, and their statistical central tendency before the event is calculated. The determination of feature association duration is achieved by analyzing the time delay correlation of multimodal data streams; for example, it is found that pressure anomalies typically appear several milliseconds earlier than texture changes detected in the image.

[0067] The multi-scale environmental feature analysis architecture is designed to address event detection needs across different time dimensions. The transient event detection layer focuses on capturing rapid changes lasting less than one second, such as the impact pressure waveform generated by a pedestrian suddenly falling. This layer employs high-pass filtering to remove low-frequency interference while retaining abrupt changes in the signal. The short-duration event analysis layer handles phenomena on timescales of one second to one minute, such as the gradual wear marks on the surface of paving stones. Feature extraction in this layer emphasizes the mid-frequency components of the signal, smoothing short-term fluctuations and highlighting trend changes through moving average techniques. The long-term environmental evolution monitoring layer focuses on slow changes on timescales of minutes and longer, such as the thermal expansion and contraction effects of paving stones during seasonal changes. This layer uses low-pass filtering and downsampling techniques to extract long-term evolution patterns from the signal.

[0068] The parameter configuration of the feature extraction layer is dynamically adjusted according to actual monitoring needs. During peak pedestrian traffic, the system enhances the processing capabilities of the transient event detection layer, increasing the sampling frequency of pressure data and the frame rate of image analysis. When environmental conditions change drastically, such as during heavy rain or extreme temperatures, the parameters of the short-term continuous event analysis layer are automatically optimized to improve sensitivity to sudden changes in temperature and humidity. The configuration of the long-term monitoring layer is adjusted according to seasonal variations; for example, the monitoring intensity of freeze-thaw cycle effects is increased in winter. This adaptive hierarchical configuration allows the system to flexibly respond to various monitoring scenarios and rationally allocate computing resources.

[0069] The feature transfer mechanism between adjacent layers employs an attention-weighted approach. The system calculates the contribution of lower-level features to the analysis of higher-level features and dynamically adjusts the weights of feature transfer. For example, when detecting the expansion of cracks in paving tiles, transient layer's minute displacement signals and short-term layer's trend change signals are given higher weights, while long-term layer's seasonal variation signals are given lower weights. This weighting mechanism ensures that each layer only transfers the most valuable feature information, reducing unnecessary data redundancy. The weight parameters are updated periodically based on historical data analysis results, gradually optimizing the efficiency of feature transfer.

[0070] The time window partitioning algorithm comprehensively considers various practical constraints. The system evaluates the computational complexity, data transmission volume, and processing latency of each potential partitioning scheme, and selects the optimal time window configuration. The window boundaries are aligned with the node sampling times as much as possible to reduce information loss caused by data resampling. The window length is dynamically adjusted according to the current environmental state; under stable conditions, the window is appropriately extended to improve processing efficiency, while during periods of high anomaly incidence, the window is shortened to enhance real-time performance. Reasonable overlap areas are set between adjacent windows to ensure continuous detection of cross-window events.

[0071] The calculation of the coordination matching degree is based on a comprehensive evaluation of multi-dimensional indicators, analyzing the degree of matching between adjacent windows in terms of data acquisition completeness, feature extraction consistency, and event detection continuity. When the matching degree is lower than a threshold, the diagnostic module analyzes the specific reasons for the mismatch, such as node clock drift, network transmission congestion, or insufficient computing resources. Based on the diagnostic results, the system takes targeted adjustment measures, such as resynchronizing node clocks, optimizing data transmission paths, or reallocating computing tasks. These adjustment measures aim to minimize system disturbances and gradually restore the ideal coordination state.

[0072] The analysis of mismatched feature distributions employs a contrastive learning approach, comparing the feature distribution of the current window with that of historical normal conditions to identify anomalous features deviating from the normal range. These anomalous features are classified into two types: transient fluctuations and systematic shifts, each handled with different strategies. Transient fluctuations are smoothed using data filtering techniques, while systematic shifts trigger a parameter recalibration process. The comparison results of feature distributions are also used to update the system's normal feature baseline, enabling it to adapt to gradual environmental changes.

[0073] Example 3: See Figure 3 The multi-scale environmental feature analysis architecture adopts a three-level hierarchical processing mode, with each level corresponding to a specific temporal resolution and feature extraction method. The first-level processing unit is configured with a millisecond-level time window, focusing on capturing transient pressure fluctuations and microscopic changes in surface texture. This level uses a high-pass digital filter to remove low-frequency environmental noise while retaining rapidly changing components in the signal. Pressure data processing includes peak detection and waveform analysis to identify abnormal impact patterns; image data is processed using inter-frame differencing techniques to extract subtle texture motion trajectories. The second-level processing unit operates on a second-level time scale, extracting mid-frequency features through sliding time window analysis. Data processing at this level includes spatial gradient calculation of pressure distribution, regional statistics of texture changes, and short-term trend analysis of temperature and humidity parameters. The third-level processing unit processes data streams at the minute level and above, using downsampling and moving average techniques to extract long-term evolution patterns, such as changes in the thermal expansion coefficient of floor tiles and diurnal fluctuation patterns of environmental parameters.

[0074] The parameter configuration of the feature extraction layers follows an adaptive adjustment principle. The system monitors the data processing load and feature output efficiency of each layer in real time, dynamically adjusting the length and overlap ratio of the processing window. During peak pedestrian traffic periods, the time window of the first-level processing unit is automatically shortened, and the sampling frequency is increased accordingly. During periods of stable environmental conditions, the window length of the third-level processing unit is appropriately extended to capture a more complete long-term variation cycle. Data transmission between layers employs compression coding technology, transmitting only the key information after feature selection to reduce internal system communication overhead. The layer switching mechanism ensures a smooth transition between analyses at different time scales, avoiding time gaps in feature extraction.

[0075] Feature interactions between adjacent levels are achieved through a weighted fusion mechanism, calculating the contribution of lower-level features to the analysis of higher-level features and generating dynamic weight coefficients. These weight coefficients are adjusted according to changes in environmental conditions and monitoring targets. For example, when detecting the expansion of cracks in paving tiles, higher weights are given to transient features at the first level; while when analyzing the overall aging trend of paving tiles, long-term features at the third level receive more attention. The weight adjustment algorithm considers multi-dimensional indicators such as the signal-to-noise ratio, temporal correlation, and spatial consistency of features to ensure that the most valuable feature information is transmitted to the appropriate analysis level.

[0076] The time window partitioning algorithm introduces the concept of perceptual coordination matching degree to evaluate the data continuity and feature consistency between adjacent windows. This matching degree is calculated based on the similarity of feature distributions in overlapping window regions and time alignment accuracy. The system maintains a dynamically updated matching degree threshold; when the actual matching degree falls below the threshold, an optimization process is triggered. The optimization process first diagnoses the cause of the declining matching degree, which may involve issues such as node clock skew, network transmission latency, or competition for computing resources. Based on the diagnostic results, the system takes targeted adjustment measures, such as resynchronizing node clocks, optimizing data transmission paths, or reallocating computing tasks. These adjustments aim to minimize system disturbances and gradually restore the ideal coordination state.

[0077] The analysis of mismatched feature distributions employs a contrastive learning framework. A baseline feature distribution under normal operating conditions is established, and the feature vectors of the current window are projected onto this baseline space to calculate the degree of deviation. Deviation analysis distinguishes between systematic shifts and random fluctuations, employing different processing strategies for each. Systematic shifts may reflect real changes in environmental conditions, triggering recalibration of system parameters; random fluctuations are handled using data smoothing techniques. The feature baseline itself is also updated periodically to adapt to gradual environmental changes, with the update cycle automatically adjusted based on feature stability.

[0078] The calculation of the time delay coefficient considers various practical factors, including the transmission delay of data from the acquisition node to the processing center, such as signal conversion time, network queuing time, and protocol processing overhead. The calculated delay depends on the current system load and task scheduling strategy. These delay parameters are quantified into a unified time metric used to predict the start time of the next time window. The prediction algorithm employs a sliding window averaging method, adjusted based on the delay trends of recent periods. Dynamic adjustment of the delay coefficient ensures that the time window division matches the actual data processing capacity, avoiding task backlog or resource idleness.

[0079] The parameter re-optimization process employs an incremental adjustment strategy. The system identifies a set of parameters requiring optimization, such as sampling frequency, data compression ratio, and transmission power, and prioritizes them according to their impact on system performance. Each optimization adjusts only a few of the top-ranked parameters, observing the effects before deciding on subsequent steps. This gradual approach avoids system instability caused by simultaneously changing too many parameters. The optimization objective is to minimize system energy consumption and communication overhead while maintaining perceived quality. Optimization decisions are based on multi-dimensional evaluation metrics, including data integrity, real-time performance, and resource utilization.

[0080] The output of multi-scale feature analysis adopts a unified space-time encoding format. Each feature point is labeled with its source level, timestamp, and spatial coordinates, forming a standardized feature description vector. These vectors are input into the fusion analysis module to participate in cross-modal correlation calculations. The feature vector encoding process preserves the precision information of the original data while controlling the data volume through quantization techniques. The encoding scheme supports flexible feature combination methods, facilitating result comparison and comprehensive judgment between different analysis levels.

[0081] The inter-level feedback mechanism enables self-improvement in the analysis process. When an upper-level analysis unit detects missing or contradictory features, it can send a supplementary data collection request to lower-level units. The lower-level units then adjust their processing parameters based on the request, such as increasing sampling density or extending the observation period, to obtain more comprehensive feature information. This two-way interaction allows the system to dynamically optimize its analysis depth and breadth, ensuring basic monitoring needs are met while implementing enhanced analysis for key areas or abnormal events. The bandwidth allocation of the feedback channel is dynamically adjusted according to request priority, ensuring timely transmission of critical information.

[0082] The dynamic adjustment algorithm for the time window considers environmental change predictions over a future period, analyzes periodic patterns and trends in historical data, and predicts potential environmental conditions in the next stage. Based on the prediction results, window parameters are pre-adjusted, such as shortening the window length in advance during periods of expected high pedestrian traffic, or increasing the monitoring intensity of environmental parameters before severe weather warnings. The prediction model employs a rolling update mechanism, continuously incorporating new observation data to correct the prediction parameters. This proactive window management strategy enables the system to prepare resources in advance and respond more readily to environmental changes. The specific calculation formulas used in the feature extraction process are as follows:

[0083] ;

[0084] in: Indicates the first and the Multi-scale feature differences between time windows The total number of feature dimensions. For the first Adaptive weights for each feature dimension It is the first The rate of change of each feature over time and These represent the two windows at the [number]th [time]. The values ​​taken on each feature dimension It represents the maximum time span of the entire analysis period. This formula quantifies the distribution distance of different time windows in the multi-feature space, and is used to guide decisions on merging or splitting windows.

[0085] The system continuously monitors the processing status and data quality at each level during operation. Monitoring metrics include the completeness of feature extraction, time alignment accuracy, and computational resource utilization. When an abnormal state is detected, such as a sustained increase in processing latency or a rise in feature loss rate at a certain level, the system initiates a diagnostic procedure to pinpoint the root cause of the problem. Possible responses include reallocating computational resources, switching to a backup processing algorithm, or temporarily reducing the priority of secondary tasks. This self-monitoring mechanism ensures that the system maintains stable analytical capabilities under various operating conditions and adapts promptly to changes in environmental conditions. Monitoring data is also used for long-term system performance optimization, identifying potential areas for improvement by analyzing historical operation records.

[0086] Example 4: See Figure 4 The implementation of cross-modal interactive fusion is based on spatial relationship modeling of multi-scale feature sets. The system constructs a dynamic feature relationship graph, where nodes represent feature vectors of different modalities at different time scales, and edge weights reflect the correlation strength between features. Feature nodes of the stress modality contain spatial statistics of the stress distribution matrix, such as gradient magnitude, asymmetry index, and concentration coefficient; image modality nodes store the frequency domain decomposition results of texture features, including low-frequency contour information and high-frequency detail components; environmental modality nodes record the time-varying characteristics of temperature and humidity parameters and their spatial gradients. The connections between these nodes are determined through feature similarity analysis, and the similarity measure comprehensively considers the spatiotemporal co-occurrence patterns and physical meaning associations of features.

[0087] Feature similarity normalization employs a multi-stage calibration method, calculating the distance matrix in the original feature space, where each element represents the degree of difference between a pair of feature vectors. The distance metric combines the advantages of Euclidean distance and cosine similarity, reflecting both feature amplitude differences and directional consistency. The original distance matrix is ​​converted into a relative ranking to mitigate the impact of extreme values. A nonlinear mapping function compresses the ranking values ​​to the [0,1] interval, yielding a standardized similarity score. This scoring process uses independent parameter settings for each modality combination, respecting the inherent characteristics of different data types. For example, the similarity calculation of stress-image features focuses more on spatial correspondence, while the similarity calculation of stress-environment features emphasizes temporal synchronization.

[0088] The dynamic weight allocation mechanism is implemented by analyzing the network topology characteristics of feature nodes, identifying key nodes in the feature graph. These nodes are typically located at the intersection of multiple modal feature flows and possess high betweenness centrality. The weights of key nodes are appropriately increased to enhance their influence in the fusion process. Weight allocation also considers the timeliness of features; recently generated features typically receive higher weight coefficients. Weight adjustment employs an iterative optimization strategy, evaluating the internal consistency of the fusion results after each iteration and fine-tuning the weight distribution based on the evaluation feedback. This dynamic adjustment allows the system to adapt to changes in feature importance under different environmental scenarios; for example, in rainy conditions, the weight of the environmental humidity feature automatically increases.

[0089] The core step in cross-modal interactive fusion is the alignment and mapping of feature spaces. This involves establishing feature correspondences between modalities and projecting feature vectors from different modalities into a unified latent space. The projection process maintains semantic consistency of features; that is, the representations of the same physical phenomenon in different modalities should be close to each other in the latent space. For example, the phenomenon of loose floor tiles manifests as localized pressure concentration in the pressure modality and as texture deformation in the image modality. These two features are mapped to adjacent regions in the latent space. The spatial alignment algorithm employs an adversarial training strategy, using a discriminator network to make features from different modalities difficult to distinguish in the latent space, thereby enhancing the fusion effect.

[0090] The iterative fusion process employs an attention mechanism for hierarchical refinement. The initial fusion result serves as the base feature map, inputting into a multi-round refinement module. Each refinement round calculates the attention weights at each location within the feature map, highlighting the feature contributions of important regions. These attention weights are determined by both feature salience and environmental context; for example, in densely populated areas, the attention weight for stress features is increased. The number of refinement rounds is adaptively adjusted, terminating the iteration when the change in the fusion results between two consecutive rounds is less than a threshold. This design balances fusion accuracy and computational efficiency, avoiding unnecessary resource consumption.

[0091] The state coordination analysis subsystem receives the fused set of environmental features and performs multi-dimensional consistency checks. These checks include time synchronization analysis to check whether different modal features exhibit related changes within the same time interval; spatial consistency assessment to verify whether the spatial distribution patterns of features conform to physical laws; and physical rationality judgment to confirm that the numerical relationships between features are within theoretical ranges. The coordination score integrates these check results to reflect the credibility of the environmental state. Low coordination signals trigger the data verification process, where the system re-examines the quality of the original data and the accuracy of feature extraction, and initiates data re-acquisition if necessary.

[0092] The fusion results are stored using a hierarchical data structure. The base layer stores the original feature vectors and fusion weights for easy traceability analysis; the intermediate layer stores the aligned and dimensionality-reduced latent space representation, supporting rapid retrieval and comparison; the application layer contains derived features for specific monitoring tasks, such as anomaly probability maps and risk heatmaps. This hierarchical storage design balances data integrity and access efficiency, establishing index relationships between different layers to support flexible query combinations. Data updates employ an incremental strategy, recalculating and storing only the changed features to reduce system overhead. See Table 1.

[0093] Table 1: Example of Multimodal Feature Fusion Weight Allocation

[0094]

[0095] During system operation, the system continuously monitors the quality indicators of the fusion effect, including parameters such as feature coverage, inter-modal consistency, and computational timeliness. When a decline in fusion quality is detected, the diagnostic module analyzes possible causes, such as reduced data quality of a particular modality, increased feature alignment deviation, or unbalanced weight allocation. Based on the diagnostic results, the system automatically adjusts the fusion strategy, such as temporarily reducing the weight of the problematic modality, switching to a backup alignment algorithm, or increasing the number of iterations for refinement. This adaptive quality control mechanism ensures that the fusion process produces reliable results under various operating conditions.

[0096] The cross-modal fusion results serve multiple downstream application modules. The anomaly detection module uses fused features to identify environmental state changes exceeding normal ranges; the predictive analysis module estimates the future condition of paving stones based on feature evolution trends; and the decision support module integrates various features to generate maintenance suggestions. Each application module periodically provides feedback on its performance, which is used to optimize the parameter settings of the fusion strategy. For example, if the anomaly detection module reports a high number of false alarms of certain types, the system will adjust the fusion weights of the relevant features to improve the discrimination accuracy. This closed-loop optimization allows the fusion process to continuously improve, better meeting actual monitoring needs.

[0097] The intermediate data and final results generated during the feature fusion process undergo rigorous version management. The system records the parameter configurations, input data, and output results for each fusion, forming a complete processing chain. Version data supports result backtracking and comparative analysis, allowing for rapid identification of problematic steps when fusion anomalies are detected. The version management system also provides a data snapshot function, saving the fusion state at specific points in time for offline analysis and algorithm improvement. These mechanisms enhance the system's maintainability and transparency, providing reliable assurance for long-term operation.

[0098] Example 5: The calculation process for environmental state perception coverage employs a gridded spatial analysis method, dividing the target monitoring area into several regular grid cells. The size of each cell is determined based on the standard specifications of pedestrian paving stones and the sensor distribution density. Grid division considers the distribution of physical obstacles in the actual environment; the location information of surface facilities such as trees and streetlights is incorporated into the grid boundary adjustment parameters. The completeness of multimodal data for each grid cell is evaluated by scanning the data records of all sensor nodes within the cell, checking the temporal continuity and spatial coverage of pressure readings, image sampling, and environmental measurements. The grid cell status is marked with three levels: complete coverage, partial coverage, or no coverage, corresponding to three situations: data acquisition meeting monitoring requirements, basically meeting requirements, and severely insufficient coverage, respectively. The coverage calculation results are visualized in the form of a two-dimensional heatmap, intuitively showing the strong coverage areas and weak links of the monitoring network.

[0099] The iterative correction mechanism for the anomaly event discrimination threshold is built on a dynamic learning framework. During system initialization, the initial threshold is set using the statistical distribution characteristics of historical data, such as taking the 95th percentile of pressure anomaly values ​​as the initial pressure threshold. In actual operation, the system continuously records the feature vector and final judgment result for each discrimination event, forming a labeled sample library. Once the sample reaches a certain scale, the threshold optimization algorithm is initiated to recalculate the cluster centers and distribution boundaries of various anomaly features. The optimization process maintains a stable recall rate for anomaly detection and improves precision by adjusting the position of the discrimination surface. Threshold updates employ a small-step, incremental strategy, making only minor adjustments to the existing threshold each time, observing the detection effect after adjustment, and then deciding on the next optimization direction. This conservative update method avoids system instability caused by sudden threshold changes.

[0100] The monitoring of the rate of change of environmental feature aggregation employs a sliding time window statistical method. The system maintains a fixed-length observation window, within which the numerical distribution of various environmental features is statistically analyzed. After each window slide, the difference in feature distribution between the old and new windows is calculated, including indicators such as mean shift, variance change, and distribution pattern differences. The rate of change algorithm assigns different sensitivity coefficients to different types of features; for example, the rate of change calculation for pressure features focuses on sudden peaks, while the rate of change calculation for temperature features emphasizes long-term trend changes. When the rate of change of feature aggregation exceeds a preset threshold, the system initiates a feature reassessment process to check for sensor drift, drastic changes in environmental conditions, or data processing anomalies. Based on the assessment results, a decision is made on whether to trigger a threshold correction process.

[0101] The collaborative optimization of sensing coverage and anomaly detection is achieved through a two-way feedback mechanism. Coverage analysis results guide the regional differentiation of anomaly detection thresholds, appropriately relaxing threshold requirements in areas with insufficient coverage to reduce the risk of false alarms. Conversely, the spatiotemporal distribution characteristics of anomalies influence the weighting of coverage assessment, raising coverage quality requirements in areas with high anomaly incidence. This two-way adjustment enables the system to optimize overall monitoring efficiency under limited resource conditions. The collaborative optimization process is executed automatically on a regular basis, with the execution frequency dynamically adjusted according to the rate of environmental change, increasing the optimization frequency during seasonal transitions or periods of drastic weather changes.

[0102] Historical data retrospective analysis provides a reference for threshold adjustment. The system establishes a complete event log database with timestamps, saving the original data, feature extraction results, and handling feedback for each abnormal alarm. Retrospective analysis focuses on two types of events: successfully verified genuine anomalies and false alarms confirmed as false alarms. By comparing the differences in feature distribution between these two types of events, the system identifies potential biases in threshold settings. The analysis results are converted into threshold adjustment suggestions, such as increasing the weight of certain features in the judgment or adjusting the fusion ratio of different modal features. The depth of retrospective analysis can be flexibly adjusted according to the importance of the event, conducting detailed full-chain analysis for major abnormal events and using a simplified analysis process for routine minor events.

[0103] The statistical method for calculating the total amount of multimodal data considers the acquisition characteristics of different data types. Stress data is recorded in time-series format, and its effective sampling points and data completeness are statistically analyzed. Image data is analyzed for the number of effective frames and coverage angles. Environmental parameters are recorded for their sampling period and numerical range. The system defines quality indicators for each type of data, such as the signal-to-noise ratio of stress data, the sharpness of image data, and the stability of environmental parameters. These quality indicators participate in the weighted calculation of the total data amount, avoiding evaluation bias caused by simple counting. The statistical results of the total data amount are compared with the spatial characteristics of the target perception area, including parameters such as area, pedestrian traffic level, and environmental complexity. Comparable coverage indicators are obtained through normalization.

[0104] The calculation of environmental feature clustering employs multi-dimensional clustering technology, projecting feature vectors from different modalities onto a unified space and identifying feature clustering regions through density clustering algorithms. For each clustered region, the statistical center point and distribution radius are calculated, serving as typical representatives of that type of feature. Monitoring changes in clustering not only focuses on the displacement of the center point but also analyzes changes in distribution morphology, such as the expansion or contraction of the distribution radius and changes in distribution density. This multi-dimensional clustering analysis can capture subtle changes in environmental conditions, providing a refined basis for adjusting system parameters.

[0105] The stability control mechanism in the threshold optimization process prevents excessive oscillations in system parameters. Before each threshold correction, the system assesses the consistency between the proposed adjustment magnitude and historical adjustment trends to avoid frequent reversals in adjustment direction between adjacent periods. An upper limit constraint is set on the correction magnitude to prevent abrupt changes in system behavior caused by excessively large single adjustments. For critical thresholds, the system retains multiple historical versions, allowing for rapid revert to a stable version if problems arise with the new threshold during actual operation. The stability control parameters themselves are also dynamically adjusted based on system operating conditions, allowing greater adjustment freedom during periods of environmental stability and employing a more conservative control strategy during periods of drastic environmental change.

[0106] The dynamic visualization of coverage results supports monitoring personnel's decision-making. Coverage calculations are overlaid on a real-world map, with gradient color levels representing coverage quality levels for different areas. The visualization interface supports timeline scrolling to display historical coverage trends; area drill-down to view detailed coverage data for specific grids; and overlay display of abnormal events to analyze the relationship between coverage quality and event detection rates. Visualization parameters are customizable based on user roles and task requirements; for example, maintenance personnel may focus on areas with weak coverage, while management personnel may emphasize overall coverage trends. This interactive visualization analysis tool transforms technical indicators into intuitive decision support information.

[0107] The system operation log fully records the entire process of coverage calculation and threshold optimization. Log content includes input parameters, intermediate results, and final output for each calculation, as well as the rationale, decision-making process, and implementation effects for threshold adjustments. Log data supports multi-dimensional retrieval by time range, regional location, or anomaly type, facilitating post-event analysis and problem troubleshooting. The log management system automatically identifies abnormal operation patterns or system errors, triggering an early warning mechanism. Regular log analysis reports summarize system performance, identify potential improvement directions, and provide a data foundation for continuous system optimization. Log recording uses a standardized format to ensure long-term storage and cross-platform compatibility.

[0108] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0109] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for environment perception of a pedestrian floor tile based on multi-modal fusion, characterized in that, The method comprises the following steps: obtain a multi-modal perception data sequence of the pedestrian floor tile in the preset monitoring area, wherein the multi-modal perception data sequence comprises a pressure distribution data sequence, a surface texture image sequence, and an environmental temperature and humidity data sequence; based on the positive balance relationship between the preset environmental parameter dynamic threshold and the spatial features of the target perception area, match the preset environmental parameter dynamic threshold with the spatial features of the target perception area to screen a dynamic perception target area that meets the preset conditions; the preset environmental parameter dynamic threshold is obtained based on the statistical distribution of historical perception data and the preset environmental anomaly judgment rule; based on the boundary position of the screened dynamic perception target area and the multi-modal perception node distribution topology, divide the perception task of the dynamic perception target area into multiple continuous perception time windows; according to the order of the continuous perception time windows, perform the perception task of each continuous perception time window in turn using a progressive multi-source data fusion method until the perception task corresponding to the last continuous perception time window is completed; the method of dividing the perception task of the dynamic perception target area into multiple continuous perception time windows based on the boundary position of the screened dynamic perception target area and the multi-modal perception node distribution topology comprises the following steps: obtain the boundary position of the dynamic perception target area and the multi-modal perception node distribution topology; the multi-modal perception node distribution topology is a spatial distribution map formed by the pressure sensor nodes, image acquisition nodes, and environmental sensor nodes in the perception area; based on the boundary position of the dynamic perception target area and the multi-modal perception node distribution topology, obtain a multi-modal data acquisition time axis of the dynamic perception target area, and divide the multi-modal data acquisition time axis into multiple continuous perception time windows.

2. The multi-modal fusion based floor tile environment perception method according to claim 1, wherein, Before the method of dividing the perception task of the dynamic perception target area into multiple continuous perception time windows based on the boundary position of the screened dynamic perception target area and the multi-modal perception node distribution topology, the method further comprises the following steps: perform multi-modal feature classification on a set of abnormal event records in a historical perception database to determine a plurality of classified abnormal event feature sets; extract environmental feature sets from the plurality of classified abnormal event feature sets to obtain a plurality of environmental feature set values and a plurality of environmental feature association durations. 3.The method of claim 2, wherein, The method of extracting environmental feature sets from the plurality of classified abnormal event feature sets to obtain a plurality of environmental feature set values and a plurality of environmental feature association durations comprises the following steps: use the plurality of environmental feature association durations as a plurality of time receptive fields of a plurality of feature extraction levels, perform multi-scale environmental feature analysis on the multi-modal perception data sequence using the plurality of configured feature extraction levels, and obtain a plurality of scale environmental feature sets.

4. The multi-modal fusion based floor tile environment perception method of claim 3, wherein, Before the method of performing the perception task of each continuous perception time window in turn according to the order of the continuous perception time windows using a progressive multi-source data fusion method, the method further comprises the following steps: based on the start time prediction values of the plurality of continuous perception time windows and a preset time association standard, calculate the perception coordination matching degree between adjacent continuous perception time windows; When the perception coordination matching degree does not reach the preset matching threshold, a re-optimization parameter is determined according to a time delay coefficient and a non-matching feature distribution; a perception parameter of a current continuous perception time window is dynamically adjusted based on the re-optimization parameter; the perception operation of each continuous perception time window is sequentially performed in the progressive multi-source data fusion manner, including: cross-modal interactive fusion is performed on the multiple-scale environment feature sets to obtain a fused environment feature set.

5. The multi-modal fusion based floor tile environment perception method according to claim 4, wherein, the cross-modal interactive fusion on the multiple-scale environment feature sets to obtain the fused environment feature set includes: a feature similarity normalization value between the different-scale environment feature sets is extracted; the feature similarity normalization value is dynamically weighted and distributed with the environment feature sets of adjacent scales; cross-modal interactive fusion is completed on all the environment feature sets through an iterative weight distribution process. 6.The method of claim 1, wherein, after the perception operation corresponding to the last continuous perception time window is completed, including: an environment state perception coverage rate is calculated based on an actual total amount of multi-modal perception data collected and a spatial feature of a target perception area.

7. The multi-modal fusion based floor tile environment perception method according to claim 6, wherein, before the environment state perception coverage rate is calculated based on the actual total amount of multi-modal perception data collected and the spatial feature of the target perception area, further including: an abnormal event discrimination threshold is iteratively corrected according to an environment feature central value until a change rate of an environment feature aggregation amount is less than a preset change rate threshold.

Citation Information

Patent Citations

  • Municipal infrastructure maintenance management system based on big data

    CN118822503A

  • Deformation monitoring and control method and system for road surface and storage medium

    CN120031392A