Solid waste intelligent monitoring method and system based on multi-source data
By using multi-source data monitoring and spatiotemporal alignment technology, the problem of single-modal recognition being susceptible to environmental interference has been solved, achieving high-precision fusion and accurate identification of multi-modal data, and improving the level of intelligence in solid waste classification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, solid waste identification relies on single-modal information, which is easily affected by environmental interference, and there is a lack of effective fusion mechanisms between different modal data, resulting in low identification accuracy.
By monitoring solid waste disposal areas using multi-source data, collecting raw multi-modal data using multi-modal sensors, and mapping the local sampling time of each modality data to the event reference time axis after preprocessing, spatiotemporal alignment is achieved. Cross-modal feature extraction and attention feature fusion are then performed to finally identify the solid waste category.
It significantly improves the accuracy and robustness of cross-modal fusion, enhances the adaptability to environmental changes, and improves the accuracy of solid waste identification.
Smart Images

Figure CN121302285B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent environmental sensing, and in particular to an intelligent monitoring method and system for solid waste based on multi-source data. Background Technology
[0002] With the acceleration of urbanization and the improvement of residents' living standards, the amount of solid waste generated continues to increase, making waste sorting and disposal a crucial aspect of urban management and environmental protection. Traditional waste sorting methods mainly rely on manual disposal and visual inspection, which are not only inefficient and costly but also easily affected by human factors, making it difficult to achieve real-time and accurate sorting monitoring. Therefore, utilizing the Internet of Things (IoT) and artificial intelligence (AI) technologies for intelligent sensing and identification of solid waste disposal points has become an important development direction for intelligent sanitation and urban management.
[0003] Currently, most garbage identification methods employ single-modal sensors, such as visual recognition methods based on video images or event detection methods based on acoustic features. While these single-modal sensor identification methods can achieve a certain degree of automatic identification under specific conditions, they still have many limitations in practical applications. On the one hand, single-modal systems are highly sensitive to environmental conditions; changes in lighting, background noise, occlusion, or signal loss can all lead to a significant decrease in identification accuracy. On the other hand, the lack of an effective spatiotemporal alignment mechanism between different sensors results in multi-source information not being accurately matched at the event level, making it difficult to achieve multimodal fusion and robust judgment. Summary of the Invention
[0004] The main objective of this invention is to provide a method and system for intelligent monitoring of solid waste based on multi-source data, aiming to solve the technical problems of existing solid waste identification relying on single-modal information, being easily affected by environmental interference, and lacking an effective fusion mechanism between different modal data, resulting in low accuracy of solid waste identification.
[0005] To achieve the above objectives, the present invention provides a method for intelligent monitoring of solid waste based on multi-source data, the method comprising the following steps:
[0006] Multi-source data monitoring is performed on the solid waste disposal area to obtain raw multimodal data, which is collected by multimodal sensors pre-installed in the solid waste disposal area;
[0007] The original multimodal data is preprocessed to obtain initial multimodal data;
[0008] The local sampling time of each modality in the initial multimodal data is mapped to the event reference time axis to generate a time mapping function. The event reference time axis includes multiple event nodes, which are generated based on the time nodes of related events of solid waste.
[0009] Based on the time mapping function and the spatial index of the initial multimodal data, the initial multimodal data is spatiotemporally aligned to obtain candidate multimodal data;
[0010] Cross-modal feature extraction is performed on the candidate multimodal data to obtain the feature embedding vector of each modality;
[0011] Attention feature fusion is performed on the feature embedding vectors of each modality data to obtain multimodal fusion features;
[0012] Based on the multimodal fusion features, the solid waste disposal area is used to identify and monitor the solid waste category.
[0013] Optionally, the step of performing spatiotemporal alignment of the initial multimodal data based on the time mapping function and the spatial index of the initial multimodal data to obtain candidate multimodal data includes:
[0014] Spatial dimension saliency analysis is performed based on the spatial index of the initial multimodal data to obtain salient parameters of local events;
[0015] Global significance analysis is performed on the initial multimodal data based on the local event saliency parameters and the time mapping function to obtain global event saliency parameters. The global event saliency parameters are used to quantify the contribution of each modality in the initial multimodal data to the saliency parameters of the event nodes in the time and spatial dimensions.
[0016] Based on the global event significance parameter, the event center node is selected from the event nodes. The event center node is the time node of a significant event whose global event significance parameter is not lower than the significance threshold.
[0017] Based on the time mapping function and the event center node, the contribution weight analysis of the initial multimodal data is performed to generate a soft alignment weight kernel function. The soft alignment weight kernel function is used to define the contribution weight of each modality data in the initial multimodal data to the event center node at each time node, so as to perform spatiotemporal alignment of the initial multimodal data.
[0018] The initial multimodal data is filtered for event relevance using the soft-aligned weight kernel function to obtain candidate multimodal data.
[0019] Optionally, the step of performing global saliency analysis on the initial multimodal data based on the local event saliency parameters and the time mapping function to obtain global event saliency parameters includes:
[0020] Signal quality analysis is performed on each modal data in the initial multimodal data to obtain the signal reliability parameters of the initial multimodal data;
[0021] A weight adjustment function is constructed based on the aforementioned signal confidence parameters. The mathematical expression of the weight adjustment function is as follows:
[0022]
[0023] in, Representing modes In time The reliability of the signal. Indicates modal index, Represents modal weights, This represents the modal index variable traversed during the summation process. Indicates all modes in time The sum of credibility at any given moment;
[0024] Global significance analysis is performed on the initial multimodal data based on the weight adjustment function, the local event significance parameters, and the time mapping function to obtain global event significance parameters, which are calculated using the following formula:
[0025]
[0026] in, Indicates a significant parameter for global events. This represents the modality set of the initial multimodal data. Representing modes Local event saliency parameters, Representing modes The inverse function of the time mapping function, Indicates a spatial index.
[0027] Optionally, the step of performing contribution weight analysis on the initial multimodal data based on the time mapping function and the event center node to generate a soft-aligned weight kernel function includes:
[0028] The response scale of the soft alignment kernel is determined based on the global event saliency parameters and the baseline bandwidth, and the response scale is calculated based on the following formula:
[0029]
[0030] in, This represents the response scale, which is used to control the degree of time spread of the kernel function. This represents the reference bandwidth, which is a modal bandwidth. Initial kernel width without event adjustment Indicates the significance of adjusting hyperparameters. Indicates a significant parameter for global events. Indicates the central node of the event;
[0031] Based on the response scale, the event center node, and the event mapping function, contribution weight analysis is performed on the initial multimodal data to generate a soft-aligned weight kernel function. The mathematical expression of the soft-aligned weight kernel function is as follows:
[0032]
[0033] in, Representing modes In time Corresponding event center Soft alignment weight kernel, Representing modes The time mapping function.
[0034] Optionally, the time mapping function includes:
[0035]
[0036]
[0037]
[0038] in, Representing modes The time mapping function, Representing modes The learnable parameters of the time-transformation network, Representing modes The local time scaling factor at time t is used to correct for differences in response speed between modes. Representing modes The local time offset is used to compensate for the time offset between different modes. Indicates the first intermediate variable. Indicates the second intermediate variable. The initial time offset constant term is represented by the first intermediate variable and the second intermediate variable, which are obtained by inputting modal data into the output of a differentiable time-series network.
[0039] Optionally, the step of performing attention feature fusion on the feature embedding vectors of each modality data to obtain multimodal fusion features includes:
[0040] The feature embedding vectors of each modality data are used as graph nodes, and a node set is constructed based on multiple graph nodes;
[0041] Event correlation analysis is performed on the feature embedding vectors of each modality data, and a dynamic edge weight function is constructed based on the analysis results:
[0042]
[0043] in, Indicates the time node of the event At that time, mode With mode Dynamic edge weights between them Indicates the time point of the event. Indicates the time decay coefficient. The exponential decay term representing the time difference, Representing modes With mode The time difference between corresponding events and Representing modes and modality At the point of time The feature vector at that location, Indicates the time node of the event Importance weight, Representing modes and modality At the point of time Cosine similarity at the location;
[0044] Determine attention scores between different modalities based on graph attention networks;
[0045] A dual-weighted attention coefficient fusion weight function is constructed based on the attention score, and the target attention weight is determined based on the dual-weighted attention coefficient fusion weight function. The mathematical expression of the dual-weighted attention coefficient fusion weight function is as follows:
[0046]
[0047] in, Represents a node For nodes Target attention weights Indicates the weight of content similarity. and This represents the attention score. Represents a node Uncertainty parameters, Represents a node The set of neighboring nodes, Represents a regular term, Indicates the neighbor node index. Represents a node Corresponding mode With nodes Corresponding mode Association weights, Represents a node Corresponding mode with neighboring nodes Corresponding mode The association weight;
[0048] Based on the dual-weighted attention coefficient fusion weights, attention feature fusion is performed on the feature embedding vectors of each modality data to obtain multimodal fusion features.
[0049] Optionally, the step of identifying and monitoring the solid waste category in the solid waste disposal area based on the multimodal fusion features includes:
[0050] The multimodal fusion features are averaged and pooled to obtain global fusion features;
[0051] The global fusion features are linearly mapped and activated to obtain intermediate features;
[0052] The intermediate features are input into the Softmax classifier, which outputs the classification prediction probability of various types of waste. Based on the classification prediction probability, the solid waste disposal area is used to identify and monitor the solid waste category.
[0053] Furthermore, to achieve the above objectives, this invention also proposes a solid waste intelligent monitoring system based on multi-source data. This system is used to implement the aforementioned solid waste intelligent monitoring method based on multi-source data. The solid waste intelligent monitoring system based on multi-source data includes:
[0054] The data monitoring module is used to monitor the solid waste disposal area from multiple sources and obtain raw multimodal data. The raw multimodal data is collected by multimodal sensors that are pre-set in the solid waste disposal area.
[0055] The data processing module is used to preprocess the raw multimodal data to obtain initial multimodal data;
[0056] The time mapping module is used to map the local sampling time of each modality in the initial multimodal data to the event reference time axis and generate a time mapping function. The event reference time axis includes multiple event nodes, which are generated based on the time nodes of related events of solid waste.
[0057] The spatiotemporal alignment module is used to perform spatiotemporal alignment on the initial multimodal data based on the time mapping function and the spatial index of the initial multimodal data to obtain candidate multimodal data;
[0058] The feature extraction module is used to perform cross-modal feature extraction on the candidate multimodal data to obtain the feature embedding vector of each modality data;
[0059] The feature fusion module is used to perform attention feature fusion on the feature embedding vectors of each modality data to obtain multimodal fusion features;
[0060] The identification and monitoring module is used to identify and monitor the solid waste category in the solid waste disposal area based on the multimodal fusion features.
[0061] Furthermore, to achieve the above objectives, this application also proposes a solid waste intelligent monitoring device based on multi-source data. The device includes: a memory, a processor, and a solid waste intelligent monitoring program based on multi-source data stored in the memory. The processor is used to run the solid waste intelligent monitoring program based on multi-source data, and the computer program is configured to implement the steps of the solid waste intelligent monitoring method based on multi-source data as described above.
[0062] In addition, to achieve the above objectives, this application also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the intelligent solid waste monitoring method based on multi-source data as described above.
[0063] This invention obtains raw multimodal data by monitoring solid waste disposal areas using multi-source data. This raw multimodal data is collected by multimodal sensors pre-installed in the solid waste disposal areas. The raw multimodal data is preprocessed to obtain initial multimodal data. The local sampling time of each modality in the initial multimodal data is mapped to an event reference time axis to generate a time mapping function. The event reference time axis includes multiple event nodes, which are generated based on the time nodes of related events of solid waste. The initial multimodal data is spatiotemporally aligned based on the time mapping function and the spatial index of the initial multimodal data to obtain candidate multimodal data. Cross-modal feature extraction is then performed on the candidate multimodal data. The invention obtains feature embedding vectors for each modality of data, performs attention feature fusion on these vectors to obtain multimodal fusion features, and uses these features to monitor and identify solid waste categories in the solid waste disposal area. Because this invention maps multimodal data to an event reference time axis, it effectively compensates for sampling differences between multimodal sensors, achieving cross-modal event synchronization. By aligning multimodal data in time and space, it can accurately capture multi-source information of the same event under different sampling rates and response delays, significantly improving the accuracy and robustness of cross-modal fusion. By focusing on significant moments centered on events, it reduces redundant data processing, enhances adaptability to environmental changes, and improves the accuracy of solid waste identification. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 This is a schematic diagram of the structure of a multi-source data-based intelligent solid waste monitoring device in the hardware operating environment of the embodiment of the present invention;
[0066] Figure 2 This is a flowchart illustrating the first embodiment of the intelligent solid waste monitoring method based on multi-source data of the present invention.
[0067] Figure 3 This is a flowchart illustrating the second embodiment of the intelligent solid waste monitoring method based on multi-source data of the present invention.
[0068] Figure 4 This is a structural block diagram of the first embodiment of the intelligent solid waste monitoring system based on multi-source data of the present invention.
[0069] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0070] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0071] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a multi-source data-based intelligent solid waste monitoring device in the hardware operating environment of an embodiment of the present invention.
[0072] like Figure 1 As shown, the intelligent solid waste monitoring device based on multi-source data may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk storage device. Optionally, the memory 1005 may also be a storage system independent of the aforementioned processor 1001.
[0073] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on intelligent monitoring devices for solid waste based on multi-source data, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0074] like Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a solid waste intelligent monitoring program based on multi-source data.
[0075] exist Figure 1In the multi-source data-based intelligent solid waste monitoring device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and memory 1005 in the multi-source data-based intelligent solid waste monitoring device of the present invention can be set in the multi-source data-based intelligent solid waste monitoring device. The multi-source data-based intelligent solid waste monitoring device calls the multi-source data-based intelligent solid waste monitoring program stored in the memory 1005 through the processor 1001 and executes the multi-source data-based intelligent solid waste monitoring method provided in the embodiment of the present invention.
[0076] This invention provides a method for intelligent monitoring of solid waste based on multi-source data, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the intelligent solid waste monitoring method based on multi-source data of the present invention.
[0077] In this embodiment, the intelligent monitoring method for solid waste based on multi-source data includes the following steps:
[0078] Step S10: Conduct multi-source data monitoring on the solid waste disposal area to obtain raw multimodal data.
[0079] It should be noted that this embodiment is applied to the classification and identification of solid waste in the environment. By mapping multimodal data to an event reference time axis, it effectively compensates for the sampling differences between multimodal sensors, realizes cross-modal event synchronization, and accurately captures multi-source information of the same event under different sampling rates and response delays through spatiotemporal alignment of multimodal data. This significantly improves the accuracy and robustness of cross-modal fusion, focuses on significant moments centered on events, reduces redundant data processing, enhances adaptability to environmental changes, and improves the identification accuracy of solid waste.
[0080] It should be understood that the executing entity of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a terminal electronic device capable of performing the above functions. The following description uses a multi-source data-based intelligent solid waste monitoring device (hereinafter referred to as the monitoring device) as an example to illustrate this embodiment and the following embodiments.
[0081] It should be noted that the raw multimodal data can be an unprocessed raw data set collected by multimodal sensors, containing information in dimensions such as vision, acoustics, gas, and space. The raw multimodal data is collected by multimodal sensors pre-installed in the solid waste disposal area.
[0082] It should be noted that solid waste disposal areas can be places used for disposing of solid waste, such as residential community garbage bins or commercial waste stations. Multimodal sensors can be sensor networks composed of multiple sensors with different modes, and these sensor networks can be combinations of sensors capable of collecting two or more different types of data (such as devices that simultaneously collect visual and acoustic data).
[0083] In practice, the monitoring equipment can be pre-positioned at the solid waste disposal point with cameras, acoustic sensors, and gas (VOC, CO, CO2, methane) sensors to collect relevant sensor information and GPS geographic information, and perform preprocessing operations on different modal data.
[0084] Step S20: Preprocess the original multimodal data to obtain initial multimodal data.
[0085] It should be noted that the initial multimodal data can be the original multimodal data after preprocessing (such as data denoising and normalization), which is used for subsequent spatiotemporal alignment and feature extraction.
[0086] In some embodiments, the monitoring equipment performs targeted preprocessing on data from different modalities. Visual data undergoes denoising, resolution adjustment, and keyframe extraction; acoustic data undergoes filtering, noise reduction, and sampling rate standardization; gas data undergoes outlier removal and smoothing; and spatial data undergoes format standardization. The preprocessed data from each modality are integrated to obtain initial multimodal data.
[0087] Step S30: Map the local sampling time of each modality data in the initial multimodal data to the event reference time axis to generate a time mapping function.
[0088] It should be noted that the event reference timeline includes multiple event nodes, which are generated based on the time nodes of related events concerning solid waste. These related events may include events such as the appearance, placement, and fall of solid waste. The event reference timeline is used to align the time dimension of multimodal data, with each event node corresponding to the time of a related event concerning solid waste.
[0089] It should be noted that the local sampling time can be the timestamp of the data collected by each modal sensor, and there are differences in time bases between different modalities. The time mapping function can be a mathematical function that converts the local sampling time of each modality to the event reference time axis, and has learnable parameters to adapt to the time differences of the sensors.
[0090] It is understood that this embodiment solves the problem of time asynchrony caused by differences in sensor hardware in multimodal data by pre-constructing an event reference time axis based on event time nodes and mapping the local sampling time of each modality data to the event reference time axis. This lays the foundation for subsequent spatiotemporal alignment and event-level analysis, and ensures the consistency of the time dimension of multimodal data.
[0091] In some embodiments, the monitoring device analyzes the local sampling time characteristics of each modality's data and constructs a differentiable time mapping function to transform the local time of each modality (such as camera frame time and acoustic sensor sampling time) to a unified event reference time axis. The event nodes of the event reference time axis are generated by identifying key time points (such as the start time and peak time of the disposal action) of related events such as solid waste disposal and impact.
[0092] Furthermore, to accurately compensate for sensor delay, response difference, clock skew, and local rate variations, in one embodiment, the time mapping function includes:
[0093]
[0094] in, Representing modes The time mapping function, that is, the original time of the mode. Mapped to a unified "aligned timeline";
[0095] Representing modes Learnable parameters of time-transformed networks;
[0096] Representing modes The local time scaling factor at time t is used to correct the difference in response speed between modes. If it is greater than 1, it means that the time axis of the mode is stretched and the event occurs "faster". If it is between 0 and 1, it means that the time axis is compressed and the event is "slowed down".
[0097] Representing modes The local time offset, i.e. the translation amount at the alignment time, can compensate for problems such as transmission delay and asynchronous sampling start between different modes.
[0098] To ensure numerical stability, let:
[0099]
[0100]
[0101] in, This represents the first intermediate variable, which is a learnable intermediate variable (an unconstrained real number).
[0102] Indicates the second intermediate variable. Both are intermediate variables, representing the derivative of the time offset, that is, the rate of change of the offset per unit time;
[0103] This represents the initial time offset constant, which is usually set to 0, or initialized based on the known clock difference of the sensor.
[0104] The first and second intermediate variables are obtained by inputting modal data to the output of a differentiable temporal network, for example, , The output of a differentiable temporal network (such as a Transformer) has the following input mode: The output mode is , .
[0105] It is understandable that this embodiment designs a differentiable and learnable time mapping. each mode The local sampling time is mapped to a unified event reference time axis, enabling cross-modal comparison and aggregation at the "event" level, while allowing compensation for sensor delay, response difference, clock skew and local rate variation.
[0106] Step S40: Based on the time mapping function and the spatial index of the initial multimodal data, perform spatiotemporal alignment on the initial multimodal data to obtain candidate multimodal data.
[0107] It should be noted that spatial indexes are information identifying the spatial location of data (such as image pixel coordinates, geographic latitude and longitude), used for spatial alignment of multimodal data. Candidate multimodal data can be a subset of multimodal data that has been spatiotemporally aligned and is highly correlated in terms of each modality, and can be used for subsequent feature extraction and fusion.
[0108] In some embodiments, the monitoring device uses a time mapping function to align the time of each modality of data to an event reference time axis. Simultaneously, it combines the spatial index of the initial multimodal data (such as visual data pixel coordinates and GPS geographic coordinates) to perform spatial dimension matching and calibration. Multimodal data corresponding to the same event in both spatiotemporal dimensions are selected to obtain candidate multimodal data. This achieves accurate matching of multimodal data in the spatiotemporal dimension, ensuring that subsequent feature extraction and fusion utilize multi-source information of the same event, improving the accuracy of multimodal data collaborative analysis, and avoiding event misjudgment caused by spatiotemporal misalignment.
[0109] Step S50: Perform cross-modal feature extraction on the candidate multimodal data to obtain the feature embedding vector of each modality data.
[0110] It should be noted that the feature embedding vector can include features of different modalities corresponding to multiple modal data. For example, the feature embedding vector can include visual modal features, acoustic modal features, gas modal features, and geographic information modal features.
[0111] In its specific implementation, the monitoring equipment performs cross-modal feature extraction on candidate multimodal data, which may include:
[0112] 1. Visual modality (camera) feature extraction:
[0113] Feature extraction is performed using the Visual Transformer technique, outputting the visual modal features at each time step t. .
[0114] The features extracted from the visual modality include:
[0115] Spatial characteristics: appearance, shape, color, texture, outline, etc. of waste.
[0116] Temporal characteristics: garbage movement, throwing behavior, and changes in accumulation.
[0117] 2. Acoustic modal (acoustic sensor) feature extraction:
[0118] First, a Mel-frequency transform is performed on the acoustic signal to obtain a time-frequency spectrum. Then, a 2D convolutional neural network (such as VGGish) is used to extract the local energy distribution and patterns in the spectrum, outputting the results for each time step. acoustic modal characteristics .
[0119] The features extracted from acoustic modes include:
[0120] Acoustic event characteristics: such as specific audio events like bottle breakage, metal impact, paper friction, and liquid spillage.
[0121] Energy characteristics: signal strength, noise level, instantaneous sound pressure level.
[0122] Spectral characteristics: The ratio of low-frequency to high-frequency energy can reflect different waste disposal behaviors.
[0123] 3. Gas mode (each gas sensor) feature extraction:
[0124] The gas concentration variation over time was extracted using a multilayer perceptron (MLP), and the multi-channel time series was processed using a Transformer Encoder to learn the correlation between different gas types, outputting the results for each time step. Gas modal characteristics .
[0125] The features extracted from gas modes include:
[0126] Gas concentration characteristics: instantaneous values and trends of VOC, CO, CO2, etc.
[0127] Dynamic change characteristics: short-term gas fluctuation rate and concentration change rate.
[0128] Environmental indicators: odor intensity, presence of burning / decomposition, etc.
[0129] 4. GPS / Geographic Information Modal Feature Extraction:
[0130] A multilayer perceptron (MLP) is used to perform a nonlinear mapping of latitude, longitude, azimuth, altitude, and region type (industrial area, residential area, etc.) to output a geographic feature vector. .
[0131] The features extracted from GPS / geographic information modal analysis include:
[0132] Spatial location information: node coordinates, relative position.
[0133] Environmental semantics: the distribution characteristics of waste types in different geographical areas (e.g., high proportion of kitchen waste in residential areas and high proportion of waste in industrial areas).
[0134] Step S60: Perform attention feature fusion on the feature embedding vectors of each modality data to obtain multimodal fusion features.
[0135] It should be noted that multimodal fusion features are comprehensive feature representations obtained by fusing multimodal feature embedding vectors. They contain key information from each modality and have stronger semantic expressive power.
[0136] It is understood that this embodiment can dynamically calculate the importance weights of each modality feature by utilizing an attention mechanism, and then fuse multimodal features according to the weights; by adaptively highlighting modality features that are important to the current task through the attention mechanism and suppressing irrelevant and redundant features, it can achieve efficient fusion of multimodal features, improve the semantic consistency and discriminativeness of features, and provide more discriminative input for subsequent classification.
[0137] Furthermore, in order to accurately fuse multimodal features and avoid ignoring temporal dependencies and spatial dynamic relationships, step S60 above may include:
[0138] Step S601: Use the feature embedding vectors of each modality data as graph nodes, and construct a node set based on multiple graph nodes.
[0139] Understandably, after modal feature extraction, the monitoring device obtains a unified embedding of each modality aligned to the time axis. Traditional methods typically employ concatenation or simple weighted summation for fusion, neglecting temporal dependencies and spatial dynamics. To address this issue, this embodiment constructs an event-driven spatiotemporal heterogeneous graph attention network (EST-HGAT) that simultaneously models temporal dynamics and cross-modal correlations, achieving spatiotemporal graph construction and cross-modal attention fusion.
[0140] In practical implementation, the monitoring equipment constructs spatiotemporal nodes, and assigns each modality... At the event node The embedding at a point is considered as a node in the graph:
[0141]
[0142] in, No. The event point in the modality The corresponding graph node below; Indicates the time point of the event Location, mode Embedded feature vectors; For the first The timestamp of each event.
[0143] Therefore, the node set is defined as follows:
[0144]
[0145] The edge set consists of two parts: temporal edges, which represent the temporal relationship between adjacent events; and modal edges, which represent the association between different modalities.
[0146] Step S602: Perform event correlation analysis on the feature embedding vectors of each modality data, and construct a dynamic edge weight function based on the analysis results.
[0147] Understandably, a dynamic edge weight function is constructed to reflect the relationship between event intensity and latency:
[0148]
[0149] in, Indicates the time node of the event At that time, mode With mode The dynamic edge weights between them reflect the strength of their correlation;
[0150] Indicates the time point of the event;
[0151] This represents the time decay coefficient, which controls the rate at which the time difference decays with respect to the weights.
[0152] The exponential decay term representing the time difference;
[0153] Representing modes With mode The time difference between corresponding events indicates a weaker correlation if the sampling of the two modalities is not synchronized.
[0154] and Representing modes and modality At the point of time The eigenvector at that location;
[0155] Indicates the time node of the event The importance weight is used to amplify the impact of key events on the graph structure;
[0156] Representing modes and modality At the point of time The cosine similarity at a given location indicates the higher the similarity, the more correlated the modes are.
[0157] The importance of an event is considered as a weighted combination of several signals: signal strength, cross-modal consistency, etc. First, an unnormalized score is calculated. Then normalize to obtain the final weights. .
[0158]
[0159] Each of them It is a characteristic component (such as signal strength component, cross-modal consistency component, scaled to the same dimension). For learnable / configurable weights; It is the number of components used.
[0160] Step S603: Determine the attention scores between each modality based on the graph attention network;
[0161] Step S604: Construct a dual-weighted attention coefficient fusion weight function based on the attention score, and determine the target attention weight based on the dual-weighted attention coefficient fusion weight function.
[0162] It is understandable that cross-modal attention fusion includes:
[0163] First, construct an attention scoring system based on the GAT format:
[0164]
[0165] in, It is a non-linear activation function; Learnable parameter vector; It is a learnable weight matrix; This involves concatenating vectors. For nodes feature.
[0166] Secondly, a dual-weighted attention coefficient fusion weight is constructed. This dual-weighted mechanism ensures that nodes with high confidence and important events gain greater influence in graph propagation.
[0167]
[0168] in, Represents a node For nodes Target attention weights Indicates the weight of content similarity. and This represents the attention score. Represents a node The uncertainty parameter is measured by the Bayesian approximation. Represents a node The set of neighboring nodes, Represents a regular term, To prevent division by zero, Indicates the neighbor node index. Represents a node Corresponding mode With nodes Corresponding mode Association weights, Represents a node Corresponding mode with neighboring nodes Corresponding mode The association weight;
[0169] Step S605: Based on the dual-weighted attention coefficient fusion weight, perform attention feature fusion on the feature embedding vectors of each modality data to obtain multimodal fusion features.
[0170] It should be noted that the above The weighting formula shows that, It is the content similarity weight. It is the structure / event weight. It is a reliability weight (the more uncertain, the smaller the contribution).
[0171] Finally, the node is updated as follows:
[0172]
[0173] in, This represents the multimodal fusion feature.
[0174] It is understood that this embodiment introduces an event-driven spatiotemporal heterogeneous graph attention network (EST-HGAT), models time dependence and intermodal correlations, and utilizes a dual-weight attention mechanism (structural weight and reliability weight) to enable the system to capture semantically reinforcing correlation signals during multimodal feature fusion, thereby improving the overall recognition accuracy.
[0175] Step S70: Based on the multimodal fusion features, perform solid waste category identification and monitoring on the solid waste disposal area.
[0176] In practical implementation, monitoring equipment can input multimodal fusion features into the classification model (such as...) The classifier outputs the probability distribution of various types of solid waste, determines the waste category based on the highest probability, and realizes real-time category identification and monitoring of solid waste disposal areas (such as distinguishing recyclables, kitchen waste, hazardous waste, etc.), thereby providing data support for waste classification management, resource recycling and environmental governance, and improving the level of intelligence in solid waste management.
[0177] Furthermore, in order to accurately identify the category of solid waste, step S70 above may include:
[0178] Step S701: Perform average pooling on the multimodal fusion features to obtain global fusion features;
[0179] Step S702: Perform linear mapping and activation processing on the global fusion features to obtain intermediate features;
[0180] Step S703: Input the intermediate features to The classifier outputs the classification prediction probability of various types of waste, and performs solid waste category identification and monitoring in the solid waste disposal area based on the classification prediction probability.
[0181] In its implementation, to achieve real-time waste category identification, this embodiment designs a lightweight classification head, including the following steps:
[0182] 1. Global aggregation:
[0183] Average pooling is applied to the event dimension to obtain the global fusion feature at the current moment, as shown in the following formula:
[0184]
[0185] in, This is for weighted average pooling.
[0186] 2. Linear mapping and nonlinear transformation, refer to the following formulas:
[0187] The global fusion features are linearly mapped and activated to obtain intermediate features:
[0188]
[0189] in, It is a linear mapping matrix; For bias terms; This is the activation function.
[0190] 3. Classified output:
[0191] Use a single layer The classifier outputs the probability distribution of each waste category, as shown in the following formula:
[0192]
[0193] in, This represents the predicted probability of various types of waste classification, with the category corresponding to the highest probability being the system's real-time identification result; This is the classification layer weight matrix; For classification layer bias terms; The function maps linear outputs to probability distributions for each category.
[0194] This embodiment obtains raw multimodal data by monitoring the solid waste disposal area using multi-source data. This raw multimodal data is collected by multimodal sensors pre-installed in the solid waste disposal area. The raw multimodal data is preprocessed to obtain initial multimodal data. The local sampling time of each modality in the initial multimodal data is mapped to an event reference time axis to generate a time mapping function. The event reference time axis includes multiple event nodes, which are generated based on the time nodes of related events of solid waste. The initial multimodal data is spatiotemporally aligned based on the time mapping function and the spatial index of the initial multimodal data to obtain candidate multimodal data. Cross-modal feature extraction is then performed on the candidate multimodal data. The method obtains feature embedding vectors for each modality of data, performs attention feature fusion on the feature embedding vectors of each modality of data to obtain multimodal fusion features, and performs solid waste category identification and monitoring on the solid waste disposal area based on the multimodal fusion features. Since this embodiment maps multimodal data to an event reference time axis, it effectively compensates for the sampling differences between multimodal sensors, realizes cross-modal event synchronization, and accurately captures multi-source information of the same event under different sampling rates and different response delays by performing spatiotemporal alignment on multimodal data. This significantly improves the accuracy and robustness of cross-modal fusion, focuses on significant moments centered on events, reduces redundant data processing, enhances adaptability to environmental changes, and improves the identification accuracy of solid waste.
[0195] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the intelligent solid waste monitoring method based on multi-source data of the present invention.
[0196] Based on the first embodiment described above, in this embodiment, step S40 further includes:
[0197] Step S401: Perform spatial dimension saliency analysis based on the spatial index of the initial multimodal data to obtain local event saliency parameters.
[0198] It should be noted that spatial significance analysis can assess the significance of regions or locations in multimodal data related to "solid waste-related events" (such as disposal or falling) from a spatial perspective. Local event significance parameters are parameters that quantify the significance of a single modality's spatial association with the event (such as target confidence, sound pressure level, and concentration gradient).
[0199] Step S402: Perform global significance analysis on the initial multimodal data based on the local event saliency parameters and the time mapping function to obtain global event saliency parameters.
[0200] It should be noted that the global event saliency parameter is used to quantify the contribution of each modality in the initial multimodal data to the saliency parameters of the event node in the time and spatial dimensions.
[0201] Furthermore, in order to accurately analyze the global event saliency of multimodal data, step S402 above may include:
[0202] Step S4021: Perform signal quality analysis on each modal data in the initial multimodal data to obtain the signal reliability parameters of the initial multimodal data;
[0203] Step S4022: Construct a weight adjustment function based on the signal confidence parameters;
[0204] Step S4023: Perform global significance analysis on the initial multimodal data based on the weight adjustment function, the local event significance parameters, and the time mapping function to obtain global event significance parameters.
[0205] Understandably, this section aims to jointly compute event saliency from multimodal time-series signals. And determine the event center through saliency peak or threshold detection. The core idea is different modalities. Events have different levels of importance, therefore variable weights are used. Adjusting the saliency of each mode The contributions of each factor form a weighted joint significance. , for soft alignment kernel Provide the basis for the weighting. Specifically:
[0206]
[0207]
[0208] in, This is the current time point, used to calculate the instantaneous significance of each mode; For spatial indexing, if the input is a spatiotemporal signal (such as a video frame or spatial distribution signal), then Representing spatial coordinates ; For modality exist The saliency of a local event is the saliency of the local event obtained from the detectors, feature energies, or network outputs of each modality. This is the modal-time inverse mapping function, which unifies time. Mapping back to mode Its own timeline is used for time alignment or compensation; Modality weights control the influence of each modality in saliency fusion; For modality A reliability metric that measures the mode's reliability over time. Signal quality or reliability, such as signal-to-noise ratio, detection confidence, etc.; Joint significance, which is the global significance score obtained by weighted summation of the significance scores of each modality. This represents the modal index variable traversed during the summation process. Indicates all modes in time The sum of credibility at any given moment.
[0209] The above formula illustrates that the saliency of each modality is weighted and fused after time alignment, with the weights... By reliability Normalization ensures that modes with high reliability account for a larger proportion. The final result is... This represents the joint saliency of the multimodal expressions at the current moment.
[0210] Step S403: Select the event center node from the event nodes based on the global event saliency parameters.
[0211] It should be noted that the event center node is the time node of a significant event whose global event significance parameter is not lower than the significance threshold.
[0212] It is understood that this embodiment aims to jointly calculate event saliency from multimodal time-series signals. And determine the event center through saliency peak or threshold detection. The core idea is different modalities. Events have different levels of importance, therefore variable weights are used. Adjusting the saliency of each mode The contributions of each factor form a weighted joint significance. , for soft alignment kernel Provide weighting criteria.
[0213] It should be understood that event center detection can be:
[0214]
[0215] in, A significance threshold is used to determine the occurrence of an event when the threshold value is exceeded. The event center is determined by using the significance curve to identify the significance threshold, resulting in a series of... This refers to the moment of significant events.
[0216] Step S404: Perform contribution weight analysis on the initial multimodal data based on the time mapping function and the event center node to generate a soft-aligned weight kernel function.
[0217] It should be noted that the soft alignment weight kernel function is used to define the contribution weight of each modality in the initial multimodal data to the event center node at each time node, so as to perform spatiotemporal alignment of the initial multimodal data.
[0218] It is understood that this embodiment dynamically adjusts the weights of each mode through reliability metrics. When a certain mode signal is distorted or missing, the system can automatically reduce its impact and strengthen the contribution of other modes, thereby enhancing its adaptability to environmental changes (such as noise, obstruction, and gas interference).
[0219] Furthermore, in order to achieve accurate alignment of event-level multimodal data, step S404 above may include:
[0220] Step S4041: Determine the response scale of the soft alignment kernel based on the global event saliency parameters and the reference bandwidth;
[0221] Step S4042: Perform contribution weight analysis on the initial multimodal data based on the response scale, the event center node, and the event mapping function to generate a soft-aligned weight kernel function.
[0222] It is understandable that the modalities are mapped over time. On the aligned timeline, it is viewed as a series of "events," with each event as the center. Define the modality at each alignment target. In time The above corresponds to the event center. Soft alignment weight kernel :
[0223]
[0224] in, Indicates the event center The "response weight" at that moment; The width of the soft alignment kernel (response scale) controls the degree of time spread of the kernel function, determining the width of the "alignment window." A larger value indicates a greater tolerance for time offsets. for:
[0225]
[0226] in, This represents the response scale, which is used to control the degree of time spread of the kernel function. This represents the reference bandwidth, which is a modal bandwidth. Initial kernel width without event adjustment This represents a significance-adjusting hyperparameter used to control significance. The strength of the effect on nuclear width Indicates a significant parameter for global events. Indicates the central node of the event. For the salience of joint events.
[0227] The above formula constructs a dynamically scalable, event-centric soft time window. When significant events occur, the kernel automatically tightens for precise capture; during noisy or ambiguous events, the kernel widens to avoid false matches. (This is related to the time mapping function.) The combination enables robust alignment across multimodal or time series with different sampling rates.
[0228] Step S405: Perform event correlation screening on the initial multimodal data using the soft alignment weight kernel function to obtain candidate multimodal data.
[0229] It should be understood that the soft alignment kernel is used for dynamic filtering of time and event correlations. It filters the event correlation of the original multimodal signals in the time domain, so that the system retains only the segments that are highly correlated with the event center, thereby achieving accurate alignment of event-level multimodal data.
[0230] The soft alignment weight kernel function is essentially a dynamic Gaussian weight function. Its function is to adjust the weighting of the timeline, after initial alignment via time mapping, around the center of the garbage disposal event, to form a modal weighting function. At any time The signal allocation contribution weight is adjusted so that the weight is higher for video frames closer to key time points (such as video frames when the delivery action occurs, impact sound) and lower for time points further away from noise (such as background images and ambient sounds when there is no delivery), thereby achieving "flexible and precise event focusing" and avoiding the rigidity and mismatch problems of traditional "hard window alignment" (fixed time range interception).
[0231] Understandably, this embodiment achieves spatiotemporal alignment of multimodal sensors (vision, acoustics, gas, and geographic information) by introducing a differentiable temporal mapping and an event-driven soft alignment kernel. This enables the system to accurately capture multi-source information of the same event under different sampling rates and response delays, significantly improving the accuracy and robustness of cross-modal fusion. Event-driven feature extraction is more targeted. Feature extraction and alignment are centered on "events," allowing feature learning to automatically focus on salient moments (such as solid waste disposal or falling) from the global time series, reducing redundant information interference and improving the model's discrimination efficiency and real-time performance.
[0232] This embodiment performs spatial saliency analysis based on the spatial index of the initial multimodal data to obtain local event saliency parameters. Based on these local event saliency parameters and the time mapping function, it performs global saliency analysis on the initial multimodal data to obtain global event saliency parameters. These global event saliency parameters quantify the contribution of each modality in the initial multimodal data to the saliency parameters of event nodes in both the time and spatial dimensions. Based on these global event saliency parameters, it selects event center nodes from the event nodes. The event center nodes are the time nodes of saliency events whose global event saliency parameters are not lower than a saliency threshold. Based on the time mapping function and the event center nodes, it performs contribution weight analysis on the initial multimodal data to generate a soft-aligned weight kernel function. The soft alignment weight kernel function is used to define the contribution weight of each modality data in the initial multimodal data to the event center node at each time node, so as to perform spatiotemporal alignment of the initial multimodal data. The soft alignment weight kernel function is used to filter the event relevance of the initial multimodal data to obtain candidate multimodal data. This implementation realizes the spatiotemporal alignment of multimodal sensors by introducing differentiable time mapping and event-driven soft alignment kernel. It can accurately capture multi-source information of the same event under different sampling rates and different response delays, which significantly improves the accuracy and robustness of cross-modal fusion. Feature extraction and alignment are performed with the event as the center, so that feature learning automatically focuses on significant event moments from the global time series, reduces redundant information interference, and improves the discrimination efficiency and real-time performance of solid waste.
[0233] Furthermore, this embodiment of the invention also proposes a computer-readable storage medium storing a solid waste intelligent monitoring program based on multi-source data. When the solid waste intelligent monitoring program based on multi-source data is executed by a processor, it implements the steps of the solid waste intelligent monitoring method based on multi-source data as described above.
[0234] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0235] The aforementioned computer-readable storage medium may be included in a multi-source data-based intelligent solid waste monitoring device; or it may exist independently and not be assembled into a multi-source data-based intelligent solid waste monitoring device.
[0236] Furthermore, this invention also proposes a computer program product, including a solid waste intelligent monitoring program based on multi-source data. When the solid waste intelligent monitoring program based on multi-source data is executed by a processor, it implements the steps of the solid waste intelligent monitoring method based on multi-source data as described above.
[0237] The specific implementation of the computer program product of the present invention is basically the same as the embodiments of the above-mentioned intelligent monitoring method for solid waste based on multi-source data, and will not be repeated here.
[0238] Reference Figure 4 , Figure 4 This is a structural block diagram of the first embodiment of the intelligent solid waste monitoring system based on multi-source data of the present invention.
[0239] like Figure 4 As shown in the embodiments of the present invention, the intelligent solid waste monitoring system based on multi-source data includes:
[0240] Data monitoring module 10 is used to perform multi-source data monitoring on the solid waste disposal area to obtain raw multimodal data. The raw multimodal data is collected by multimodal sensors pre-set in the solid waste disposal area.
[0241] Data processing module 20 is used to preprocess the raw multimodal data to obtain initial multimodal data;
[0242] The time mapping module 30 is used to map the local sampling time of each modality data in the initial multimodal data to the event reference time axis and generate a time mapping function. The event reference time axis includes multiple event nodes, which are generated based on the time nodes of the associated events of solid waste.
[0243] The spatiotemporal alignment module 40 is used to perform spatiotemporal alignment on the initial multimodal data based on the time mapping function and the spatial index of the initial multimodal data to obtain candidate multimodal data;
[0244] Feature extraction module 50 is used to perform cross-modal feature extraction on the candidate multimodal data to obtain feature embedding vectors for each modality data;
[0245] The feature fusion module 60 is used to perform attention feature fusion on the feature embedding vectors of each modality data to obtain multimodal fusion features;
[0246] The identification and monitoring module 70 is used to identify and monitor the solid waste category in the solid waste disposal area based on the multimodal fusion features.
[0247] This embodiment obtains raw multimodal data by monitoring the solid waste disposal area using multi-source data. This raw multimodal data is collected by multimodal sensors pre-installed in the solid waste disposal area. The raw multimodal data is preprocessed to obtain initial multimodal data. The local sampling time of each modality in the initial multimodal data is mapped to an event reference time axis to generate a time mapping function. The event reference time axis includes multiple event nodes, which are generated based on the time nodes of related events of solid waste. The initial multimodal data is spatiotemporally aligned based on the time mapping function and the spatial index of the initial multimodal data to obtain candidate multimodal data. Cross-modal feature extraction is then performed on the candidate multimodal data. The method obtains feature embedding vectors for each modality of data, performs attention feature fusion on the feature embedding vectors of each modality of data to obtain multimodal fusion features, and performs solid waste category identification and monitoring on the solid waste disposal area based on the multimodal fusion features. Since this embodiment maps multimodal data to an event reference time axis, it effectively compensates for the sampling differences between multimodal sensors, realizes cross-modal event synchronization, and accurately captures multi-source information of the same event under different sampling rates and different response delays by performing spatiotemporal alignment on multimodal data. This significantly improves the accuracy and robustness of cross-modal fusion, focuses on significant moments centered on events, reduces redundant data processing, enhances adaptability to environmental changes, and improves the identification accuracy of solid waste.
[0248] The intelligent solid waste monitoring system based on multi-source data provided in this application, employing the intelligent solid waste monitoring method based on multi-source data in the above embodiments, can solve the technical problems of intelligent solid waste monitoring based on multi-source data. Compared with the prior art, the beneficial effects of the intelligent solid waste monitoring system based on multi-source data provided in this application are the same as those of the intelligent solid waste monitoring method based on multi-source data provided in the above embodiments, and other technical features of the intelligent solid waste monitoring system based on multi-source data are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0249] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0250] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0251] In addition, for technical details not described in detail in this embodiment, please refer to the intelligent monitoring method for solid waste based on multi-source data provided in any embodiment of the present invention, which will not be repeated here.
[0252] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0253] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0254] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0255] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for solid waste intelligent monitoring based on multi-source data, characterized in that, The solid waste intelligent monitoring method based on multi-source data comprises: Multi-source data monitoring is performed on a solid waste dropping area to obtain original multi-modal data, which is collected by multi-modal sensors pre-set in the solid waste dropping area; The original multi-modal data is pre-processed to obtain initial multi-modal data; Local sampling time of each modal data in the initial multi-modal data is mapped to an event reference time axis to generate a time mapping function, the event reference time axis comprises a plurality of event nodes, the event nodes are generated based on time nodes of associated events of solid waste, and the time mapping function comprises: wherein, denotes a modality time mapping function, denotes a modality time transformation network's learnable parameters, denotes a modality local time stretching coefficient at time instant, which is used to correct the response speed difference between modalities, denotes a modality local time offset, which is used to compensate the time offset problem between different modalities, denotes a first intermediate variable, denotes a second intermediate variable, denotes an initial time offset constant term, the first intermediate variable and the second intermediate variable are obtained by inputting modality data into a differentiable time series network. Based on the time mapping function and the spatial index of the initial multi-modal data, the initial multi-modal data is spatio-temporally aligned to obtain candidate multi-modal data; Cross-modal feature extraction is performed on the candidate multi-modal data to obtain feature embedding vectors of each modal data; Attention feature fusion is performed on the feature embedding vectors of each modal data to obtain multi-modal fusion features; Based on the multi-modal fusion features, the solid waste dropping area is monitored for solid waste classification.
2. The multi-source data based solid waste intelligent monitoring method of claim 1, wherein, The time mapping function and the spatial index of the initial multi-modal data are used to spatio-temporally align the initial multi-modal data to obtain candidate multi-modal data, which comprises: Based on the spatial index of the initial multi-modal data, spatial dimension saliency analysis is performed to obtain local event saliency parameters; Based on the local event saliency parameters and the time mapping function, global saliency analysis is performed on the initial multi-modal data to obtain global event saliency parameters, which are used to quantify the saliency parameter contribution degree of each modal data in the initial multi-modal data in the time dimension and the spatial dimension to the event nodes; Based on the global event saliency parameters, event center nodes are selected from the event nodes, the event center nodes are time nodes of salient events with global event saliency parameters not lower than a saliency threshold; Based on the time mapping function and the event center nodes, contribution weight analysis is performed on the initial multi-modal data to generate a soft alignment weight kernel function, which is used to define the contribution weight of each modal data in the initial multi-modal data to the event center nodes at each time node, so as to spatio-temporally align the initial multi-modal data; The initial multi-modal data is event correlation filtered by the soft alignment weight kernel function to obtain candidate multi-modal data.
3. The multi-source data based solid waste intelligent monitoring method of claim 2, wherein, The global saliency analysis of the initial multi-modal data based on the local event saliency parameters and the time mapping function to obtain global event saliency parameters comprises: Signal quality analysis is performed on each modal data in the initial multi-modal data to obtain signal reliability parameters of the initial multi-modal data; A weight adjustment function is constructed based on the signal reliability parameters, and the mathematical expression of the weight adjustment function is: wherein, represents a modality At time the signal credibility, represents a modality index, represents a modality weight, represents a modality index variable traversed in the summation process, represents the credibility sum of all modalities at time instance. performing global saliency analysis on the initial multi-modal data based on the weight adjustment function, the local event saliency parameter and the time mapping function to obtain a global event saliency parameter, wherein the global event saliency parameter is calculated according to the following formula: wherein, denotes a global event saliency parameter, denotes a set of modalities of the initial multi-modal data, denotes a modality of the local event saliency parameter, denotes an inverse function of the temporal mapping function of the modality , denotes a spatial index.
4. The multi-source data based solid waste intelligent monitoring method of claim 3, wherein, performing contribution weight analysis on the initial multi-modal data based on the time mapping function and the event center node to generate a soft alignment weight kernel function, including: determining a response scale of the soft alignment kernel based on the global event saliency parameter and a reference bandwidth, wherein the response scale is calculated according to the following formula: wherein, denotes a response scale for controlling the degree of temporal diffusion of the kernel function, denotes a reference bandwidth, which is a modal initial kernel width at the time of event-free adjustment, denotes a saliency adjustment hyperparameter, denotes a global event saliency parameter, denotes an event center node; performing contribution weight analysis on the initial multi-modal data based on the response scale, the event center node and the event mapping function to generate a soft alignment weight kernel function, wherein the mathematical expression of the soft alignment weight kernel function is: wherein, denotes a modality at time corresponding event center a soft alignment weight kernel, denotes a modality a time mapping function.
5. The multi-source data based solid waste intelligent monitoring method according to any one of claims 1 to 4, characterized in that, performing attention feature fusion on the feature embedding vectors of the multi-modal data to obtain multi-modal fusion features, including: taking the feature embedding vectors of the multi-modal data as graph nodes and constructing a node set based on the plurality of graph nodes; performing event correlation analysis on the feature embedding vectors of the multi-modal data and constructing a dynamic edge weight function based on the analysis result; wherein, denotes an event time node , a modality , a dynamic edge weight between modalities , denotes an event time node, denotes a time decay coefficient, denotes an exponential decay term of time difference, denotes a modality , a modality , a time difference between corresponding events, and denote a modality and a modality , respectively, a feature vector at a time point , denotes an importance weight of an event time node , denotes a modality and a modality , a cosine similarity at a time point ; determining attention scores between the modalities based on the graph attention network; constructing a double-weight attention coefficient fusion weight function based on the attention scores and determining target attention weights based on the double-weight attention coefficient fusion weight function, wherein the mathematical expression of the double-weight attention coefficient fusion weight function is: wherein, denotes a node a target attention weight of a node , denotes a content similarity weight, and denotes an attention score, denotes an uncertainty parameter of a node , denotes a set of neighbor nodes of a node , denotes a regularization term, denotes a neighbor node index, denotes a node a corresponding modality a node a corresponding modality a corresponding modality a node a corresponding modality a corresponding modality a corresponding modality a corresponding modality performing attention feature fusion on the feature embedding vectors of the multi-modal data based on the double-weight attention coefficient fusion weight to obtain multi-modal fusion features.
6. The multi-source data based solid waste intelligent monitoring method according to any one of claims 1 to 4, wherein, performing solid waste category identification monitoring on the solid waste disposal area based on the multi-modal fusion features, including: performing average pooling on the multi-modal fusion features to obtain global fusion features; performing linear mapping and activation processing on the global fusion features to obtain intermediate features; inputting the intermediate features into a Softmax classifier to output classification prediction probabilities of each type of waste, and performing solid waste category identification monitoring on the solid waste disposal area based on the classification prediction probabilities.
7. A solid waste intelligent monitoring system based on multi-source data, characterized in that, The system is used to implement the multi-source data-based solid waste intelligent monitoring method according to any one of claims 1 to 6, and the multi-source data-based solid waste intelligent monitoring system includes: a data monitoring module configured to perform multi-source data monitoring on a solid waste disposal area to obtain original multi-modal data, which is collected by multi-modal sensors pre-set in the solid waste disposal area; a data processing module configured to pre-process the original multi-modal data to obtain initial multi-modal data; a time mapping module configured to map local sampling times of multi-modal data in the initial multi-modal data to an event reference time axis to generate a time mapping function, wherein the event reference time axis includes a plurality of event nodes, and the event nodes are generated based on time nodes of associated events of solid waste; The spatio-temporal alignment module is configured to perform spatio-temporal alignment on the initial multi-modal data based on the time mapping function and a spatial index of the initial multi-modal data, to obtain candidate multi-modal data. The feature extraction module is configured to perform cross-modal feature extraction on the candidate multi-modal data, to obtain feature embedding vectors of multi-modal data. The feature fusion module is configured to perform attention feature fusion on the feature embedding vectors of multi-modal data, to obtain multi-modal fusion features. The recognition monitoring module is configured to perform solid waste category recognition monitoring on the solid waste dropping area based on the multi-modal fusion features.
8. A solid waste intelligent monitoring device based on multi-source data, characterized in that, The solid waste intelligent monitoring device based on multi-source data comprises a memory, a processor, and a solid waste intelligent monitoring program based on multi-source data stored in the memory. The processor is configured to run the solid waste intelligent monitoring program based on multi-source data. The solid waste intelligent monitoring program based on multi-source data is configured to implement the solid waste intelligent monitoring method based on multi-source data according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a solid waste intelligent monitoring program based on multi-source data. The solid waste intelligent monitoring program based on multi-source data is executed by the processor to implement the solid waste intelligent monitoring method based on multi-source data according to any one of claims 1 to 6.
Citation Information
Patent Citations
A multi-sensor fusion adaptive detection method for stone picking machine and intelligent stone picking machine
CN119760468A
Environment supervision method and system based on edge cloud architecture
CN121078074A