Monitoring information identification method and server based on Internet of Things

By obtaining multimodal perceptual data of IoT monitoring nodes for cross-modal feature extraction and fusion processing, a comprehensive spatiotemporal feature map is generated, and dynamic pattern matching is performed, the problem of data complementarity in the IoT monitoring system is solved, and accurate identification and efficient response to abnormal events are achieved.

CN120145321BActive Publication Date: 2025-08-12LESHAN YONGXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510631656.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-12
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing IoT monitoring system cannot fully explore the potential connections and complementarity between different modal data, resulting in the inability to form a comprehensive and comprehensive understanding of the target area, it is difficult to accurately judge the types of abnormal situations and their temporal and spatial distribution, and traditional methods lack flexibility and adaptability, and are prone to misjudgment or misjudgment.

Method used

Obtain multimodal perception data of multiple IoT monitoring nodes, including image, sound and environmental parameters, perform cross-modal feature extraction and fusion processing, generate a comprehensive spatio-temporal feature map, and dynamic pattern matching is performed through the preset abnormal mode library, generate an adapted regulatory instruction set, and distribute it to the IoT execution device cluster to eliminate the impact of abnormal events.

Benefits of technology

It realizes accurate identification of abnormal event types and spatiotemporal distribution information in the target monitoring area, improves the accuracy and response speed of abnormal detection, optimizes resource utilization efficiency, and realizes the automation and intelligent closed-loop operation of the monitoring and regulation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145321B_ABST
    Figure CN120145321B_ABST
Patent Text Reader

Abstract

The present invention provides a monitoring information identification method and server based on the Internet of Things. First, multimodal perception data containing images, sounds and environmental parameters collected in real time by multiple Internet of Things monitoring nodes in a target monitoring area is obtained. Then, cross-modal feature extraction and fusion processing are performed on the multimodal perception data to generate a comprehensive spatiotemporal feature map. Then, based on a preset abnormal pattern library, the type of abnormal event in the area and its spatiotemporal distribution information are determined through dynamic pattern matching. Based on this, an adaptive control instruction set is generated to clarify the control parameters and execution priorities of different Internet of Things execution devices. Finally, the instruction set is distributed to the associated Internet of Things execution device cluster, triggering them to execute the control parameters according to priority, eliminating the impact of abnormal events, and realizing comprehensive and intelligent monitoring and abnormality processing of the target monitoring area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things, and in particular to a monitoring information identification method and server based on the Internet of Things. Background Art

[0002] With the rapid development of Internet of Things technology, its application in various fields is becoming increasingly extensive. Effective monitoring of target areas has become a key requirement in many scenarios, such as industrial production monitoring, urban public area security monitoring, and intelligent building environment management.

[0003] Related technologies often analyze different types of data independently, failing to fully explore the potential connections and complementarities between different modalities. This isolated data processing approach prevents a holistic and comprehensive understanding of the target area, making it difficult to accurately determine whether anomalies exist within the area and the specific characteristics of such anomalies.

[0004] Traditional methods for identifying abnormal events are mostly based on fixed rules or simple thresholds. These methods lack sufficient flexibility and adaptability in the complex and ever-changing real-world monitoring scenarios. Once the monitoring environment changes or new abnormal patterns emerge, the existing rules may become invalid, leading to misidentification or omission of abnormal events and an inability to accurately determine the type of abnormal event or its distribution in time and space. Summary of the Invention

[0005] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a monitoring information identification method based on the Internet of Things, the method comprising:

[0006] Acquire multimodal sensing data collected in real time by multiple IoT monitoring nodes within a target monitoring area, wherein the multimodal sensing data includes image monitoring data, sound monitoring data, and environmental parameter monitoring data;

[0007] Performing cross-modal feature extraction and fusion processing on the multimodal sensing data to generate a comprehensive spatiotemporal feature map corresponding to the target monitoring area;

[0008] Based on a preset abnormal pattern library, dynamic pattern matching is performed on the comprehensive spatiotemporal feature map to determine at least one abnormal event type and its spatiotemporal distribution information existing in the target monitoring area;

[0009] Generate, according to the abnormal event type and the spatiotemporal distribution information, a set of control instructions adapted to the abnormal event type, the set of control instructions including control parameters and execution priorities for different IoT execution devices;

[0010] The control instruction set is distributed to the associated IoT execution device cluster within the target monitoring area, triggering the IoT execution device cluster to execute the control parameters according to the execution priority to eliminate the impact of the abnormal event.

[0011] On the other hand, an embodiment of the present invention also provides a server, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0012] Based on the above aspects, the embodiment of the present application obtains multimodal perception data from multiple IoT monitoring nodes, covering image monitoring data, sound monitoring data, and environmental parameter monitoring data, and then performs cross-modal feature extraction and fusion processing on the multimodal perception data to generate a comprehensive spatiotemporal feature map. It fully utilizes the complementarity between different modal data, can discover the complex internal correlations between the data, and thus generate a more representative and comprehensive feature map, which more accurately reflects the true state of the target monitoring area. Based on the preset abnormal pattern library, dynamic pattern matching is performed on the comprehensive spatiotemporal feature map, which realizes the accurate identification of the abnormal event type and its spatiotemporal distribution information in the target monitoring area, can adapt to different monitoring scenarios and real-time changes, and has higher flexibility and adaptability than traditional static matching or simple threshold judgment methods. It can more accurately capture complex and changeable abnormal events, greatly improving the accuracy and reliability of abnormal detection. Based on the type of abnormal event and its temporal and spatial distribution, an adaptive set of control instructions is generated, and the control parameters and execution priorities for different IoT execution devices are clearly defined. This allows for precise control strategies to be tailored for different execution devices based on specific abnormal situations, ensuring that each device plays its maximum role at the right time. By rationally setting execution priorities, resource utilization efficiency is further optimized, making the entire control process more efficient and orderly, and enabling the impact of abnormal events to be quickly and effectively eliminated. Finally, the control instruction set is distributed to the associated IoT execution device cluster and triggered to execute the control parameters according to priority, achieving automated, intelligent, closed-loop operation of the entire monitoring and control system. This greatly improves the response speed and processing effect of abnormal events, reduces the need for manual intervention, reduces human errors and response delays, and significantly improves the overall efficiency of the IoT monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a schematic diagram of the execution flow of the monitoring information identification method based on the Internet of Things provided by an embodiment of the present invention.

[0014] Figure 2 Schematic diagram of the hardware architecture of the server provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0015] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a monitoring information identification method based on the Internet of Things provided by an embodiment of the present invention. The monitoring information identification method based on the Internet of Things is introduced in detail below.

[0016] Step S110: acquiring multimodal sensing data collected in real time by multiple IoT monitoring nodes within a target monitoring area, wherein the multimodal sensing data includes image monitoring data, sound monitoring data, and environmental parameter monitoring data.

[0017] Taking the industrial security monitoring scenario as an example, multiple IoT monitoring nodes are deployed in a target monitoring area such as an industrial plant. These IoT monitoring nodes have multiple sensing capabilities and are used to obtain multimodal sensing data.

[0018] Specifically, for image monitoring data, high-definition cameras installed at key locations throughout the factory continuously capture surveillance footage. For example, cameras are located above assembly lines in the production workshop, in warehouse aisles, and at factory entrances and exits. These cameras capture the operating status of production equipment, worker operations, and the stacking and handling of goods. This continuous recording generates real-time image monitoring data.

[0019] Sound monitoring data is collected by sound sensors installed in various areas of the factory. In a noisy industrial environment, these sensors can detect the sounds of machine operation, unusual noises from equipment failures, the sounds of workers operating tools, and any unusual collisions. For example, when a component on a large stamping machine begins to wear, it may emit a sharp grinding sound, which will be captured by the sound sensor.

[0020] For environmental parameter monitoring, temperature and humidity sensors are distributed throughout the factory to monitor ambient temperature and humidity. In temperature- and humidity-sensitive production processes, such as electronic component manufacturing, temperature and humidity fluctuations can affect product quality. Furthermore, smoke sensors monitor smoke production, which is crucial for preventing fire hazards. For example, if an electrical device shorts and emits smoke, the smoke sensor will immediately detect changes in smoke concentration.

[0021] The above-mentioned different types of sensors jointly collect environmental parameter monitoring data, providing multi-faceted source data information for comprehensive monitoring of the safety status of industrial plants.

[0022] Step S120: performing cross-modal feature extraction and fusion processing on the multimodal sensing data to generate a comprehensive spatiotemporal feature map corresponding to the target monitoring area.

[0023] Specifically, multi-scale spatial feature extraction is first performed on the image monitoring data to obtain the visual semantic feature vector within the target monitoring area. Taking a production workshop as an example, the image frame sequence captured by the camera is first divided into preset time intervals. Assuming each interval is 5 seconds, multiple image frame sequences are obtained. These image frame sequences are then subjected to denoising and illumination equalization to eliminate noise caused by lighting variations and equipment interference. For example, flickering lights in the workshop or vibrations caused by machine operation can affect image quality. After processing, a clear and stable image frame sequence is obtained.

[0024] Next, a convolutional neural network is used to perform hierarchical feature extraction on the processed image frame sequence. During this process, neural networks at different levels can identify local texture feature maps and global scene feature maps at different scales. For example, for products on a production line, the neural network can identify subtle texture features on the product surface while also understanding global scene features such as the layout of the entire production line and the overall arrangement of equipment.

[0025] The spatial pyramid pooling layer then performs multi-resolution feature fusion on the local texture feature map and the global scene feature map to generate a fused spatial semantic feature matrix. For example, for products on an assembly line, information such as the texture features of different parts and the overall positional relationship on the assembly line are integrated into this matrix. The spatial semantic feature matrix is then aggregated using a sliding window in the time dimension to extract the dynamic trend characteristics of the target monitoring area over a continuous time period. Assuming a sliding window step size of 1 second and a window size of 10 seconds, dynamic trends such as product movement and changes in processing status on the production line within 10 seconds can be captured. Finally, the dynamic trend characteristics are concatenated with the spatial semantic feature matrix within the current time window to generate a visual semantic feature vector.

[0026] For sound monitoring data, voiceprint spectrum decomposition is performed to extract the characteristic vectors of acoustic events within the target monitoring area. In an industrial plant, the sounds emitted by different equipment have different voiceprint spectra. For example, the sound spectrum of a stamping machine is significantly different from that of a conveyor belt. Through voiceprint spectrum decomposition, the sound emitted by a specific piece of equipment can be accurately identified. When an abnormality in the sound of the equipment occurs, the acoustic characteristics of the abnormality are extracted.

[0027] Time series analysis of environmental parameter monitoring data is performed to extract characteristic vectors of environmental fluctuations within the target monitoring area, and a spatiotemporal correlation matrix is constructed corresponding to these characteristic vectors. For example, temperature and humidity sensors continuously collect data. Through time series analysis, patterns of temperature and humidity fluctuations throughout the day can be identified. If the temperature or humidity suddenly rises or falls abnormally within a certain period of time, these fluctuations are extracted, and a spatiotemporal correlation matrix is constructed based on the temperature and humidity variations at different locations in the factory. For example, the temperature and humidity variations in areas near heating equipment may differ from those in areas further away.

[0028] The visual semantic feature vector, acoustic event feature vector, and spatiotemporal correlation matrix are input into a cross-modal fusion network. The attention weight allocation layer within the cross-modal fusion network dynamically assigns the contribution weights of each modal feature in the spatiotemporal dimension. For example, when determining whether there is a potential equipment failure, image monitoring data may initially have a greater contribution weight to determining the device's appearance. However, when the device begins to make abnormal sounds, the contribution weight of the sound monitoring data may increase. At the same time, if temperature changes in environmental parameter monitoring data are related to equipment failure, their weight will also be adjusted accordingly. Based on these contribution weights, the visual semantic feature vector, acoustic event feature vector, and spatiotemporal correlation matrix are weighted and concatenated to generate a comprehensive spatiotemporal feature map. This comprehensive spatiotemporal feature map contains the spatial event distribution and environmental parameter evolution trends of the target monitoring area within a continuous time window, such as the spatial distribution of equipment operating status within the factory and the temporal trends of temperature and humidity.

[0029] Step S130 , based on a preset abnormal pattern library, dynamic pattern matching is performed on the comprehensive spatiotemporal feature map to determine at least one abnormal event type and its spatiotemporal distribution information existing in the target monitoring area.

[0030] In this embodiment, the pre-set abnormal pattern library is constructed based on historical monitoring data. For example, in past industrial security monitoring, a large piece of equipment suddenly stopped functioning. The multimodal sensing data at that time included images of smoke coming out of the equipment captured by the camera, unusual noise picked up by the sound sensor, and a sharp rise in the temperature around the equipment detected by the temperature and humidity sensor. The event type was marked as equipment failure, and the disposition record indicated that maintenance personnel performed emergency repairs, and the equipment returned to normal operation.

[0031] Multiple abnormal event templates are loaded from a pre-set abnormal pattern library. Each abnormal event template includes the reference spatiotemporal characteristics of historical abnormal events, event type labels, and impact range parameters. In this industrial plant scenario, different types of abnormal event templates may exist, such as equipment failure, fire hazard, and illegal intrusion.

[0032] Calculate the similarity matrix between the comprehensive spatiotemporal feature map and the reference spatiotemporal features of each abnormal event template. Suppose the current comprehensive spatiotemporal feature map shows a sudden temperature increase and unusual equipment noise in a certain area. By calculating the similarity matrix with the reference spatiotemporal features of the equipment failure abnormal event template, it is found that the two have a high degree of similarity. Based on the similarity matrix, a set of candidate abnormal event templates with a similarity above a preset threshold is selected. For example, if the threshold is set to 0.8 and the calculated similarity exceeds 0.8, the equipment failure abnormal event template is included in the set of candidate abnormal event templates.

[0033] Perform spatiotemporal alignment on each candidate abnormal event template in the candidate abnormal event template set, mapping the timestamps of the integrated spatiotemporal feature map and the timestamps of the candidate abnormal event templates to a unified time coordinate system. For example, when determining equipment failure, it is important to ensure that the times of temperature rise and abnormal sound in the integrated spatiotemporal feature map correspond to the times in the candidate abnormal event templates to accurately determine whether they are the same type of equipment failure.

[0034] Based on the feature differences after spatiotemporal alignment, a target abnormal event template that matches the comprehensive spatiotemporal feature map is identified from the candidate abnormal event template set. The event type label and impact range parameters of the target abnormal event template are used as the abnormal event type and spatiotemporal distribution information. If the abnormal event is determined to be an equipment failure, the spatiotemporal distribution information may include the specific location of the faulty equipment within the factory (the starting location), the possible spread path of the fault's impact (e.g., whether it affects adjacent production lines), and the fault impact intensity attenuation curve (e.g., whether the fault's impact gradually increases or decreases over time).

[0035] Step S140: generating a control instruction set adapted to the abnormal event type according to the abnormal event type and the spatiotemporal distribution information, wherein the control instruction set includes control parameters and execution priorities for different IoT execution devices.

[0036] In this embodiment, the emergency response policy associated with the abnormal event type can be called from the IoT device policy library based on the abnormal event type. Taking device failure as an example, the emergency response policy in the IoT device policy library includes device linkage rules, parameter adjustment rules, and execution condition constraints. Device linkage rules may involve actions such as notifying maintenance personnel and shutting down associated devices to prevent the fault from escalating. Parameter adjustment rules may include adjusting the operating parameters of other devices related to the faulty device, such as reducing the power of adjacent devices to reduce the load on the faulty device. Execution condition constraints may specify that certain actions are only executed after the temperature around the faulty device reaches a certain threshold or the faulty device stops operating.

[0037] Based on the impact intensity attenuation curve from the spatiotemporal distribution information, the optimal startup time and operating duration of each execution device in the IoT execution device cluster are determined. If the impact intensity of a device failure gradually increases over time, the optimal startup time for the maintenance equipment and related auxiliary equipment should be as soon as possible, and the operating duration should be long enough to ensure that the fault is completely repaired.

[0038] Control parameters are assigned to each execution device based on device linkage and parameter adjustment rules. For example, for a maintenance robot, its operating power should be adjusted based on the difficulty of repairing the faulty device, its direction of action should be directed to the specific location of the faulty device, and its coverage should ensure that it can reach all possible fault points of the device.

[0039] Combined with execution constraints and optimal startup time, the system calculates an estimated resource consumption for each execution device within the target monitoring area. Control parameters are then optimized and calibrated based on this estimated resource consumption. For example, a maintenance robot consumes different amounts of energy at different power levels. If lowering power while still meeting maintenance requirements can reduce energy consumption, then the operating power level is optimized and calibrated.

[0040] Based on the priority weights of the impact intensity attenuation curve and the estimated resource consumption, the control parameters of each execution device are assigned an execution priority, and a set of control instructions is generated. For example, devices that can quickly reduce the impact intensity of a fault are given a higher execution priority. Resource consumption is also taken into consideration. If a device helps reduce the impact intensity but consumes too many resources, its execution priority may be lowered.

[0041] Step S150: Distribute the control instruction set to the IoT execution device cluster associated with the target monitoring area, triggering the IoT execution device cluster to execute the control parameters according to the execution priority to eliminate the impact of the abnormal event.

[0042] In this embodiment, the control instruction set can be divided into multiple instruction batches based on execution priority, and each instruction batch can be sent to the corresponding execution device sub-cluster in batch order. For example, instructions related to emergency stopping faulty equipment can be sent to the corresponding execution device sub-cluster (such as the device control system) as the first batch of instructions, and instructions related to maintenance equipment preparation can be sent to the maintenance device sub-cluster as the second batch of instructions.

[0043] When sending each instruction batch, the response status of the execution device sub-cluster is monitored in real time. If the target execution device fails to return a response confirmation signal within the preset time, a backup device replacement strategy is initiated. For example, if a maintenance robot fails to return a response confirmation signal within the specified time, it may be faulty and unable to execute the instruction. A set of candidate backup devices with the same functional attributes as the target execution device (maintenance robot), such as other maintenance robots of the same model, is selected from the IoT execution device cluster. The current operating status, remaining resource capacity, and physical distance from the target monitoring area (the location of the faulty device) of each candidate backup device are obtained. If a backup maintenance robot is currently idle, has sufficient remaining resource capacity, and is close to the faulty device, its availability score is higher. Based on these factors, a comprehensive replacement priority list is generated, and the candidate backup device with the highest score is selected as the target backup device. The control parameters corresponding to the target execution device are scaled according to a preset ratio and assigned to the target backup device. The scaled control parameters and new execution priority are sent to the target backup device, and the device identifier and parameter record in the control instruction set are updated.

[0044] After the execution device subcluster completes the parameter execution of the current instruction batch, it collects real-time feedback data from the target monitoring area and evaluates the rate of change of the impact intensity of the abnormal event based on the real-time feedback data. If it is a device failure, the sensor will re-detect the temperature around the device, the operating status of the device, and other information to calculate whether the impact intensity of the fault is attenuating as expected. If the rate of change of the impact intensity does not reach the expected attenuation rate, the control parameters and execution priority of subsequent instruction batches are dynamically adjusted based on the real-time feedback data. For example, if the impact intensity of the fault does not decrease as expected, it may be necessary to increase the operating power of the maintenance equipment or adjust the maintenance strategy, and accordingly increase the execution priority of certain execution devices.

[0045] After all instruction batches are executed, a handling report for the abnormal event is generated and uploaded to the IoT monitoring platform for archiving. This report includes information such as the type of abnormal event, the time and location of occurrence, the operation of the executing device during the handling process, and the final handling results of the abnormal event. This information facilitates subsequent query and analysis, providing a reference for handling similar abnormal events in the future.

[0046] Based on the above steps, the embodiment of the present application obtains multimodal perception data from multiple IoT monitoring nodes, covering image monitoring data, sound monitoring data, and environmental parameter monitoring data, and then performs cross-modal feature extraction and fusion processing on the multimodal perception data to generate a comprehensive spatiotemporal feature map. It fully utilizes the complementarity between different modal data and can discover the complex internal correlations between the data, thereby generating a more representative and comprehensive feature map that more accurately reflects the true state of the target monitoring area. Based on the preset abnormal pattern library, dynamic pattern matching is performed on the comprehensive spatiotemporal feature map to achieve accurate identification of the abnormal event type and its spatiotemporal distribution information in the target monitoring area. It can adapt to different monitoring scenarios and real-time changes. Compared with traditional static matching or simple threshold judgment methods, it has higher flexibility and adaptability, can more accurately capture complex and changeable abnormal events, and greatly improves the accuracy and reliability of abnormality detection. Based on the type of abnormal event and its temporal and spatial distribution, an adaptive set of control instructions is generated, and the control parameters and execution priorities for different IoT execution devices are clearly defined. This allows for precise control strategies to be tailored for different execution devices based on specific abnormal situations, ensuring that each device plays its maximum role at the right time. By rationally setting execution priorities, resource utilization efficiency is further optimized, making the entire control process more efficient and orderly, and enabling the impact of abnormal events to be quickly and effectively eliminated. Finally, the control instruction set is distributed to the associated IoT execution device cluster and triggered to execute the control parameters according to priority, achieving automated, intelligent, closed-loop operation of the entire monitoring and control system. This greatly improves the response speed and processing effect of abnormal events, reduces the need for manual intervention, reduces human errors and response delays, and significantly improves the overall efficiency of the IoT monitoring system.

[0047] In a possible implementation, step S120 includes:

[0048] Step S121 , performing multi-scale spatial feature extraction on the image monitoring data to obtain a visual semantic feature vector within the target monitoring area, and performing voiceprint spectrum decomposition on the sound monitoring data to extract an acoustic event feature vector within the target monitoring area.

[0049] For example, taking equipment monitoring in a production workshop as an example, multi-scale spatial feature extraction first requires preprocessing of the image monitoring data. Cameras installed at various locations in the workshop continuously collect image data, which contains information such as the equipment's appearance, operating status, and surrounding environment. This image data is then divided into image frame sequences at preset time intervals, for example, every 10 seconds. Because industrial environments may experience image noise and unstable lighting due to factors such as uneven lighting and equipment vibration, each image frame sequence must be denoised and illuminated with equalization. This processed image frame sequence more accurately reflects the actual conditions within the workshop.

[0050] Next, a convolutional neural network is used for hierarchical feature extraction. During this process, the different layers of the convolutional neural network can identify local texture feature maps and global scene feature maps at different scales. For large-scale production equipment in a workshop, the local texture feature map may reveal microscopic features such as surface wear and tear and component connection status, while the global scene feature map can reveal macroscopic information such as the entire equipment's location in the workshop, its relative layout to other equipment, and surrounding human activity.

[0051] The local texture feature map and the global scene feature map are then fused at multiple resolutions using a spatial pyramid pooling layer. For example, different components of production equipment on an assembly line may have distinct features at different resolutions. This fusion method integrates these features at different resolutions into a fused spatial semantic feature matrix, encompassing semantic information about each component and the entire machine within the workshop.

[0052] The spatial semantic feature matrix is then subjected to sliding window aggregation along the time dimension to extract the dynamic change trend characteristics of the target monitoring area within a continuous time period. Assume that the sliding window step size is 5 seconds and the window size is 30 seconds. During this time period, production equipment may experience different states such as startup, stable operation, and shutdown. Sliding window aggregation can capture the dynamic changes in equipment status during this period, such as whether the equipment's operating speed is stable and whether there are any abnormal pauses. Finally, the dynamic change trend characteristics are spliced with the spatial semantic feature matrix within the current time window to obtain the visual semantic feature vector within the target monitoring area. This visual semantic feature vector can comprehensively reflect the visual characteristics of the equipment and environment in the workshop and their changes over time.

[0053] At the same time, the sound monitoring data is subjected to voiceprint spectrum decomposition to extract acoustic event feature vectors. Industrial plants are plagued by a variety of sound sources. Different equipment, due to their varying structures, operating principles, and working conditions, emits sounds with specific voiceprint spectra. For example, cutting equipment emits a stable sound with a specific frequency and amplitude during normal operation. As the cutting tool wears, the frequency and amplitude of the sound change, and the voiceprint spectrum also changes accordingly. Voiceprint spectrum decomposition technology can decompose the collected sound monitoring data into spectral components of varying frequencies and amplitudes, accurately extracting feature vectors associated with specific acoustic events. These acoustic event feature vectors can reflect the characteristics of various sound events within the workshop, such as whether the equipment is operating normally and whether any parts are loose or collided.

[0054] Step S122 , performing time series analysis on the environmental parameter monitoring data, extracting the environmental fluctuation characteristic vectors within the target monitoring area, and constructing a spatiotemporal correlation matrix corresponding to the environmental fluctuation characteristic vectors.

[0055] In industrial plants, environmental parameters such as temperature, humidity, and smoke concentration have a significant impact on production processes and equipment operation. Temperature, humidity, and smoke sensors continuously collect data on these environmental parameters. For example, temperature fluctuates throughout the day due to equipment heat dissipation, ventilation system operation, and the external environment. Time series analysis can transform these time-varying temperature data into environmental fluctuation feature vectors. Furthermore, due to varying equipment distribution and ventilation conditions in different areas of the plant, temperature fluctuations also vary spatially. Constructing a spatiotemporal correlation matrix corresponding to these environmental fluctuation feature vectors can describe the relationship between temperature at different temporal and spatial locations. For example, temperatures near large heat-generating equipment may fluctuate significantly, while temperatures farther away may remain relatively stable. The spatiotemporal correlation matrix accurately reflects these spatial and temporal temperature correlations.

[0056] In step S123, the visual semantic feature vector, the acoustic event feature vector, and the spatiotemporal correlation matrix are input into a cross-modal fusion network, and the contribution weight of each modal feature in the spatiotemporal dimension is dynamically allocated through the attention weight allocation layer in the cross-modal fusion network.

[0057] Step S124: Based on the contribution weights, weighted concatenation is performed on the visual semantic feature vector, the acoustic event feature vector, and the spatiotemporal correlation matrix to generate the comprehensive spatiotemporal feature map. The comprehensive spatiotemporal feature map includes the spatial event distribution and environmental parameter evolution trends of the target monitoring area within a continuous time window.

[0058] In a cross-modal fusion network, the attention weight allocation layer plays a key role in dynamically assigning the contribution weights of each modal feature in the spatiotemporal dimension. In industrial plant monitoring scenarios, different abnormal events may have varying degrees of dependence on features from different modalities. For example, when determining whether equipment has external damage, the visual semantic feature vector may have a higher contribution weight. However, when determining whether an internal fault has caused abnormal sound, the acoustic event feature vector's contribution weight increases. If the fault is related to environmental factors, such as excessive temperature, the weight of the spatiotemporal correlation matrix corresponding to the environmental fluctuation feature vector becomes more important. After the attention weight allocation layer dynamically adjusts the weights of each modal feature based on the actual situation, the visual semantic feature vector, acoustic event feature vector, and spatiotemporal correlation matrix are weighted and concatenated based on these contribution weights to generate a comprehensive spatiotemporal feature map. This comprehensive spatiotemporal feature map contains the spatial event distribution and environmental parameter evolution trends of the target monitoring area within a continuous time window. For example, it can display the operating status distribution of equipment at different times within the workshop, the movement trajectories of personnel, and the changing trends of environmental parameters such as temperature and humidity.

[0059] In a possible implementation, step S130 includes:

[0060] Step S131 : Load multiple abnormal event templates from the preset abnormal pattern library, each of the abnormal event templates including reference spatiotemporal features, event type labels, and impact range parameters corresponding to historical abnormal events.

[0061] In this embodiment, the preset abnormal pattern library is constructed based on the historical monitoring data of the industrial plant. In the past industrial security monitoring process, a large amount of abnormal event data was recorded. For example, a short circuit fault occurred in an electrical control cabinet. The monitoring data at that time included image monitoring data showing smoke and flashes of fire in the control cabinet, sound monitoring data capturing the strong arc sound during the short circuit, and environmental parameter monitoring data showing a sharp increase in temperature and smoke concentration in the area. The event was marked as an electrical equipment short circuit fault type, and the handling process was recorded, such as cutting off the power supply, using fire extinguishing equipment, and maintenance personnel performing repairs.

[0062] Abnormal event templates contain reference spatiotemporal features, event type labels, and impact range parameters corresponding to various historical abnormal events. In industrial plant scenarios, abnormal event templates may include types such as equipment failures (such as mechanical and electrical failures), fire hazards, and hazardous gas leaks. The reference spatiotemporal features in each template summarize the characteristics of historical abnormal events at different times and spaces. The event type label specifies the type of abnormal event, and the impact range parameter describes the area and number of devices that may be affected by the abnormal event.

[0063] Step S132 , calculating a similarity matrix between the comprehensive spatiotemporal feature map and the reference spatiotemporal features of each of the abnormal event templates, and screening out a set of candidate abnormal event templates whose similarity is higher than a preset threshold based on the similarity matrix.

[0064] Suppose the current comprehensive spatiotemporal feature map shows elevated temperatures, smoke, and unusual equipment noise in a certain area of the workshop. By calculating a similarity matrix with the reference spatiotemporal features of various abnormal event templates (such as fire hazard templates and equipment failure templates), the degree of similarity between them can be quantified. Based on this similarity matrix, a set of candidate abnormal event templates with similarities exceeding a preset threshold is selected. For example, if the similarity threshold is set to 0.8, and the similarity with the fire hazard template reaches 0.85 and the similarity with the equipment failure template is 0.7, then the fire hazard template will be selected as a candidate abnormal event template.

[0065] Step S133 , performing spatiotemporal alignment processing on each candidate abnormal event template in the candidate abnormal event template set, and mapping the timestamp of the comprehensive spatiotemporal feature map and the timestamp of the candidate abnormal event template to a unified time coordinate system.

[0066] Different historical abnormal events may have differences in time and space scales when they are recorded. For example, for fire hazard events, the time in the historical records may be in minutes, while the time scale of the current comprehensive spatiotemporal feature map may be in seconds; spatially, the position in the historical records may be relative to an old coordinate system, while the current coordinate system may have been updated. Through spatiotemporal alignment processing, the timestamps of the comprehensive spatiotemporal feature map and the timestamps of the candidate abnormal event templates are mapped to a unified time coordinate system. For example, the current timestamp in seconds is converted to a time scale corresponding to minutes in the historical records, and the current spatial coordinates are converted to the same coordinate system as in the historical records to ensure comparison within the same spatiotemporal framework.

[0067] Step S134: According to the feature difference after spatiotemporal alignment, a target abnormal event template that matches the comprehensive spatiotemporal feature map is determined from the candidate abnormal event template set, and the event type label and impact range parameters of the target abnormal event template are used as the abnormal event type and the spatiotemporal distribution information.

[0068] The spatiotemporal distribution information includes the starting position, diffusion path and impact intensity attenuation curve of the abnormal event in the target monitoring area.

[0069] For example, if, after spatiotemporal alignment, the feature difference from the fire hazard template is minimal, the abnormal event is determined to be a fire hazard. Its spatiotemporal distribution information includes the fire hazard's initial location within the workshop, which could be an area with concentrated electrical equipment; its diffusion path, which could be along ventilation ducts or near flammable material storage; and its impact intensity attenuation curve, which could indicate the rate at which the fire hazard's impact will expand over time if no action is taken, such as the area spread per minute. This spatiotemporal distribution information is crucial for subsequent targeted response measures.

[0070] In a possible implementation, step S140 includes:

[0071] Step S141: Based on the abnormal event type, an emergency response strategy associated with the abnormal event type is called from an IoT device strategy library. The emergency response strategy includes device linkage rules, parameter adjustment rules, and execution condition constraints.

[0072] Taking a fire hazard in an industrial plant as an example, the emergency response policy in the IoT device policy library includes rules for device linkage, parameter adjustment, and execution constraints. For abnormal events like fire hazards, device linkage rules specify the coordinated operations between related devices. For example, linkages exist between the fire alarm system, fire-fighting equipment (such as fire extinguishers and sprinklers), ventilation equipment, and production equipment associated with hazardous areas. Upon detecting a fire hazard, the fire alarm system should immediately trigger the preparation of fire-fighting equipment. Ventilation equipment should adjust its ventilation mode based on the intensity of the fire and smoke, perhaps increasing ventilation to expel smoke while avoiding intensifying the fire. Production equipment associated with the hazardous area should be shut down according to predefined rules to prevent further damage from the fire.

[0073] Parameter adjustment rules have different parameter settings for different equipment when dealing with fire hazards. The operating power of the fire sprinkler needs to be adjusted according to the severity of the fire hazard. If the temperature in the fire hazard area rises quickly and the smoke concentration is high, it may mean that the fire is large, so the operating power of the fire sprinkler must be increased accordingly to ensure that enough fire extinguishing agent is sprayed. The direction of action of the fire sprinkler must accurately point to the starting location of the fire hazard and its possible spread path, and the coverage range must ensure that it can cover the area where the fire may spread. The operating power of the ventilation equipment must also be adjusted according to the spread of smoke. Its direction of action must be conducive to the discharge of smoke, and the coverage range must include the entire area that may be affected by the smoke.

[0074] Execution constraints specify the conditions under which a device must perform an action. For example, a fire sprinkler might activate when the fire alarm system detects a temperature exceeding a certain threshold and smoke concentration reaching a certain standard. Ventilation equipment might adjust based on smoke concentration changes reported by a smoke sensor within a certain timeframe after a fire alarm.

[0075] Step S142: determining the optimal startup time and action duration of each execution device in the IoT execution device cluster based on the impact intensity attenuation curve in the spatiotemporal distribution information.

[0076] For example, for fire hazard events, the impact intensity attenuation curve reflects the development of the fire over time if no measures are taken or different measures are taken. Assume that the impact intensity attenuation curve shows that within the first 5 minutes after the fire occurs, if effective measures are not taken, the fire will spread at a faster rate. According to this impact intensity attenuation curve, the optimal activation time of the fire extinguishing equipment should be as short as possible after the fire is detected, for example, within 1 minute. The duration of the fire extinguishing equipment should be determined based on the expected development of the fire. If the fire is expected to be effectively controlled within 10 minutes, the duration of the fire extinguishing equipment should be set to at least 10 minutes to ensure that the fire is completely extinguished. The optimal activation time of the ventilation equipment should also be as early as possible to exhaust the smoke as soon as possible, and its duration of action should continue until the smoke concentration drops below the safety standard.

[0077] Step S143: allocating the control parameters to each of the execution devices according to the device linkage rule and the parameter adjustment rule, wherein the control parameters include operating power, action direction, and coverage range.

[0078] For example, fire sprinklers operate at a lower power level based on the severity of the fire. If the fire is in its early stages, the power level can be set to a lower level, such as 5 liters per minute. Once the fire grows to a certain level, the power level is increased to 10 liters per minute. The direction of fire sprinkler action is adjusted based on the fire's origin and spread. For example, if a fire starts in the southeast corner of a workshop and tends to spread northward, the sprinkler's direction of action should be adjusted to the southeast, covering the area extending northward from the southeast corner. The coverage area should be set based on the potential spread of the fire, such as a circular area with a radius of 5 meters. For ventilation equipment, the power level is adjusted based on smoke concentration. If smoke concentration is high, the power level is increased to 80% of the maximum ventilation rate, directed toward the area with the thickest smoke, and the coverage area should cover the entire workshop.

[0079] Step S144 , calculating the estimated resource consumption of each execution device in the target monitoring area in combination with the execution condition constraint and the optimal startup time, and optimizing and calibrating the control parameters based on the estimated resource consumption.

[0080] Taking a fire sprinkler as an example, if it is started within 1 minute of the optimal start-up time and runs for 10 minutes at the set operating power, the total amount of fire extinguishing agent consumed can be calculated. This is the estimated resource consumption of the fire sprinkler. At the same time, considering the energy supply of the fire sprinkler (if it is an electric sprinkler) or the fire extinguishing agent reserves, if the resource consumption estimate is too high, it may lead to insufficient resources during the subsequent fire fighting process. Then it is necessary to optimize and calibrate the control parameters, such as reducing the operating power or adjusting the duration of action. The same is true for ventilation equipment. The estimated resource consumption of electricity consumption is calculated based on its optimal start-up time, operating power and duration of action. If it exceeds the energy reserve or energy supply limit of the equipment, it is necessary to adjust the control parameters such as the operating power or duration of action.

[0081] Step S145 : allocating the execution priority to the control parameter of each execution device according to the priority weight of the impact intensity attenuation curve and the estimated resource consumption value, and generating the control instruction set.

[0082] For example, in the case of a fire hazard, fire-fighting equipment plays the most direct and critical role in controlling the fire and reducing the intensity of the impact, so it has a higher priority weight in the impact intensity attenuation curve. If the estimated resource consumption of the fire-fighting equipment is within an acceptable range, then its execution priority will be set to the highest. Although ventilation equipment is also important, its role in attenuating the intensity of the impact is slightly less than that of fire-fighting equipment, and if its resource consumption is high, its execution priority may be slightly lower. Based on these factors, an execution priority is assigned to each execution device, thereby generating a control instruction set. This control instruction set contains the control parameters and execution priority of each execution device, which is used to guide the IoT execution device cluster in responding to fire hazard events.

[0083] Wherein, step S150 includes:

[0084] Step S151 : dividing the control instruction set into a plurality of instruction batches according to the execution priority, and sending each instruction batch to a corresponding execution device sub-cluster in batch order.

[0085] For example, in the event of a fire hazard, commands related to fire extinguishing equipment are sent to the fire extinguishing equipment sub-cluster as the first batch of commands, as fire extinguishing equipment has the highest execution priority. This batch of commands contains control parameters such as the fire extinguishing equipment's activation time, operating power, direction of action, and coverage. Then, commands related to ventilation equipment are sent to the ventilation equipment sub-cluster as the second batch of commands, containing these control parameters.

[0086] Step S152: When sending each batch of instructions, the response status of the execution device sub-cluster is monitored in real time. If it is detected that the target execution device does not return a response confirmation signal within a preset time, the backup device replacement strategy is activated to reallocate the control parameters corresponding to the target execution device to the backup execution device.

[0087] For example, if a command is sent to a fire sprinkler and no response confirmation signal is received from it within the preset 10 seconds, it indicates that the sprinkler may be faulty. At this point, a backup device replacement strategy is initiated. A set of backup sprinklers with the same functional attributes as the specified sprinkler is selected from the IoT execution device cluster. Each backup sprinkler's current operating status (e.g., whether it is idle), remaining resource capacity (e.g., the amount of remaining fire extinguishing agent), and physical distance from the target monitoring area (fire hazard area) are obtained. If a backup sprinkler is currently idle, has sufficient remaining fire extinguishing agent, and is close to the fire hazard area, its availability score is high. Based on these factors, a comprehensive replacement priority list is generated, and the backup sprinkler with the highest score in the comprehensive replacement priority list is selected as the target backup device. The control parameters corresponding to the faulty sprinkler are scaled according to a preset ratio and assigned to the target backup device. For example, if the original control parameter was 10 liters of fire extinguishing agent per minute, the control parameter would be scaled to 8 liters per minute based on the remaining resource capacity of the target backup device. The scaled control parameters and new execution priority are sent to the target standby device, and the device identification and parameter records in the control instruction set are updated.

[0088] Step S153 : After the execution device sub-cluster completes parameter execution of the current instruction batch, real-time feedback data of the target monitoring area is collected, and the impact intensity change rate of the abnormal event is evaluated based on the real-time feedback data.

[0089] For example, in the case of a fire hazard, after the fire-fighting equipment subcluster executes its first batch of commands, it collects real-time feedback data from the target monitoring area (the fire hazard area) through devices such as temperature sensors and smoke sensors. The rate of change of the fire's impact intensity is calculated, for example, by measuring the rate of temperature drop or smoke concentration reduction. If the temperature does not drop significantly or the smoke concentration does not decrease as expected within a period of time after the fire-fighting equipment executes the command, the rate of change of the impact intensity has not reached the expected decay rate.

[0090] Step S154: If the impact intensity change rate does not reach the expected attenuation rate, dynamically adjust the control parameters and the execution priority in subsequent instruction batches according to the real-time feedback data.

[0091] For example, if the intensity of a fire doesn't decrease as expected, it may be necessary to increase the operating power of the fire-fighting equipment or adjust its direction of action. For example, if a fire in a certain area is found to be under control, the operating power of the fire sprinklers near that area may need to be increased from 8 liters of extinguishing agent per minute to 12 liters per minute. At the same time, because the role of fire-fighting equipment is more critical, its execution priority may need to be increased to ensure that it can more effectively implement control parameters. For ventilation equipment, control parameters such as operating power and direction of action may also need to be adjusted based on the spread of smoke.

[0092] In step S155, after all the instruction batches are executed, a processing report of the abnormal event is generated, and the processing report is uploaded to the Internet of Things monitoring platform for archiving.

[0093] For example, a handling report contains detailed information about a fire hazard, such as its origin, discovery time, impact area, emergency response strategy, operational status of the executing devices (including the execution of control parameters such as each executing device's start-up time, operating power, direction of action, and coverage), and the final handling results of the abnormal event (such as whether the fire was extinguished, extinguishing time, and any equipment damage). This handling report is uploaded to the IoT monitoring platform for archiving and future query and analysis, providing a reference for handling similar fire hazard incidents in the future.

[0094] In a possible implementation, step S121 includes:

[0095] Step S1211 : dividing the image monitoring data into a plurality of image frame sequences according to a preset time interval, and performing denoising and illumination equalization processing on each of the image frame sequences to obtain a processed image frame sequence.

[0096] For example, cameras installed at various locations throughout a factory continuously collect image data, which contains a wide range of information, including production equipment, personnel, and goods. This image data is divided into multiple image frame sequences at predetermined intervals, such as every five seconds. In industrial environments, images can be noisy and have uneven lighting due to factors such as flickering lighting and vibrations from equipment operation. Using specialized image processing algorithms, each image frame sequence is denoised to remove noise such as speckles and color variations caused by environmental interference. Lighting is also balanced to achieve more balanced brightness and contrast across different image areas. The resulting processed image frame sequence more accurately reflects the actual conditions within the factory.

[0097] In step S1212, a convolutional neural network is used to perform hierarchical feature extraction on the processed image frame sequence to obtain local texture feature maps and global scene feature maps at different scales.

[0098] For example, when monitoring a production workshop, different layers of a convolutional neural network can identify features at different scales. For example, when identifying production equipment, shallower layers can capture local texture features on the equipment surface, such as tiny scratches on the equipment casing and worn markings. Deeper layers, on the other hand, can capture global scene features, including macroscopic scene information such as the equipment's location and layout within the workshop, its relative position to other equipment, and the distribution of surrounding personnel.

[0099] Step S1213 , performing multi-resolution feature fusion on the local texture feature map and the global scene feature map through a spatial pyramid pooling layer to generate a fused spatial semantic feature matrix.

[0100] In actual factory monitoring, features at different resolutions contain different semantic information. For example, when inspecting products on an assembly line, local details are clearer at high resolution, while their overall positional relationships on the assembly line are better represented at low resolution. The spatial pyramid pooling layer fuses these local texture feature maps at different resolutions with the global scene feature map, integrating features at various scales into a fused spatial semantic feature matrix. This matrix contains semantic information about various objects in the factory, from the microscopic to the macroscopic, and from the local to the global, such as the status of equipment components and their spatial relationships within the factory.

[0101] Step S1214 , performing sliding window aggregation on the spatial semantic feature matrix in the time dimension, and extracting dynamic change trend features of the target monitoring area in a continuous time period.

[0102] During industrial production, conditions within the factory floor constantly change over time. The sliding window step and window size are set, for example, to 2 seconds for a step and 10 seconds for a window. Within this 10-second window, production equipment may experience different states, such as startup, acceleration, stable operation, and deceleration, and workers may move between different areas. Sliding window aggregation can capture the dynamic changes of these objects within this continuous time period, such as trends in equipment speed and worker movement trajectories.

[0103] Step S1215 , concatenating the dynamic change trend feature with the spatial semantic feature matrix in the current time window to generate the visual semantic feature vector.

[0104] In a factory monitoring scenario, this visual semantic feature vector fully encompasses the visual information within the target monitoring area and its dynamic changes over time. For example, it includes both spatial semantic information such as the appearance and position of production equipment at a specific moment, as well as temporal information such as the changing trends in the equipment's operating status over a period of time.

[0105] In one possible implementation, the method for constructing the preset abnormal pattern library includes:

[0106] Step S210 , collecting a plurality of historical case data marked as abnormal events in the historical monitoring data, each of the historical case data includes multimodal perception data, event type labels and handling records when the abnormality occurs.

[0107] Over the long-term operation of industrial plants, a vast amount of historical monitoring data has been accumulated. To build a pre-set abnormal pattern library, we first need to collect multiple historical case data marked as abnormal events from this historical monitoring data. This historical case data covers a wide range of possible abnormal situations within the plant.

[0108] For example, a short circuit occurred in an electrical device in a production workshop. The multimodal sensing data at the time included image monitoring data showing smoke and flashes of fire coming from the electrical control cabinet, sound monitoring data capturing strong arcing sounds and unusual humming noises from the equipment, and environmental parameter monitoring data showing a sharp rise in temperature and smoke concentration in the area. The event was labeled as an electrical equipment short circuit, and the handling record included a series of actions. For example, maintenance personnel arrived at the scene three minutes after receiving the alarm, first disconnecting the power supply to the equipment, then using a fire extinguisher to extinguish the flames that could have caused a fire. After 15 minutes of inspection and replacement of the short-circuited electrical component, the equipment resumed normal operation.

[0109] Step S220: performing cross-modal feature extraction on the multimodal perception data in the historical case data to generate a historical spatiotemporal feature map corresponding to each of the historical case data.

[0110] Taking the electrical equipment short-circuit fault mentioned earlier as an example, for image monitoring data, multi-scale spatial feature extraction yields a visual semantic feature vector. This vector contains visual feature information such as the location of smoke in the electrical control cabinet and the range of the fire. For sound monitoring data, the acoustic event feature vector obtained after voiceprint spectrum decomposition can reflect the frequency and amplitude of arcing sound, as well as the characteristics of abnormal humming sound. For environmental parameter monitoring data, the environmental fluctuation feature vector and spatiotemporal correlation matrix obtained through time series analysis can reflect the rate of temperature rise and the spatial diffusion of smoke concentration. These feature vectors and matrices of different modes are fused to generate a historical spatiotemporal feature map corresponding to the electrical equipment short-circuit fault event. This historical spatiotemporal feature map fully describes the multimodal characteristics of this abnormal event within a specific time and space.

[0111] Step S230 : clustering the historical spatiotemporal feature maps into different abnormal event categories according to the event type labels, and generating an initial abnormal event template for each abnormal event category.

[0112] Industrial plants experience a variety of abnormal event types, including equipment failures (including mechanical and electrical failures), fire hazards, and hazardous gas leaks. All historical spatiotemporal feature maps marked as electrical equipment short-circuit failures are clustered into the abnormal event category of equipment failure. Then, for this category, the common features within these historical spatiotemporal feature maps are analyzed, such as the similarity of visual features such as smoke and flames in images of equipment failures, the commonality of abnormal sounds, and the patterns of temperature rise and smoke generation in environmental parameters. Based on these common features, an initial abnormal event template is generated. This initial abnormal event template includes the typical spatiotemporal features of the equipment failure type, the event type label (equipment failure - electrical short circuit), and preliminary estimated impact range parameters, such as the number of adjacent devices potentially affected and the area of the surrounding area.

[0113] Step S240 , based on the treatment effect evaluation index in the treatment record, calibrate the impact range parameter in the initial abnormal event template so that the calibrated impact range parameter is negatively correlated with the treatment effect evaluation index.

[0114] For example, evaluation metrics for the effectiveness of handling electrical equipment short-circuit faults might include the repair time for the faulty equipment, whether a secondary fault occurs, and the extent of the impact on surrounding equipment and production. If, during the handling process, it is discovered that the number of adjacent devices actually affected is less than initially estimated, the repair time is short, and there is no significant impact on surrounding production, this indicates that the impact range parameter in the initial abnormal event template may be too large. Based on these evaluation metrics, the impact range parameter is calibrated to more accurately reflect the actual situation. For example, the number of adjacent devices potentially affected can be adjusted from 5 to 3, and the affected area of the surrounding area can be adjusted from 20 square meters to 15 square meters.

[0115] Step S250: storing each calibrated abnormal event template into the preset abnormal pattern library, and establishing a mapping relationship between the abnormal event template and the emergency response strategy in the IoT device strategy library.

[0116] For example, the preset anomaly pattern library stores various calibrated anomaly event templates, such as those for equipment failure, fire hazards, and hazardous gas leaks. Simultaneously, a mapping relationship is established with the emergency response policies in the IoT device policy library. For example, in the case of an electrical equipment short-circuit failure, a corresponding emergency response policy exists in the IoT device policy library. For example, a device linkage rule stipulates that the short-circuit alarm in the electrical control cabinet should immediately trigger a partial shutdown of the main power supply to prevent further damage caused by excessive short-circuit current. A parameter adjustment rule specifies the activation parameters for fire extinguishing equipment and the adjustment parameters for ventilation equipment. An execution condition constraint stipulates that maintenance personnel can only perform maintenance operations after the short-circuit current stabilizes or decreases to a certain level. This mapping relationship allows monitoring to quickly invoke the corresponding emergency response policy when an anomaly matching an electrical equipment short-circuit failure template is detected.

[0117] In one possible implementation, the method for updating the IoT device policy library includes:

[0118] Step S310: monitoring the device operation logs and abnormal event handling effect data after the IoT execution device cluster executes the control instruction set.

[0119] For example, during the handling of an electrical equipment short-circuit incident, the equipment operation log records the operation status of each executing device. The equipment operation log for fire-extinguishing equipment displays its startup time, operating power changes, and fire extinguishing agent consumption; the equipment operation log for ventilation equipment records the ventilation volume adjustment process and continuous operation time; and the equipment operation log for maintenance equipment includes the start time of the maintenance operation and parameter adjustments during the maintenance process. Abnormal event handling effect data reflects the results of the entire handling process, such as whether the electrical equipment was fully repaired, the time required for repair, and any additional impact on surrounding equipment and production.

[0120] Step S320: extract the parameter execution deviation value and equipment energy consumption data in the equipment operation log, and calculate the abnormality elimination efficiency and resource utilization rate in the treatment effect data.

[0121] For example, for fire-fighting equipment, the parameter execution deviation might be the difference between the actual operating power and the set operating power in the control instructions. If the set operating power is 10 liters of fire extinguishing agent per minute, but in actual operation, due to equipment aging or other reasons, only 8 liters of fire extinguishing agent are discharged per minute, the parameter execution deviation is -2 liters / minute. Equipment energy consumption data records the energy consumption of the fire-fighting equipment throughout its operation, such as the amount of electricity consumed or the reduction in fire extinguishing agent reserves. Regarding treatment effectiveness data, the efficiency of abnormality resolution can be measured by the time it takes for electrical equipment to resume normal operation. If a repair should normally take 10 minutes but actually takes 15 minutes, the abnormality resolution efficiency is relatively low. Resource utilization considers the effective utilization of various resources (such as energy, equipment, and manpower) during the treatment process. For example, whether the fire extinguishing agent in the fire-fighting equipment is fully utilized and whether the maintenance personnel's working time is allocated reasonably.

[0122] Step S330 , performing strategy optimization on the device linkage rule and the parameter adjustment rule in the emergency response strategy according to the parameter execution deviation value and the abnormality elimination efficiency, and generating updated device linkage rule and parameter adjustment rule.

[0123] If the parameter execution deviation of the fire extinguishing equipment is large and the abnormality elimination efficiency is low, the equipment linkage rules may need to be adjusted. For example, it was originally stipulated that the fire extinguishing equipment would be activated within 1 minute after the electrical equipment short circuit alarm, but due to the actual startup delay, the startup time can be brought forward to within 30 seconds. Regarding the parameter adjustment rules, based on the actual operation of the fire extinguishing equipment and the abnormality elimination efficiency, if it is found that the fire cannot be effectively controlled according to the original parameter settings, the operating power range of the fire extinguishing equipment can be adjusted, and the maximum operating power can be increased from 10 liters of fire extinguishing agent per minute to 12 liters of fire extinguishing agent per minute to ensure more effective fire extinguishing operations in similar electrical equipment short circuit faults.

[0124] Step S340 : Based on the device energy consumption data and the resource utilization rate, dynamically adjust the resource allocation threshold in the execution condition constraint, and recalculate the calibration coefficient of the resource consumption estimate.

[0125] For ventilation equipment, if the equipment energy consumption data shows that the energy consumption is too high during the handling of a short-circuit fault of an electrical device, it may be necessary to re-evaluate its resource allocation threshold in the execution condition constraint. For example, the power resource allocation threshold for ventilation equipment during fault handling was originally set at 30% of the total power resources, but the actual consumption exceeded this threshold and reached 40%. Based on the equipment energy consumption data and resource utilization, the resource allocation threshold can be adjusted to 35%. At the same time, the calibration coefficient of the resource consumption estimate is recalculated. For example, when calculating the resource consumption estimate of ventilation equipment, the calibration coefficient is adjusted to take into account the actual operating efficiency of the equipment, environmental factors, etc., so that the resource consumption estimate more accurately reflects the actual situation.

[0126] Step S350: Synchronize the updated device linkage rules, parameter adjustment rules, and adjusted resource allocation thresholds to the IoT device policy library to replace the original policies.

[0127] For example, optimized device linkage rules (such as advancing the activation time of fire extinguishing equipment and adjusting the linkage relationship between ventilation equipment and other equipment), parameter adjustment rules (such as adjusting the operating power range of fire extinguishing equipment and the ventilation volume parameters of ventilation equipment), and adjusted resource allocation thresholds (such as adjusting the power resource allocation threshold of ventilation equipment) can be synchronized to the IoT device policy library, replacing the original corresponding policies. This way, when encountering similar abnormal events in the future, IoT devices can use the updated policies to perform more efficient and accurate emergency response operations.

[0128] In a possible implementation, step S152 includes:

[0129] Step S1521 : screening a set of candidate backup devices having the same functional attributes as the target execution device from the IoT execution device cluster.

[0130] For example, in a factory ventilation system, suppose the target execution device is a large axial flow fan, providing ventilation services for a specific area of the workshop. Within the IoT execution device cluster, all backup ventilators with the same functional attributes as this axial flow fan are selected. These backup ventilators are similar in design, function, and application scenarios to the target ventilator and can, to a certain extent, replace the target ventilator.

[0131] Step S1522: Obtain the current working status, remaining resource capacity, and physical distance from the target monitoring area of each candidate backup device.

[0132] For example, the current operating status of these candidate backup ventilators may vary. Some may be completely idle and ready for use at any time, while others may be performing auxiliary ventilation tasks but still have some spare capacity to replace the target ventilator. Remaining resource capacity is also an important factor. For ventilators, this can be reflected in factors such as remaining power supply capacity, acceptable operating hours, and ventilation volume adjustment range. For example, a backup ventilator with a power supply that can support two hours of maximum power operation and a wide ventilation volume adjustment range indicates a relatively high remaining resource capacity. Furthermore, the physical distance from the target monitoring area (i.e., the specific area of the workshop that the target ventilator is responsible for ventilating) can also affect the feasibility of its replacement. If a backup ventilator is far from the target monitoring area, this may lead to complex ventilation duct connections or reduced ventilation efficiency. Backup ventilators that are closer have advantages in this regard.

[0133] Step S1523 : Calculate the availability score of each candidate backup device according to the current working status and the remaining resource capacity, and generate a comprehensive replacement priority list in combination with the physical distance.

[0134] For each candidate backup ventilator, a quantitative assessment is performed based on its current working status and remaining resource capacity. For example, if a backup ventilator is completely idle and has a high remaining resource capacity, it can be given a higher basic score; if it is performing some tasks but has limited remaining resource capacity, it can be given a relatively low score. This score is then combined with the physical distance factor. Assuming that the ventilators closer to the target monitoring area have higher scores on the physical distance factor, the basic score and the physical distance score are combined through a specific algorithm to obtain the availability score of each candidate backup ventilator. Based on these availability scores, a comprehensive replacement priority list is generated, and the candidate backup ventilators are sorted from high to low according to their availability scores in this comprehensive replacement priority list.

[0135] Step S1524 : Select the candidate backup device with the highest score in the comprehensive replacement priority list as the target backup device, and allocate the control parameters corresponding to the target execution device to the target backup device after scaling them according to a preset ratio.

[0136] In the ventilator example above, if a backup ventilator scores the highest in the comprehensive replacement priority list, it will be selected as the target backup device. The control parameters of the target execution device (i.e., the axial flow fan with the problem) need to be adjusted according to the actual situation of the target backup device. For example, the original operating power of the target execution device is 10 kilowatts. Due to the remaining resource capacity or performance characteristics of the target backup device, the operating power may need to be scaled according to a preset ratio. Assuming the preset ratio is 0.8, the operating power allocated to the target backup device will be adjusted to 8 kilowatts. At the same time, other control parameters of the target execution device, such as the ventilation direction, ventilation volume adjustment range, etc., will also be adjusted accordingly based on the characteristics of the target backup device.

[0137] Step S1525: Send the scaled control parameters and new execution priority to the target standby device, and update the device identification and parameter record in the control instruction set.

[0138] In this embodiment, the adjusted operating power, ventilation direction, and other control parameters, along with the new execution priority, can be sent to the target backup device, enabling it to operate according to the new instructions. Simultaneously, in the control instruction set, the device identifier originally associated with the target execution device is replaced with the identifier of the target backup device, and the corresponding parameter records are updated. This ensures that the entire control instruction set matches the actual execution device, allowing subsequent monitoring and management operations to proceed accurately and without error.

[0139] In a possible implementation, step S1214 includes:

[0140] Step S1214 - 1 : setting the step length and window size of the sliding window, and dividing the spatial semantic feature matrix into a plurality of overlapping feature sub-matrices in time sequence according to the step length and window size of the sliding window.

[0141] For example, in monitoring a production workshop, the spatial semantic feature matrix contains the spatial semantic information of the equipment, personnel, and other items within the workshop, as well as their temporal changes. The sliding window step size is set to 5 seconds, and the window size is 30 seconds. With this setting, starting from the start time of the spatial semantic feature matrix, a 30-second window is created every 5 seconds, resulting in multiple overlapping feature sub-matrices. Because the step size is 5 seconds, adjacent feature sub-matrices overlap by 25 seconds. This overlapping setting helps capture subtle changes within a continuous timeframe.

[0142] Step S1214 - 2 , performing a temporal convolution operation on each of the feature sub-matrices to extract the feature change gradients and directions of adjacent time segments.

[0143] For example, for each 30-second feature submatrix, the temporal convolution operation can analyze the characteristic changes between adjacent time segments (e.g., one segment every 5 seconds) for various monitored objects within the workshop (such as the operating status of equipment and personnel activities) within that 30-second period. For example, using a piece of automated production equipment within the workshop, the temporal convolution operation can reveal the gradient of the equipment's operating speed between adjacent time segments—that is, whether the speed is accelerating or decelerating, and whether the direction of change is positive (increasing) or negative (decreasing). For personnel activities, the temporal convolution operation can also reveal how characteristics such as the speed and direction of movement within the workshop change between adjacent time segments.

[0144] Step S1214-3: perform sequence modeling on the feature change gradient and direction through a long short-term memory network to generate trend prediction vectors of the target monitoring area in different time segments.

[0145] For example, in long-term monitoring of a workshop, the operating status of equipment and the movement patterns of personnel often exhibit certain temporal patterns. Long-Short-Term Memory (LSTM) networks can model the sequence of feature change gradients and directions, such as those obtained through temporal convolution operations, such as those for equipment speed and personnel movement directions. For example, for production equipment, the LSTM network can predict whether the equipment's speed will remain stable, gradually increase, or potentially experience a malfunction leading to a sudden drop in speed over the next few time periods based on the changing speed trends over the past period, generating corresponding trend prediction vectors. For personnel movements, trend prediction vectors can also be used to predict personnel aggregation within the workshop (e.g., whether they will gather in a specific area) and movement patterns (e.g., whether they will move along a fixed route).

[0146] Step S1214-4: concatenate the trend prediction vectors of each time segment in chronological order to form the dynamic change trend feature, wherein the dynamic change trend feature is used to characterize the temporal regularity of the movement pattern, aggregation state or morphological change of the monitored objects in the target monitoring area.

[0147] For example, within the entire monitoring cycle of a workshop, different time segments have corresponding trend prediction vectors. These trend prediction vectors are spliced together in chronological order to form a dynamic trend signature. This dynamic trend signature comprehensively reflects the temporal patterns of movement patterns, aggregation states, or changes in the morphology of monitored objects (equipment and personnel) within the target monitoring area (i.e., the production workshop). For example, this dynamic trend signature can be used to analyze the changing patterns of equipment operating status at different time periods during a day's production process, as well as the changes in personnel activity patterns and aggregation states at different work stages, providing a powerful basis for industrial plant management and the prevention of abnormal events.

[0148] Figure 2 The schematic diagram shows exemplary hardware and software components of a server 100 that can implement the concepts of the present application according to some embodiments of the present application. For example, a processor 120 can be used on the server 100 to perform the functions of the present application.

[0149] The server 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the monitoring information identification method based on the Internet of Things of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0150] For example, the server 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the server 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application may be implemented based on these program instructions. The server 100 also includes an input / output (I / O) interface 150 between the computer and other input / output devices.

[0151] For ease of explanation, only one processor is described in the server 100. However, it should be noted that the server 100 in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the server 100 performs step A and step B, it should be understood that step A and step B may also be performed jointly by two different processors or individually in one processor. For example, the first processor performs step A and the second processor performs step B, or the first processor and the second processor perform steps A and B together.

[0152] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned monitoring information identification method based on the Internet of Things is implemented.

[0153] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A monitoring information identification method based on the Internet of Things, characterized in that: The method comprises: Acquire multimodal sensing data collected in real time by multiple IoT monitoring nodes within a target monitoring area, wherein the multimodal sensing data includes image monitoring data, sound monitoring data, and environmental parameter monitoring data; Performing cross-modal feature extraction and fusion processing on the multimodal sensing data to generate a comprehensive spatiotemporal feature map corresponding to the target monitoring area; Based on a preset abnormal pattern library, dynamic pattern matching is performed on the comprehensive spatiotemporal feature map to determine at least one abnormal event type and its spatiotemporal distribution information existing in the target monitoring area; Generate, according to the abnormal event type and the spatiotemporal distribution information, a set of control instructions adapted to the abnormal event type, the set of control instructions including control parameters and execution priorities for different IoT execution devices; Distributing the control instruction set to the associated IoT execution device cluster within the target monitoring area, triggering the IoT execution device cluster to execute the control parameters according to the execution priority to eliminate the impact of the abnormal event; The method of performing dynamic pattern matching on the comprehensive spatiotemporal feature map based on a preset abnormal pattern library to determine at least one abnormal event type and its spatiotemporal distribution information existing in the target monitoring area includes: Loading multiple abnormal event templates from the preset abnormal pattern library, each of the abnormal event templates includes reference spatiotemporal features, event type labels, and impact range parameters corresponding to historical abnormal events; Calculating a similarity matrix between the comprehensive spatiotemporal feature map and the reference spatiotemporal features of each of the abnormal event templates, and screening out a set of candidate abnormal event templates having a similarity higher than a preset threshold based on the similarity matrix; Performing spatiotemporal alignment processing on each candidate abnormal event template in the candidate abnormal event template set, and mapping the timestamp of the comprehensive spatiotemporal feature map and the timestamp of the candidate abnormal event template to a unified time coordinate system; According to the feature difference after spatiotemporal alignment, a target abnormal event template that matches the comprehensive spatiotemporal feature map is determined from the candidate abnormal event template set, and the event type label and impact range parameter of the target abnormal event template are used as the abnormal event type and the spatiotemporal distribution information; The spatiotemporal distribution information includes the starting position, diffusion path and impact intensity attenuation curve of the abnormal event in the target monitoring area.

2. The monitoring information identification method based on the Internet of Things according to claim 1 is characterized in that: The cross-modal feature extraction and fusion processing of the multimodal sensing data to generate a comprehensive spatiotemporal feature map corresponding to the target monitoring area includes: Performing multi-scale spatial feature extraction on the image monitoring data to obtain a visual semantic feature vector within the target monitoring area, and performing voiceprint spectrum decomposition on the sound monitoring data to extract an acoustic event feature vector within the target monitoring area; Performing time series analysis on the environmental parameter monitoring data, extracting environmental fluctuation characteristic vectors within the target monitoring area, and constructing a spatiotemporal correlation matrix corresponding to the environmental fluctuation characteristic vectors; Inputting the visual semantic feature vector, the acoustic event feature vector, and the spatiotemporal correlation matrix into a cross-modal fusion network, and dynamically allocating the contribution weights of each modal feature in the spatiotemporal dimension through an attention weight allocation layer in the cross-modal fusion network; Based on the contribution weights, weighted concatenation is performed on the visual semantic feature vector, the acoustic event feature vector, and the spatiotemporal correlation matrix to generate the comprehensive spatiotemporal feature map; Among them, the comprehensive spatiotemporal feature map includes the spatial event distribution and environmental parameter evolution trend of the target monitoring area in a continuous time window.

3. The monitoring information identification method based on the Internet of Things according to claim 1 is characterized in that: The step of generating a control instruction set adapted to the abnormal event type according to the abnormal event type and the spatiotemporal distribution information includes: According to the abnormal event type, calling the emergency response strategy associated with the abnormal event type from the IoT device strategy library, the emergency response strategy including device linkage rules, parameter adjustment rules and execution condition constraints; Determining the optimal startup time and action duration of each execution device in the IoT execution device cluster based on the impact intensity attenuation curve in the spatiotemporal distribution information; Allocating the control parameters to each of the execution devices according to the device linkage rules and the parameter adjustment rules, the control parameters including operating power, action direction, and coverage range; In combination with the execution condition constraints and the optimal startup time, calculating an estimated resource consumption value of each of the execution devices in the target monitoring area, and optimizing and calibrating the control parameters based on the estimated resource consumption value; assigning the execution priority to the control parameter of each of the execution devices according to the priority weight of the impact intensity attenuation curve and the estimated resource consumption value, and generating the control instruction set; The step of distributing the control instruction set to an associated IoT execution device cluster within the target monitoring area, and triggering the IoT execution device cluster to execute the control parameters according to the execution priority, includes: Dividing the control instruction set into a plurality of instruction batches according to the execution priority, and sending each instruction batch to a corresponding execution device subcluster in batch order; When sending each batch of instructions, the response status of the execution device sub-cluster is monitored in real time. If it is detected that the target execution device does not return a response confirmation signal within a preset time, a backup device replacement strategy is activated to reallocate the control parameters corresponding to the target execution device to the backup execution device; After the execution device subcluster completes parameter execution of the current instruction batch, real-time feedback data of the target monitoring area is collected, and the impact intensity change rate of the abnormal event is evaluated based on the real-time feedback data; If the rate of change of the impact intensity does not reach the expected decay rate, dynamically adjusting the control parameters and the execution priority in subsequent instruction batches according to the real-time feedback data; After all the instruction batches are executed, a processing report of the abnormal event is generated and the processing report is uploaded to the Internet of Things monitoring platform for archiving.

4. The monitoring information identification method based on the Internet of Things according to claim 2 is characterized in that: The performing multi-scale spatial feature extraction on the image monitoring data to obtain a visual semantic feature vector within the target monitoring area includes: Dividing the image monitoring data into a plurality of image frame sequences at preset time intervals, and performing denoising and illumination equalization processing on each of the image frame sequences to obtain a processed image frame sequence; Using a convolutional neural network to perform hierarchical feature extraction on the processed image frame sequence to obtain local texture feature maps and global scene feature maps at different scales; Performing multi-resolution feature fusion on the local texture feature map and the global scene feature map through a spatial pyramid pooling layer to generate a fused spatial semantic feature matrix; Performing sliding window aggregation on the spatial semantic feature matrix in the time dimension to extract dynamic change trend characteristics of the target monitoring area within a continuous time period; The dynamic change trend feature is spliced with the spatial semantic feature matrix in the current time window to generate the visual semantic feature vector.

5. The monitoring information identification method based on the Internet of Things according to claim 3 is characterized in that: The method for constructing the preset abnormal pattern library includes: Collecting multiple historical case data marked as abnormal events in historical monitoring data, each of the historical case data includes multimodal perception data, event type labels, and handling records when the abnormality occurs; Performing cross-modal feature extraction on the multimodal perception data in the historical case data to generate a historical spatiotemporal feature map corresponding to each of the historical case data; Clustering the historical spatiotemporal feature maps into different abnormal event categories according to the event type labels, and generating an initial abnormal event template for each abnormal event category; Based on the treatment effect evaluation index in the treatment record, calibrate the impact range parameter in the initial abnormal event template so that the calibrated impact range parameter is negatively correlated with the treatment effect evaluation index; The calibrated abnormal event templates are stored in the preset abnormal pattern library, and a mapping relationship between the abnormal event templates and the emergency response strategies in the Internet of Things device strategy library is established.

6. The monitoring information identification method based on the Internet of Things according to claim 3 is characterized in that: The method for updating the IoT device policy library includes: Monitor the device operation logs and abnormal event handling effect data after the IoT execution device cluster executes the control instruction set; Extracting parameter execution deviation values and equipment energy consumption data from the equipment operation log, and calculating the abnormality elimination efficiency and resource utilization rate in the treatment effect data; Optimizing the device linkage rules and the parameter adjustment rules in the emergency response strategy according to the parameter execution deviation value and the abnormality elimination efficiency, and generating updated device linkage rules and parameter adjustment rules; Based on the device energy consumption data and the resource utilization rate, dynamically adjusting the resource allocation threshold in the execution condition constraint, and recalculating the calibration coefficient of the resource consumption estimate; The updated device linkage rules, parameter adjustment rules and adjusted resource allocation thresholds are synchronized to the IoT device policy library to replace the original policies.

7. The monitoring information identification method based on the Internet of Things according to claim 3 is characterized in that: The starting of the backup device replacement strategy to reallocate the control parameters corresponding to the target execution device to the backup execution device includes: Screening a set of candidate backup devices having the same functional attributes as the target execution device from the IoT execution device cluster; Obtaining the current working status, remaining resource capacity, and physical distance of each candidate backup device from the target monitoring area; Calculating the availability score of each candidate backup device based on the current working status and the remaining resource capacity, and generating a comprehensive replacement priority list in combination with the physical distance; Selecting the candidate backup device with the highest score in the comprehensive replacement priority list as the target backup device, and allocating the control parameters corresponding to the target execution device to the target backup device after scaling according to a preset ratio; The scaled control parameters and the new execution priority are sent to the target standby device, and the device identification and parameter records in the control instruction set are updated.

8. The monitoring information identification method based on the Internet of Things according to claim 4 is characterized in that: The step of performing sliding window aggregation on the spatial semantic feature matrix in a time dimension to extract dynamic change trend features of the target monitoring area within a continuous time period includes: Setting a step length and a window size of a sliding window, and dividing the spatial semantic feature matrix into a plurality of overlapping feature sub-matrices in chronological order according to the step length and the window size of the sliding window; Performing a temporal convolution operation on each of the feature submatrices to extract the feature change gradients and directions of adjacent time segments; Perform sequence modeling on the feature change gradient and direction through a long short-term memory network to generate trend prediction vectors of the target monitoring area in different time segments; Splicing the trend prediction vectors of each time segment in chronological order to form the dynamic change trend feature; The dynamic change trend feature is used to characterize the temporal regularity of the movement pattern, aggregation state or morphological change of the monitored objects in the target monitoring area.

9. A server, characterized in that: The server includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the monitoring information identification method based on the Internet of Things as described in any one of claims 1 to 8 above.

Citation Information

Patent Citations

  • Multi-mode building intelligent monitoring system based on IOT

    CN118842815A

  • Intelligent monitoring method and system based on Internet of Things, medium and program product

    CN119135741A