Intelligent monitoring method and system, embedded device and storage medium
By integrating multi-source data and dynamic 3D environment modeling, combined with time series prediction algorithms, the problem of target monitoring in complex environments was solved, achieving high-precision target positioning and motion trajectory prediction, and improving the robustness and practicality of the monitoring system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies struggle to achieve continuous and effective monitoring of targets in complex environments. In particular, when there is high-temperature interference, the difference between the target and background thermal radiation is not obvious, or the target is obscured, infrared images cannot be displayed normally, resulting in blurred and difficult-to-distinguish target features, making it difficult for the monitoring system to track the target.
Using multi-source data fusion technology, environmental depth data, visible light images, ultraviolet images, and infrared images from the fire scene are received. Through feature extraction and multimodal data fusion, a dynamic three-dimensional environment model is constructed. The target's motion trajectory is predicted by combining time series prediction algorithms, and the target is matched and located in the three-dimensional model using human feature information.
It achieves high-precision, real-time monitoring and tracking of targets in complex environments, improves the reliability and comprehensiveness of feature extraction, ensures accurate target positioning and trajectory prediction in situations such as smoke obscuring the target, and reduces misjudgment and missed judgment.
Smart Images

Figure CN121033759B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent video surveillance, and more particularly to an intelligent surveillance method, system, embedded device, and storage medium. Background Technology
[0002] In modern security monitoring and industrial surveillance, there is a high demand for accurate monitoring and tracking of targets in complex environments. For example, in fire rescue scenarios, it is necessary to promptly and accurately determine the location and movement of trapped personnel or important equipment in order to carry out rescue operations efficiently.
[0003] In related technologies, infrared imagery is often relied upon for target monitoring. Infrared sensors capture the thermal radiation information of a target, converting invisible thermal signals into visualized infrared images. Technicians then use image recognition algorithms to accurately identify and locate the target based on the thermal characteristics presented in the image, such as temperature distribution differences and thermal profile morphology.
[0004] However, when environmental conditions include high-temperature interference, insignificant differences in thermal radiation between the target and background, or complete obstruction of the target, the monitored target cannot be displayed properly in the infrared image, and its features become blurred and difficult to discern. Consequently, the monitoring system struggles to continue tracking the target, causing it to slip out of the monitoring system. This technique cannot achieve continuous and effective monitoring of targets in complex and ever-changing environments, and it fails to meet the practical application requirements for full-process tracking of monitored targets. Summary of the Invention
[0005] This application provides an intelligent monitoring method, system, embedded device, and storage medium to solve the technical problem of difficulty in accurately monitoring targets in complex environments, and to achieve high-precision, real-time monitoring and tracking of targets.
[0006] Firstly, this application provides an intelligent monitoring method applied to a monitoring system. The method includes: receiving environmental depth data, visible light images, ultraviolet images, and infrared images from a fire scene within the monitoring system; extracting features from the visible light image, the ultraviolet image, and the infrared image to obtain human feature information and initial location area features of the monitored target, wherein the human feature information includes color features, thermal features, and posture features, and the monitored target includes rescue personnel and trapped personnel; and, when the monitored target disappears from the monitoring screen, constructing a dynamic three-dimensional environmental model of the disappeared area based on the environmental depth data and the initial location area features, wherein the dynamic three-dimensional environmental model includes... The system identifies an initial smoke-covered area, a flame distribution area, and an initial location area. It then extracts the spatial distribution characteristics of the smoke in the initial smoke-covered area from the dynamic 3D environment model. Based on the initial location area characteristics, the human body feature information, and the smoke spatial distribution characteristics, a time-series prediction algorithm is used to predict the movement trajectory of the monitored target within the smoke-covered area. Based on the human body feature information, a matching area is checked in the dynamic 3D environment model to determine if it falls within the range of the movement trajectory. This matching area includes areas with similar features to the human body feature information. If so, the spatial location of the monitored target is determined based on the matching area, yielding the real-time location and direction of movement of the monitored target.
[0007] By employing the aforementioned technical solution, the system first receives environmental depth data, visible light images, ultraviolet images, and infrared images from the fire scene. These multiple data types reflect the characteristics of the target and the environment from different dimensions. Visible light images provide visual information such as the target's color and texture; ultraviolet images help detect the distribution of gases such as smoke; infrared images capture the target's thermal features; and environmental depth data acquires the three-dimensional spatial structure of the scene. Next, feature extraction is performed on these images to obtain human feature information, including color, thermal features, and posture features, as well as initial location area features. The fusion of multi-source data avoids the limitations of a single sensor. For example, when dense smoke causes infrared images to be confused by thermal features or the target's thermal signal to be obscured, making it impossible to identify the monitored target, this method, by fusing edge texture information from visible light images, gas distribution features from ultraviolet images, and spatial structure information from environmental depth data, can still achieve continuous target identification and localization through cross-validation and feature complementarity of multimodal data. The combination of multiple features makes the extracted human feature information more comprehensive and accurate. The initial location area features can also more accurately locate the target's specific location in the environment, laying a solid data foundation for subsequent monitoring and tracking in complex environments. It effectively solves the problem of insufficient information from a single sensor in complex environments, improves the reliability and comprehensiveness of feature extraction for monitored targets, and thus enables a more accurate grasp of the target's initial state, providing accurate raw data for subsequent processing in complex scenarios.
[0008] In conjunction with some embodiments of the first aspect, in some embodiments, feature extraction is performed on the visible light image, the ultraviolet image, and the infrared image to obtain human feature information and initial location area features of the monitored target. Specifically, this includes: extracting features from the visible light image, the ultraviolet image, and the infrared image to obtain image feature data; integrating the image feature data and the environmental depth data into a unified multimodal image data stream using a multimodal data fusion algorithm, wherein the multimodal image data stream fuses visible light, ultraviolet, infrared, and depth information; and extracting human feature information and initial location area features of the monitored target from the multimodal image data stream.
[0009] By employing the above technical solution, image feature data is first extracted, and then integrated with environmental depth data using a multimodal data fusion algorithm. Different images provide information from different spectra, while depth data assigns spatial coordinates. The fused data stream contains both spectral and spatial depth information. For example, visible light combined with depth data determines the target's size and location, while infrared and ultraviolet data are combined to determine the target's environment. Extracting features from the multimodal data stream leverages data complementarity, compensating for the limitations of single-modality approaches in complex environments, providing more accurate data for subsequent modeling and prediction, and improving the system's feature extraction capabilities and environmental perception accuracy.
[0010] In conjunction with some embodiments of the first aspect, in some embodiments, when the monitored target disappears from the monitoring screen, a dynamic three-dimensional environmental model of the disappearance area is constructed based on the environmental depth data and the initial location area features. Specifically, this includes: when the monitored target disappears from the monitoring screen, transforming the two-dimensional coordinates into a three-dimensional spatial coordinate system based on the environmental depth data and the visible light image to generate an initial three-dimensional environmental model; marking the initial location area in the initial three-dimensional environmental model to obtain a three-dimensional environmental model of the area where the monitored target disappears; marking the initial smoke-covered area in the three-dimensional environmental model using the ultraviolet image and the environmental depth data; and identifying the flame distribution area using the infrared image to obtain a dynamic three-dimensional environmental model in which the initial smoke-covered area and the flame distribution area are spatially superimposed.
[0011] By employing the above technical solution, an initial 3D environment model is constructed by combining environmental depth data and visible light images when the target disappears, providing a spatial framework for subsequent modeling. The initial location area is marked to clarify the starting coordinates, and smoke and flame areas are labeled using ultraviolet and infrared images respectively, forming a dynamic 3D environment model. This model reflects the environmental state and distribution of hazardous areas when the target disappears. In complex fire environments, it can update the spatial distribution of environmental factors in real time, providing accurate environmental constraints for predicting the target's trajectory, avoiding the limitations of 2D image modeling, and improving the system's ability to model complex dynamic environments.
[0012] In conjunction with some embodiments of the first aspect, in some embodiments, based on the initial location area features, the human body feature information, and the smoke spatial distribution features, a time series prediction algorithm is used to predict the movement trajectory of the monitored target within the smoke-covered area. Specifically, this includes: filtering out the passable paths of the monitored target based on the initial location area features and the smoke spatial distribution features; matching the passable paths with the human body feature information to obtain the suspected movement trend of the monitored target; using the time series prediction algorithm in conjunction with the suspected movement trend to predict the suspected movement trajectory of the monitored target in the dynamic three-dimensional environment model; and filtering the suspected movement trajectory in the dynamic three-dimensional environment model to obtain a reasonable movement trajectory of the monitored target within the smoke-covered area.
[0013] By employing the above technical solution, passable paths are screened based on initial location and smoke features, movement trends are determined by combining human characteristics, and a time-series prediction algorithm is used to generate potential movement trajectories, which are then further filtered to obtain reasonable trajectories. Initial location and smoke features eliminate impassable areas, while human characteristics infer movement direction. The time-series algorithm considers historical patterns, and the filtering process eliminates unreasonable trajectories. Compared to traditional prediction methods, this approach can more accurately simulate the movement of a target in complex environments, solving the problem of large prediction deviations under complex movement conditions and improving the accuracy and reliability of trajectory prediction.
[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the suspected motion trajectory is screened in the dynamic three-dimensional environment model to obtain a reasonable motion trajectory of the monitored target within the smoke-covered area. Specifically, this includes: extracting the spatial coordinates of the suspected motion trajectory based on the suspected motion trajectory and the dynamic three-dimensional environment model; comparing the spatial coordinates of the obstacle positions in the dynamic three-dimensional environment model with the spatial coordinates of the suspected motion trajectory to obtain overlapping trajectory points; marking the trajectory points as invalid points to obtain the remaining points in the suspected motion trajectory; and re-determining a reasonable motion trajectory based on the remaining points.
[0015] By employing the aforementioned technical solution, the spatial coordinates of suspected motion trajectories and obstacles are extracted and compared. Overlapping invalid points are marked, and a reasonable trajectory is determined based on the remaining points. In complex fire scenes, traditional methods easily generate unreasonable trajectories that pass through obstacles. This step, through precise comparison, eliminates invalid points, ensuring that the trajectory conforms to the constraints of the physical environment. The filtered trajectory more realistically reflects the target's movement path, avoids prediction errors, improves the rationality and practicality of the trajectory, and provides a reliable basis for target localization.
[0016] In conjunction with some embodiments of the first aspect, in some embodiments, based on the human body feature information, checking whether the matching region is within the range of the motion trajectory in the dynamic three-dimensional environment model specifically includes: expanding the motion trajectory into a volume space with a region to obtain the range of the motion trajectory; marking the range of the motion trajectory in the dynamic three-dimensional environment model to obtain a marked three-dimensional environment model; based on the human body feature information, filtering matching regions with similar features in the marked three-dimensional environment model; and determining whether the matching region is within the range of the motion trajectory through positioning information.
[0017] By adopting the above technical solution, the motion trajectory is expanded into a volumetric space and labeled into a 3D environment model. Matching regions are then selected based on human features, and it is determined whether they fall within the trajectory range. Expanding the trajectory into a volumetric space takes into account the actual volume of the target and motion deviations, improving matching tolerance. Combining human feature selection avoids false positives, ensuring that the matching region both conforms to the target characteristics and falls within a reasonable trajectory range. This improves the accuracy and reliability of target positioning, reduces false positives and false negatives, and makes the monitoring system more stable and effective in complex environments.
[0018] In conjunction with some embodiments of the first aspect, in some embodiments, after the step of checking whether the matching area is within the range of the motion trajectory in the dynamic three-dimensional environment model based on the human feature information, the method further includes: if not, calculating the distance value from the center point of the matching area to the range of the motion trajectory based on the center coordinates of the matching area; sorting the matching areas according to the distance value to obtain the matching area with the smallest distance; and obtaining the real-time position and motion direction of the monitored target based on the positioning information of the matching area with the smallest distance.
[0019] By employing the above technical solution, when the matching area is outside the trajectory range, the distance from it to the trajectory range is calculated, sorted, and the area with the smallest distance is selected to determine the target position and direction. In complex environments, the matching area may deviate from the trajectory. This step can handle such deviations, selecting the most reasonable candidate position through distance sorting, avoiding target loss, improving the system's robustness in uncertain environments, ensuring timely and accurate target information in emergency scenarios, and providing support for rescue operations.
[0020] In a second aspect, a monitoring system includes one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, wherein the one or more processors invoke the computer instructions to cause the monitoring system to perform a method as described in the first aspect and any possible implementation thereof.
[0021] Thirdly, an embedded device, when running on a monitoring system, causes the monitoring system to perform the method described in the first aspect and any possible implementation thereof.
[0022] Fourthly, a computer-readable storage medium storing computer instructions that, when executed on a monitoring system, cause the monitoring system to perform the method described in the first aspect and any possible implementation thereof.
[0023] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0024] 1. By employing multi-source data fusion and multi-modal feature extraction techniques (fusing visible light, ultraviolet, and infrared images with environmental depth data to extract human information including color, thermal features, and posture features, as well as initial position area features), visible light images provide visual features such as target texture and color; ultraviolet images can detect smoke distribution; infrared images capture target thermal signals; and environmental depth data provides spatial coordinates. Multi-source data fusion integrates this information from different dimensions, effectively solving the technical problems of information loss and inaccurate feature recognition in complex environments such as smoke obstruction by a single visible light sensor in existing technologies. This enables accurate extraction of multi-dimensional features of the monitored target and reliable initial position localization.
[0025] 2. Employing a dynamic 3D environment modeling technology based on multimodal data (combining environmental depth data and visible light images to construct a 3D spatial coordinate system, overlaying smoke areas marked by ultraviolet images and flame distribution areas identified by infrared images), the system first uses environmental depth data and visible light images to convert 2D information into a 3D spatial coordinate system, constructing an initial 3D environment model and building a spatial framework for subsequent modeling. Next, based on the sensitivity of ultraviolet images to smoke, the initial smoke-covered area is marked using environmental depth data to clarify the distribution range of smoke in 3D space; the spatial location of high-temperature hazard areas is determined by identifying flame distribution areas through infrared images. The smoke, flame areas, and the initial target location area are then overlaid in the 3D environment model to form a dynamic 3D environment model. This effectively solves the technical problems of existing technologies that rely on 2D image modeling, resulting in incomplete environmental information and an inability to accurately reflect the distribution of complex spatial obstacles and hazard areas. It achieves three-dimensional modeling of smoke-covered areas, flame distribution areas, and the initial target location at a fire scene, providing realistic and dynamic 3D environmental constraints for target trajectory prediction and improving the system's environmental perception capabilities in complex scenes.
[0026] 3. By employing techniques of motion trajectory volumetric spatial expansion and multimodal feature matching (converting the motion trajectory into a volumetric model encompassing a spatial range, and combining human body color, thermal features, and posture features to filter matching regions and determine trajectory correlation), the motion trajectory is expanded into a volumetric space. This fully considers the actual space occupied by the target and deviations during movement, thus broadening the search range. In the dynamic 3D environment model, regions that may match the target are selected by combining human body color, thermal features, and posture features. These features describe the target's characteristics from different angles, reducing the impact of interference factors such as smoke on target recognition. By determining whether the matching region is within the expanded motion trajectory volumetric range, its correlation with the target's motion trajectory is verified. When the target is obscured by factors such as smoke, resulting in incomplete sensor data, this 3D search and feature correlation verification method effectively solves the technical problems of existing technologies where single target feature matching is easily affected by smoke interference and the positioning range is limited to the trajectory line, leading to missed detections. This enables 3D search and feature correlation verification of potential target locations in complex environments such as smoke obscuration, improving the fault tolerance and accuracy of target positioning and reducing misjudgments and missed detections caused by environmental noise or incomplete sensor data. Attached Figure Description
[0027] Figure 1 This is an exemplary scenario diagram of monitoring a fire scene in an embodiment of this application;
[0028] Figure 2 This is a flowchart illustrating an intelligent monitoring method in an embodiment of this application;
[0029] Figure 3 This is an exemplary scenario diagram illustrating the prediction of the movement trajectory of a monitored target within a smoke-covered area in this application embodiment;
[0030] Figure 4 This is an exemplary scenario diagram illustrating the search for and monitoring of regions that match the feature information of the target in this application.
[0031] Figure 5 This is another flowchart illustrating the intelligent monitoring method in the embodiments of this application;
[0032] Figure 6 This is a schematic diagram of the hardware architecture of an electronic device in an embodiment of this application. Detailed Implementation
[0033] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0034] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0035] Figure 1 This is an exemplary scenario diagram of monitoring a fire scene in an embodiment of this application.
[0036] Please see Figure 1 At the fire scene, drones were used to intelligently monitor the situation.
[0037] Existing technologies utilize a single type of sensor to acquire target-related information. For example, a visible light camera is used to capture images of the target, and simple edge detection algorithms are applied to the visible light images to identify the general outline of the target and attempt to determine its location. The determination of the target's trajectory is based on inferences made from simple changes in the target's position in adjacent frames. For instance, the displacement of the target on the image plane is calculated, and the velocity is estimated based on time intervals to predict its possible subsequent location.
[0038] The intelligent monitoring method employed in this application receives environmental depth data, visible light images, ultraviolet images, and infrared images from the fire scene. These multiple data types reflect the characteristics of the target and environment from different dimensions. Visible light images provide visual information such as the target's color and texture; ultraviolet images help detect the distribution of gases such as smoke; infrared images capture the target's thermal characteristics; and environmental depth data reveals the three-dimensional spatial structure of the scene. Feature extraction is then performed on these images to obtain human feature information, including color, thermal features, and posture features, as well as initial position area features. The fusion of multi-source data avoids the limitations of a single sensor. The combination of multiple features makes the extracted human feature information more comprehensive and accurate, laying a solid data foundation for subsequent monitoring and tracking in complex environments. This effectively solves the problem of insufficient information from a single sensor in complex environments, improving the reliability and comprehensiveness of feature extraction for monitored targets.
[0039] The following is in conjunction with the above. Figure 1 The schematic diagram shown illustrates the intelligent monitoring method in this application embodiment: receiving environmental depth data, visible light images, ultraviolet images, and infrared images from the fire scene; extracting human characteristics such as color, heat, and posture of the monitored targets (rescuers and trapped personnel) as well as initial position area features; when the target disappears from the screen, constructing a dynamic three-dimensional environment model including smoke obstruction, flame distribution, and the target's initial position based on the depth data and initial position; extracting the spatial distribution features of the smoke; combining multiple features and using a time series algorithm to predict the target's trajectory in the smoke area; and finally, checking whether the matching area of similar features in the three-dimensional model is within the trajectory range based on human characteristics, thereby achieving accurate monitoring and real-time positioning of targets in complex environments.
[0040] This embodiment involves a visible light camera, an ultraviolet sensor, an infrared sensor, and a depth sensor. The following is a detailed description of these hardware devices:
[0041] Visible light cameras capture visible light images of fire scenes, containing rich color, texture, and shape information, which can be used to identify the color of rescuers' clothing, the approximate outline of trapped people, etc. In the feature extraction stage, convolutional neural networks are used to extract the texture features of the target (such as the texture of clothing fabrics), color blocks (such as the red area of flames), and human contour pose, providing visual semantic information for multimodal fusion.
[0042] Ultraviolet (UV) sensor: Acquires UV images, which can clearly show the location, shape, and spread trend of flames. Especially in environments where smoke obscures visible light, the UV image, despite its weak penetration, is sensitive to flames, aiding in the location of the core area of the fire source. UV images can also sensitively detect the presence and distribution of smoke. By thresholding (setting a smoke grayscale threshold of 150-255) to mark smoke areas, the geometric features (area, perimeter, eccentricity) and texture features (grayscale co-occurrence matrix) of smoke clumps are extracted, providing a two-dimensional contour basis for 3D smoke modeling.
[0043] Infrared sensors acquire infrared images, which capture the infrared radiation emitted by objects and convert temperature differences into a visual image. Infrared images directly reflect the surface temperature of objects and can identify high-temperature areas (such as flames, overheated equipment), as well as the thermal signals of humans or animals (body temperature is higher than ambient temperature). Especially in dense smoke or dark environments, targets can be located through thermal signals without relying on visible light. They also capture the thermal characteristics of targets (such as human body temperature, high-temperature areas of flames), identifying target locations through thermal signals when obscured by smoke. Employing thermal imaging technology (such as uncooled microbolometers), they output a temperature matrix with a resolution typically of 640×512, supporting a temperature range of -20℃ to 1500℃, and capable of identifying temperature differences as small as 0.1℃.
[0044] Depth sensor: Provides environmental depth data, acquires depth information (distance from the sensor) of each point in the scene, maps two-dimensional image coordinates to a three-dimensional spatial coordinate system, and outputs point cloud data (X, Y, Z coordinates) with an accuracy of up to millimeters and a frame rate synchronized with visible light cameras.
[0045] Figure 2 This is a flowchart illustrating an intelligent monitoring method in an embodiment of this application.
[0046] Please see Figure 2 The intelligent monitoring method is described in detail below:
[0047] 201. Receive environmental depth data, visible light images, ultraviolet images, and infrared images from the fire scene in the monitoring system;
[0048] Visible light images display conventional visual information of the monitored scene, including color features (such as clothing color and skin tone) and posture features (such as movements and limb shapes) of the monitored targets (rescuers, trapped persons), as well as visible details of the surrounding environment (such as object outlines and scene layout), providing a foundation for extracting human appearance and posture information. Ultraviolet (UV) images display the radiation characteristics of flames, high-temperature objects, or discharge phenomena in the UV band at a fire scene. This helps identify the UV reflection and radiation characteristics of flame distribution areas and smoke environments, and is used to construct the flame distribution portion in a dynamic 3D environment model, indirectly providing environmental constraints for location. Infrared (IR) images capture differences in the thermal radiation of objects, displaying the thermal characteristics of the monitored targets (such as the temperature comparison between human body temperature and the surrounding environment) and the environmental temperature distribution. This is used to distinguish personnel from the background, identify heat sources (such as human bodies) hidden in smoke, and provide crucial data for thermal feature extraction and motion trajectory prediction.
[0049] The system is equipped with visible light cameras, ultraviolet sensors, infrared sensors, and depth sensors. The visible light cameras can capture visible light images of the fire scene, which contain rich color, texture, and shape information, and can be used to identify the color of rescuers' clothing, the approximate outline of trapped people, etc.
[0050] In terms of data transmission, data collected by various sensors is transmitted in real time to the system's central processing unit via wired or wireless communication modules. Wired transmission, such as Ethernet, features high stability and fast transmission speed, making it suitable for connecting fixed monitoring equipment to the central processing unit. Wireless transmission, such as Wi-Fi and 4G / 5G, can meet the real-time data transmission needs of mobile monitoring equipment such as drones, ensuring that the system can obtain the latest data in a timely manner in complex fire scene environments, regardless of the location of the monitoring equipment.
[0051] When receiving data, the system performs preliminary preprocessing, including data format conversion and timestamp marking, to ensure consistency in time and space between different types of data, laying the foundation for subsequent multimodal data fusion. For example, although visible light images, ultraviolet images, infrared images, and environmental depth data are collected by different sensors, they correspond to the same fire scene. By using timestamp marking, the system can accurately correlate various types of data at the same time, avoiding analysis errors caused by time asynchrony.
[0052] 202. Perform feature extraction on the visible light image, the ultraviolet image and the infrared image to obtain human body feature information and initial position area features of the monitored target. The human body feature information includes color features, thermal features and posture features. The monitored target includes rescue personnel and trapped personnel.
[0053] The system performs single-modal feature extraction on visible light, ultraviolet, and infrared images respectively. For visible light images, the system employs algorithms such as convolutional neural networks (CNNs) from deep learning to extract target features such as color (e.g., the specific color of rescuers' clothing, the color of trapped people's clothing), texture features (e.g., the texture structure of clothing, skin texture), and shape features (e.g., the outline of the human body, the approximate outline of the posture). For ultraviolet images, the system mainly extracts the distribution features of smoke, such as the edge contour of the smoke area and the gradient changes in smoke concentration. These features help determine the extent and shape of the smoke-obscured area. For infrared images, the system utilizes thermal imaging feature extraction algorithms to obtain the thermal features of the target, such as the temperature distribution of the human body and the distribution of high-temperature areas of flames. Thermal features have unique advantages in smoke-obscured environments, enabling the system to accurately identify the location of people and flames even when the visible light image cannot clearly display the target.
[0054] 203. When the monitored target disappears from the monitoring screen, a dynamic three-dimensional environment model of the disappearance area is constructed based on the environmental depth data and the initial position area features. The dynamic three-dimensional environment model includes the initial smoke obscuring area, the flame distribution area and the initial position area.
[0055] The disappearance of a monitored target typically refers to situations where it cannot be directly detected or identified in visible light, ultraviolet, or infrared images, especially in scenarios where infrared images fail. Specific examples include: high-concentration smoke strongly scatters and absorbs infrared light, causing human thermal signals to be masked by smoke noise; strong infrared radiation from flames or hot objects in high-temperature environments obscuring human thermal characteristics; rescuers wearing infrared camouflage materials or high-insulation equipment, or trapped individuals in extremely low-temperature areas, resulting in human thermal radiation signals below the infrared detection threshold; the target being completely obscured by physical obstacles or entering a monitoring blind spot; and the high overlap between the ultraviolet or infrared radiation characteristics of flames, high-temperature smoke, and the human body, making it impossible for multimodal data to distinguish the target from the environment. In such cases, it is necessary to utilize environmental depth data and historical location features to construct a dynamic three-dimensional environmental model to further predict the target's trajectory.
[0056] The system transforms two-dimensional coordinates into a three-dimensional spatial coordinate system based on environmental depth data and visible light images to generate an initial three-dimensional environment model. Environmental depth data provides depth information for each point in the scene (i.e., the distance to the distance sensor), while visible light images provide two-dimensional visual information about the scene. Through camera calibration technology, the system establishes a mapping relationship between two-dimensional image coordinates and three-dimensional spatial coordinates, mapping each pixel in the visible light image to a specific location (X, Y, Z) in three-dimensional space.
[0057] After generating the initial 3D environment model, the system marks the initial location area where the monitored target disappeared within the initial 3D environment model based on the extracted initial location area features. For example, if the monitored target was located in a corner of a room before disappearing, the system will accurately mark the 3D coordinate range of that corner in the 3D environment model, forming the initial location area. This marking process provides a clear starting point for the target in subsequent modeling, ensuring that the model accurately reflects the specific location where the target disappeared.
[0058] Next, the system uses ultraviolet images and environmental depth data to annotate the initial smoke-occupied areas in the 3D environment model. Ultraviolet images can sensitively detect the presence of smoke. Through image segmentation algorithms, the system can extract the two-dimensional contour of the smoke from the ultraviolet images. Then, combined with environmental depth data, the two-dimensional smoke contour is extended into three-dimensional space to determine the volume range and distribution pattern of the smoke in three-dimensional space.
[0059] Simultaneously, the system identifies the flame distribution area using infrared images. Infrared images can clearly show the location of high-temperature objects, and flames, as high-temperature areas, exhibit obvious high-brightness characteristics in infrared images. The system uses target detection algorithms to identify the location and extent of the flame from the infrared images, and then combines this with environmental depth data to map the flame area into three-dimensional space, determining the specific location and size of the flame in the three-dimensional environment.
[0060] 1. Extract the spatial distribution characteristics of the smoke in the initial smoke-covered area from the dynamic three-dimensional environment model;
[0061] The system performs 3D data analysis on the initial smoke-covered area in a dynamic 3D environment model. This area is constructed by combining ultraviolet images with environmental depth data. The ultraviolet images provide the 2D contour information of the smoke, while the environmental depth data provides its 3D spatial coordinates. The system uses voxelization technology to divide the smoke-covered area into a dense 3D mesh. Each voxel corresponds to a tiny 3D spatial unit, storing attributes such as smoke concentration, temperature, and flow direction within that unit.
[0062] Next, the system extracts the spatial geometric features of the smoke, including the volume, surface area, and spatial occupancy of the smoke region (e.g., maximum and minimum coordinates on the X, Y, and Z axes), as well as its relative position to surrounding obstacles (e.g., walls, beams, and columns). The system also analyzes the dynamic diffusion characteristics of the smoke, monitoring boundary changes in the smoke region using continuously received ultraviolet images and depth data, and calculating the smoke's diffusion speed, diffusion direction, and concentration change rate.
[0063] 205. Based on the initial location area features, the human body feature information, and the smoke spatial distribution features, a time series prediction algorithm is used to predict the movement trajectory of the monitored target within the smoke-covered area;
[0064] Please see Figure 3 This is an exemplary scenario diagram illustrating the prediction of the movement trajectory of a monitored target within a smoke-covered area in this application embodiment.
[0065] The system filters passable paths for monitored targets based on initial location area features and smoke spatial distribution features. The initial location area marks the three-dimensional coordinate range when the target disappears, while the smoke spatial distribution features clarify the area and concentration of smoke obscured, in order to exclude paths that are completely blocked by smoke or contain high-risk areas such as flames.
[0066] Next, the system matches passable paths with human feature information to deduce the suspected movement trend of the monitored target. Posture features within human features directly reflect the target's movement state, while color and thermal features are used to distinguish different people, allowing the system to further optimize movement trend judgment. After generating suspected movement trends, the system uses a time-series prediction algorithm (Long Short-Term Memory network) to model the target trajectory. These algorithms can capture the time dependence of target movement, combining historical movement data and current environmental features to predict position coordinates within future time periods. For example, an LSTM-based prediction model takes the target's historical movement speed and direction, as well as the geometric features of the current passable path, as input, and outputs possible position coordinates every 0.5 seconds within the next 5 seconds.
[0067] 206. Based on the human body feature information, check whether the matching region is within the range of the motion trajectory in the dynamic three-dimensional environment model, wherein the matching region includes regions with similar features to the human body feature information;
[0068] Please see Figure 4 This is an exemplary scenario diagram illustrating the search for and monitoring of regions that match target feature information in this application embodiment.
[0069] The system fully utilizes human characteristics, such as height, weight, movement speed, and posture (standing or crouching), to perform a detailed search within a constantly changing 3D environment model. Through analysis of multimodal data, it identifies areas in the environment that resemble human features.
[0070] The system compares these found similar regions with pre-predicted human movement trajectories. It checks whether these matching regions fall within the area covered by the human movement trajectory, that is, whether they are located in the space that the human movement path might involve. If the matching region is within the trajectory range, it indicates a strong correlation between the two in dynamic space, which is very likely the location of the target human.
[0071] If the matching area is within the range of the motion trajectory, proceed to step 207.
[0072] If the matching area is not within the range of the motion trajectory, the real-time position and direction of the monitored target are determined by calculating and sorting the distance from the center point of the matching area to the motion trajectory, and taking the positioning information of the area with the smallest distance.
[0073] 207. Based on the matching area, locate the spatial position of the monitored target to obtain the real-time position and direction of movement of the monitored target.
[0074] When the matching area is determined to be within the movement trajectory range, the system directly determines the real-time position and movement direction of the monitored target based on the positioning information of that area. For example, if the center point coordinates of the matching area are (10.5, 5.2, 1.8), and the area is continuously distributed along the positive X-axis in the trajectory volume space, then it is inferred that the target is moving in the northeast direction, and the real-time position is the current center point coordinates.
[0075] The intelligent monitoring method described in this application integrates visible light, ultraviolet, and infrared images and environmental depth data through multi-source data fusion technology. Combined with dynamic 3D environment modeling and time-series prediction algorithms, it achieves high-precision real-time positioning and trajectory prediction of monitored targets (rescuers and trapped individuals) in complex fire scenarios. This not only solves the technical challenges of traditional single sensors in smoke-covered and complex lighting environments, resulting in information loss and large positioning deviations, but also significantly improves the system's dynamic perception capability of target movement trends through multimodal feature matching and trajectory volume space verification mechanisms. This ensures that even when the target briefly leaves the monitoring frame, it can still be continuously tracked based on environmental constraints and historical features, providing reliable technical support for emergency rescue scenarios and effectively enhancing the robustness and practicality of the monitoring system in extreme environments.
[0076] The following is in conjunction with the above. Figure 1 The schematic diagram shown illustrates another intelligent monitoring method in this application embodiment: receiving multi-source data from the fire scene, extracting human features and initial position through multimodal fusion; after the target disappears, constructing a dynamic three-dimensional environment model, extracting smoke features, combining multiple features and using time series algorithms to predict the trajectory, filtering out invalid trajectory points containing obstacles; then expanding the trajectory into a volume space, labeling the model and filtering matching areas, and determining the real-time position and direction of movement of the monitored target based on whether it is within the trajectory range.
[0077] Figure 5 This is another flowchart illustrating the intelligent monitoring method in the embodiments of this application.
[0078] Please see Figure 5 The other intelligent monitoring method is described in detail below:
[0079] 501. Receive environmental depth data, visible light images, ultraviolet images, and infrared images from the fire scene in the monitoring system (this step has been described in 201 and will not be repeated here).
[0080] 502. Extract features from the visible light image, the ultraviolet image, and the infrared image to obtain image feature data;
[0081] The system employs a hierarchical feature extraction strategy tailored to the physical characteristics of images of different modalities. In visible light image processing, a convolutional neural network is first used to extract multiple layers of features: shallow convolutional layers capture low-level features such as edges and corners, middle-level networks extract textures (such as the patterns of clothing fabrics) and color blocks (such as the red areas of flames), and high-level semantic layers identify human contours, poses, and movements through object detection algorithms.
[0082] The core of ultraviolet imaging is smoke feature extraction. The system first marks the smoke region by threshold segmentation (setting the smoke grayscale threshold to 150-255), and then uses connected component analysis to calculate geometric features: the area, perimeter, and eccentricity of the smoke clumps (measuring the regularity of the shape), as well as texture features (contrast and entropy value of the grayscale co-occurrence matrix).
[0083] The system binds features of the same target across different modalities through spatiotemporal registration. For example, the pixel coordinates (500, 600) of the human contour center detected in a visible light image are mapped to three-dimensional spatial coordinates (X=6.2 m, Y=4.8 m, Z=1.6 m) using depth data. Simultaneously, an infrared image detects a thermal signal of 36.8℃ near these coordinates, and an ultraviolet image shows a moderate smoke concentration in the area (grayscale value 180), forming a multimodal feature vector containing color (RGB histogram), thermal radiation (0.9 W / m²), smoke concentration (180), and three-dimensional coordinates. The system verifies the effectiveness of the features and eliminates redundant information through mutual information analysis.
[0084] 503. By using a multimodal data fusion algorithm, the image feature data and the environmental depth data are integrated into a unified multimodal image data stream, wherein the multimodal image data stream integrates visible light, ultraviolet, infrared and depth information;
[0085] First, camera calibration technology is used to establish the intrinsic and extrinsic parameter matrices of each sensor, converting the two-dimensional image coordinates into three-dimensional world coordinates. The specific process is as follows: The visible light, ultraviolet, and infrared cameras are calibrated to obtain the rotation matrix R and translation vector T, converting the pixel coordinates (u,v) into three-dimensional coordinates (Xc,Yc,Zc) in the camera coordinate system. Then, through world coordinate system origin calibration, the global coordinates (Xw,Yw,Zw) are obtained. The depth sensor directly outputs point cloud data in the world coordinate system, which is aligned with the image data using a timestamp synchronization algorithm (dynamic time warping) to ensure that each frame of multimodal data corresponds to the same spatiotemporal scene.
[0086] During the feature fusion stage, the system employs a cross-modal attention mechanism to enhance feature complementarity. Visible light color features (such as the orange of a rescue suit) and infrared thermal features (human body temperature) are dynamically allocated through attention weights: in areas heavily obscured by smoke, the attention weight of infrared thermal features is increased to 70%, dominating target recognition; in clearly visible areas, the weight of visible light color features is increased to 60%, improving detail resolution. During the fusion process, the three-dimensional coordinates provided by the depth data serve as spatial anchors, mapping each modal feature to a unified spatial coordinate system, forming data units that contain spectral, thermal (temperature), geometric, and spatial information.
[0087] The decision-level fusion employs Bayesian inference, enhancing target detection confidence through joint probability calculation. For example, when a visible light image detects a suspected human silhouette (confidence 0.6), an infrared image detects human thermal signals in the same area (confidence 0.8), and depth data shows that the area's height matches human dimensions (confidence 0.7), the system calculates a joint probability of 0.6 × 0.8 × 0.7 ÷ (0.6 × 0.8 + 0.8 × 0.7 + 0.6 × 0.7 - 0.6 × 0.8 × 0.7) = 0.91, significantly higher than the single-modal detection result. The fused multimodal data stream possesses powerful environmental description capabilities: it can accurately determine the real-time location of rescue personnel, construct a three-dimensional smoke diffusion model, and mark high-temperature hazard areas of flames.
[0088] 504. Extract human body feature information and initial location region features of the monitored target from the multimodal image data stream;
[0089] First, for the fused data stream, the system adopts a layered processing strategy: first, it uses a convolutional neural network to extract features from the visible light image, separating the color features of the target (such as the RGB value distribution of orange on the rescuer's uniform and the texture gradient of the trapped person's clothing); then, it performs thermal imaging analysis on the infrared image, uses a Gaussian mixture model to segment the human body thermal signal, and extracts thermal features such as thermal center coordinates and thermal radiation intensity; finally, it uses a pose estimation algorithm combined with depth data to obtain the three-dimensional coordinates of the human body's joints, calculates the joint angles to determine the pose features.
[0090] In the initial location region feature extraction, the system first locates the 2D pixel coordinates of the last frame before the target disappears, maps the depth data to 3D space, and constructs a bounding box with the target's centroid as the origin (for example, the bounding box for an adult target is set as a cuboid of 1.8m × 0.5m × 0.3m). Simultaneously, the system queries environmental elements within 1.5 meters of this bounding box, such as wall coordinates and passageway directions. For example, if the target disappears at a corridor corner, the system records the 3D coordinate range of that area (X: 4.2-5.1m, Y: 2.8-3.6m, Z: 0.5-2.2m), and marks the east side as an open passage and the west side as a wall obstacle, providing spatial constraints for subsequent modeling.
[0091] 505. When the monitored target disappears from the monitoring screen, based on the environmental depth data and the visible light image, the two-dimensional coordinates are transformed into a three-dimensional spatial coordinate system to generate an initial three-dimensional environment model.
[0092] When the monitored target does not appear in the frame for three consecutive frames, the system triggers the 3D modeling module. Through camera calibration technology, the system establishes a mapping relationship between two-dimensional image coordinates and three-dimensional spatial coordinates, mapping each pixel in the visible light image to a specific location in three-dimensional space.
[0093] Taking a fire scene as an example, the system processes multi-frame data collected by drones: for each visible light image, pixels (u,v) are used to calculate world coordinates (Xw,Yw,Zw) using the depth value z, generating point cloud data that includes room layout and staircase location. After voxelization and downsampling, a mesh is generated, clearly showing the spatial location of corridors, doors, and windows, providing a millimeter-level precision 3D spatial framework for subsequent target location marking. The system also registers continuous frame point clouds to ensure that model deviations caused by device movement are less than 5 centimeters, guaranteeing the spatiotemporal consistency of the environmental model.
[0094] 506. Mark the initial location area in the initial three-dimensional environment model to obtain a three-dimensional environment model of the area where the monitored target disappears;
[0095] The system directly maps the target centroid coordinates (e.g., X=7.3m, Y=4.1m, Z=1.7m) and bounding box parameters extracted in step 504 to the world coordinate system of the initial 3D model in step 505. Centered on these coordinates, a cuboid region (length × width × height = 1.2m × 0.8m × 2.0m) encompassing the range of human activity is expanded and highlighted in red. Simultaneously, the system retrieves environmental features within this region: if the bounding box overlaps with an open door (coordinates X=7.0-7.6m, Y=4.0m, Z=0-2.3m), the region is marked as a "passable exit," and the door's orientation (negative Y-axis direction) and passage width are recorded; if obstacles exist nearby (e.g., a pillar with coordinates X=6.5m, Y=4.5m, Z=0-3m), an obstacle coordinate matrix is generated for collision detection during subsequent trajectory prediction.
[0096] 507. Using the ultraviolet image and the environmental depth data, mark the initial smoke-covered area in the three-dimensional environment model;
[0097] The system first preprocesses the ultraviolet image, uses an adaptive threshold segmentation algorithm to extract the two-dimensional contour of the smoke, and combines morphological operations to remove noise, generating a smoke mask image. For each pixel in the mask image, the corresponding three-dimensional coordinates are obtained through depth data, and voxelization technology is used to expand the two-dimensional contour into a cubic mesh in three-dimensional space (voxel size 0.1m×0.1m×0.1m) to construct the initial volume model of the smoke.
[0098] The system assigns a concentration value (0-100%) to each voxel based on the grayscale value of the ultraviolet image (reflecting the smoke concentration), and combines this with the temperature data from the infrared image to mark areas with a concentration >70% and a temperature >50℃ as "high-risk smoke areas".
[0099] 508. By identifying the flame distribution area through the infrared image, a dynamic three-dimensional environment model is obtained by spatially superimposing the initial smoke-covered area and the flame distribution area.
[0100] The system preprocesses the infrared image, separates high-temperature regions through adaptive threshold segmentation, and eliminates noise points at the flame edges using morphological closing operations to generate a two-dimensional contour mask for the flame. Since flames appear as high-brightness pixels in infrared images (temperatures typically above 500℃), the system can quickly locate suspected flame regions by setting a temperature threshold (e.g., ≥300℃). A second verification is then performed using a deep learning model (e.g., an improved Faster R-CNN) to reduce the false detection rate (e.g., excluding interference from high-temperature metal objects).
[0101] The system maps the two-dimensional flame outline to three-dimensional space. For each flame pixel (u, v), its corresponding world coordinates (X, Y, Z) are obtained through depth data, and a 0.5m × 0.5m × 0.8m cube (simulating the three-dimensional combustion range of the flame) is extended from this point. If a high-temperature signal is detected in the same area for multiple consecutive frames and the area expands, the system determines it to be a stable burning flame region, uses Kalman filtering to predict its three-dimensional expansion trend (e.g., the flame height increases at 0.5m / s), and dynamically updates the volume range of the flame.
[0102] During the overlay process, the system performs spatial Boolean operations on the flame distribution area, the initial smoke-covered area, and the initial location area: for overlapping areas that contain both smoke and flame (such as the flame burning area below the smoke), they are marked as "extremely high-risk areas" and given the highest priority access restrictions (prohibiting any path from crossing); for areas that contain only flame, they are rendered as red semi-transparent bodies with flashing warnings at the edges; for areas where smoke and flame are adjacent but do not overlap, the system calculates the safe distance between them (such as the area within 1.5m of the flame perimeter being a no-passage zone).
[0103] 509. Extract the spatial distribution features of the smoke in the initial smoke-covered area from the dynamic three-dimensional environment model (this step has been explained in 204 and will not be repeated here).
[0104] 510. Based on the initial location area features and the spatial distribution features of the smoke, select the passable paths of the monitored target;
[0105] The system constructs a gridded passage cost map within the 3D environment model, starting from the centroid of the initial location region. Each voxel (resolution 0.2m × 0.2m × 0.2m) is assigned a passage cost: flame areas have an infinite cost (completely impassable), high-concentration smoke areas (concentration > 80%) have a cost of 100, medium-low concentration smoke areas (30%-80%) have a cost of 50, and areas without smoke or obstacles have a cost of 10. Obstacles (such as walls and pillars) also have an infinite voxel cost.
[0106] During the screening process, the system simultaneously analyzes the dynamic diffusion characteristics of the smoke. If smoke is detected spreading towards the initial location area at a speed of 0.5 m / s, the system updates the passage cost of the affected area in real time, prioritizing paths that move against the direction of smoke diffusion.
[0107] 511. Match the passable path with the human feature information to obtain the suspected movement trend of the monitored target;
[0108] The system matches the spatial attributes (such as direction) of each passable path with human posture features to obtain the possible direction of movement of the search target.
[0109] Secondly, the system distinguishes target types (rescuers and trapped personnel) by thermal and color characteristics. The thermal characteristics of rescuers typically include equipment heat sources (such as the heating point of oxygen cylinders), and the color characteristic is highly visible uniforms. Their movement trend is more likely to point near the fire source (performing firefighting tasks) or to areas where trapped personnel are gathered. The thermal characteristics of trapped personnel are singular (human body temperature), while their color characteristics are diverse. Their movement trend is more likely to move towards exits or well-ventilated areas.
[0110] 512. Using a time series prediction algorithm combined with the suspected motion trend, the suspected motion trajectory of the monitored target in the dynamic three-dimensional environment model is predicted;
[0111] After determining the suspected movement trend of the target, the system uses time series prediction algorithms (such as LSTM neural networks) to generate multiple sets of possible movement trajectories, and then filters them in combination with a three-dimensional environment model.
[0112] The system collects historical motion data of the target before it disappears (including at least the 3D coordinates, velocity, and acceleration of the first 5 frames) and constructs a time-series input matrix. For rescuers, the historical data may contain a periodic movement pattern of "going back and forth between the fire source and the supply area"; for trapped personnel, it may show a linear trend of "moving along the wall to find an exit". At the same time, the geometric parameters of the currently passable path (such as direction vector, length, and coordinates of key turning points) are converted into feature vectors and input into the prediction model.
[0113] Taking the LSTM model as an example, the network structure contains two hidden layers (128 neurons each). The input dimensions are historical coordinates (X, Y, Z), velocity, and attitude angle, and the output dimension is the predicted coordinates for the next 10 time steps (0.5 seconds each). During the training phase, the model has learned motion patterns in different scenarios (such as acceleration changes when moving straight, turning, and avoiding obstacles), and can generate trajectories that conform to physical laws based on the current motion trend.
[0114] 513. Based on the suspected motion trajectory and the dynamic three-dimensional environment model, extract the spatial coordinates of the suspected motion trajectory;
[0115] After receiving the suspected motion trajectory output by the time series prediction algorithm, the trajectory data needs to be structured and parsed to extract the three-dimensional spatial coordinate information. Specifically, the suspected motion trajectory is usually indexed by a timestamp, storing the estimated position coordinates (X, Y, Z) of the target at each prediction time point, forming a trajectory dataset containing both time and spatial dimensions. The system first needs to confirm that the coordinate system of the trajectory data is consistent with the coordinate system of the dynamic three-dimensional environment model (e.g., uniformly using the world coordinate system or a local coordinate system). If there are coordinate system differences, coordinate normalization processing needs to be performed using a preset transformation matrix to ensure the accuracy of subsequent spatial comparisons.
[0116] At the data structure level, suspected motion trajectories may be stored as lists, arrays, or tensors. Each trajectory point contains three-dimensional coordinates accurate to the millimeter level and a corresponding timestamp. For example, the predicted motion trajectory of rescue personnel in a drone monitoring scenario may contain a sequence of data in the form of [(t1, x1, y1, z1), (t2, x2, y2, z2), ..., (tn, xn, yn, zn)], where ti is the timestamp and (xi, yi, zi) are the predicted spatial coordinates. The system iterates through this sequence, extracts the (x, y, z) coordinates of all trajectory points, and forms an independent three-dimensional coordinate set P = {(x1, y1, z1), (x2, y2, z2), ..., (xn, yn, zn)}, which serves as input data for subsequent spatial comparisons.
[0117] 514. Compare the spatial coordinates of the obstacle positions in the dynamic three-dimensional environment model with the spatial coordinates of the suspected motion trajectory to obtain overlapping trajectory points;
[0118] After acquiring the spatial coordinate set P of suspected motion trajectories, the system retrieves the spatial coordinate data of obstacles from the dynamic 3D environment model. Obstacles in the dynamic 3D environment model include fixed structures (such as walls, beams, columns, and furniture) and dynamic obstacles (such as collapsed components and moving flame zones), and their positional information is stored in the form of 3D volumetric data. For each obstacle, the system extracts its occupied spatial coordinate range, including the minimum and maximum X, Y, and Z coordinates (i.e., x_min, x_max, y_min, y_max, z_min, z_max), forming a spatial bounding volume set O = {O1, O2, ..., Om}, where each Oi corresponds to the spatial range of an obstacle.
[0119] Next, the system performs obstacle collision detection on each point (x, y, z) in the trajectory point set P. The specific detection logic is as follows: For the current trajectory point p, iterate through all obstacles Oi and determine whether p satisfies x_min ≤ x ≤ x_max, y_min ≤ y ≤ y_max, and z_min ≤ z ≤ z_max. If satisfied, it is determined that the trajectory point p and obstacle Oi have spatial overlap, meaning the point is located within the space occupied by the obstacle and is therefore an invalid trajectory point.
[0120] Taking wall obstacles in a fire scenario as an example, assuming the spatial range of a wall is x∈[2, 4], y∈[1,3], z∈[0, 3] (unit: meters), if the coordinates of a trajectory point are (3, 2, 1.5), then this point falls within the bounding box of the wall and is determined to be an overlapping point; if another trajectory point is (5, 2, 1.5), then it does not fall within the wall's range and is determined to be a valid point. For dynamic obstacles (such as flame areas), the system needs to update their spatial coordinate range in real time (e.g., dynamically adjust the flame bounding box based on infrared image detection results) to ensure the real-time performance of collision detection.
[0121] 515. Mark the trajectory points as invalid points to obtain the remaining points in the suspected motion trajectory;
[0122] For each trajectory point, a validity identifier (e.g., a boolean value "valid", initially set to true) is generated. When an overlap between the point and an obstacle is detected, "valid" is marked as false. Subsequently, the system filters the trajectory point set based on the validity identifiers, retaining all points where "valid" is true, forming the remaining valid point set P_valid.
[0123] At the data storage level, the system can maintain a mask array `mask` corresponding to the set of trajectory points, where `mask[i]` corresponds to the validity status of trajectory point P[i]. For example, if the original trajectory contains 100 points, and 10 of them overlap with obstacles, then the positions in the `mask` array corresponding to these 10 points are set to `false`, and the rest are set to `true`. Through mask operations, the system can quickly extract valid points, avoiding direct modification of the original trajectory data and facilitating subsequent backtracking and verification.
[0124] Furthermore, the system needs to handle the continuity of invalid points. If multiple consecutive trajectory points are marked as invalid (e.g., the predicted trajectory of the target passes through a thick wall), it needs to determine whether these invalid points constitute an impassable area. In fire rescue scenarios, if the predicted trajectory of rescuers continuously passes through walls, the system should identify this segment of the trajectory as unreasonable and regenerate a path around the obstacle in subsequent steps using a trajectory correction algorithm. Simultaneously, for isolated invalid points (such as a single point falling into an obstacle corner), the system can directly filter them, retaining the valid points before and after them to provide basic data for subsequent trajectory fitting.
[0125] 516. Based on the remaining points, a reasonable motion trajectory is re-determined;
[0126] First, the points in the valid point set are sorted by timestamp to ensure the correct temporal order of the trajectory points. If there are discontinuous timestamps (such as a point at a certain moment being marked as invalid), the system needs to determine whether interpolation is needed to supplement intermediate points (e.g., using linear interpolation or polynomial interpolation) based on the time interval between valid points and the movement trend before and after, in order to ensure the temporal continuity of the trajectory.
[0127] To avoid abrupt changes in the predicted trajectory (such as instantaneous velocity changes exceeding physical limits), the system employs a smoothing algorithm to fit valid points. For example, cubic spline interpolation is used to generate a smooth curve passing through all valid points, ensuring the continuity of the first derivative (velocity) and second derivative (acceleration) of the trajectory, conforming to the physical laws of human motion. In fire rescue scenarios, the movement speed of rescuers typically does not exceed 6 meters per second, and their acceleration does not exceed the acceleration due to gravity. The system can set these physical constraints to verify the fitted trajectory and eliminate outliers that do not meet the constraints.
[0128] 517. Expand the motion trajectory into a volumetric space with a region to obtain the range of the motion trajectory;
[0129] The basic volume parameters are determined based on the anthropometric features of the monitored target (such as height, shoulder width, and body shape). For example, the average size of an adult human body can be modeled as a cuboid with a height of 1.75 meters, a width of 0.5 meters, and a thickness of 0.3 meters. For rescue personnel, the volume of protective equipment (such as oxygen cylinders) needs to be considered, expanding the model to a cuboid with a height of 1.8 meters and a width of 0.6 meters. The system maintains a target volume template library and dynamically selects the corresponding volume model based on extracted anthropometric features (such as body shape level identified through visible light images, and standing or bending posture features).
[0130] Centered on each valid point (x_i, y_i, z_i) on the trajectory line, the target volume is expanded by half its size along each axis of the three-dimensional space to form a local bounding volume. For continuous trajectories, adjacent bounding volumes are merged to form the overall motion trajectory volume space.
[0131] By binding volume space with timestamps, a dynamic "spatiotemporal volume" is formed. For example, the volume space corresponding to the trajectory point (t1, x1, y1, z1) is the effective search range within t∈[t1-Δt, t1+Δt] (considering data transmission delay), ensuring the temporal continuity of the target's motion.
[0132] 518. Mark the range of the motion trajectory in the dynamic three-dimensional environment model to obtain the marked three-dimensional environment model;
[0133] The volumetric space of the motion trajectory (such as a cylinder, ellipsoid, or custom shape) is converted into a geometric data structure in three-dimensional space. For example, if the trajectory range is a cylinder with radius R centered on the sequence of trajectory points, then the cylinder corresponding to each trajectory point can be represented as a geometric entity with the coordinates of the center of the base (x_i, y_i, z_i), radius R, and height H (considering human height). For continuous trajectories, adjacent cylinders can be merged into a continuous tubular volumetric space, forming a motion channel that runs through the three-dimensional environment.
[0134] Ensure that the coordinate system of the motion trajectory volume space is completely consistent with the dynamic three-dimensional environment model (e.g., uniformly adopt the world coordinate system, with the unit being meters).
[0135] The trajectory range is integrated into the environment model in the following ways:
[0136] Mesh voxelization: Divide the dynamic 3D environment model into a voxel mesh with a resolution of 0.1m × 0.1m × 0.1m. Determine whether each voxel belongs to the motion trajectory volume space (such as falling within the cylinder) and mark it as a "trajectory-associated voxel".
[0137] Bounding box annotation: Generate axis-aligned bounding boxes (AABB) or oriented bounding boxes (OBB) for the entire trajectory range, which are visualized as semi-transparent red grids in the 3D model to facilitate rapid identification by subsequent algorithms;
[0138] Attribute Attachment: Add metadata to the labeled trajectory range, such as timestamp (indicating the predicted time period corresponding to the trajectory), confidence level (the output confidence value based on the time series prediction algorithm), and target type (rescuers or trapped persons, distinguished by clothing color or thermal features in human body feature information).
[0139] 519. Based on the human body feature information, select matching regions with similar features in the labeled 3D environment model;
[0140] Human feature information (including visible light color features, infrared thermal features, and posture features) extracted through multimodal data fusion is used to accurately filter target regions within an annotated 3D environment model. First, for visible light images, the system extracts the target's clothing color features and enhances the robustness of color recognition through ultraviolet image dehazing. For infrared images, the system extracts the human body's thermal radiation features (e.g., areas with body temperatures between 30 and 35°C), excluding high-temperature areas (above 200°C) of flames and low-temperature areas (below 25°C). For posture features, a skeletal keypoint detection algorithm is used to obtain the target's joint coordinates (e.g., torso angle, arm swing amplitude), forming a posture feature vector. Subsequently, the system divides the region in the annotated 3D environment model associated with the trajectory range into multiple subspaces (e.g., a search cube every 1 meter along the trajectory path), narrowing the search range to improve efficiency. Within each subspace, the system performs feature matching on a voxel or region-by-region basis: For visible light data, it calculates the Barcol distance between the candidate region's color histogram and the target's color histogram; a distance less than a threshold (e.g., 0.2) is considered a color match. For infrared data, only thermal signal regions within the body temperature range are retained. For posture features, a machine learning model (e.g., support vector machine) is used to determine whether the keypoint distribution of the candidate region is consistent with the target's historical posture, excluding static objects or non-human heat sources. For example, in a smoke-covered scenario, the system can locate thermal signal regions of 30-35°C using infrared thermal features, and then combine this with ultraviolet imagery to help identify the residual orange clothing color features, ultimately selecting candidate regions that match the characteristics of rescue personnel.
[0141] 520. Determine whether the matching area is within the range of the motion trajectory based on the positioning information;
[0142] After selecting candidate matching regions, it is necessary to verify whether their spatial location is consistent with the range of the motion trajectory to ensure the accuracy and relevance of the positioning results. First, the system parses the positioning information of each matching region, including the geometric center coordinates (obtained by calculating the average of the coordinates of all voxels within the region), the spatial range (minimum and maximum X, Y, Z coordinates), and the data acquisition timestamp (which needs to be aligned with the time window of trajectory prediction). Subsequently, the system performs spatial inclusion judgment: for the trajectory volume space (such as a cylinder), it determines whether the geometric center of the matching region falls within its bounding box (i.e., satisfying x_min≤x_center≤x_max, y_min≤y_center≤y_max, z_min≤z_center≤z_max), and further calculates whether the distance from the center to the trajectory centerline is less than a preset radius (such as 0.3 meters, to accommodate normal human movement offset). If the trajectory range is a complex geometry (such as an ellipsoid that changes with posture), a spatial geometric algorithm (such as ray intersection detection) is used to determine whether the matching region intersects with the trajectory volume space. Meanwhile, the system verifies timestamp consistency to ensure that the acquisition time of the matching area is within the trajectory prediction time window (e.g., t ± 1 second), avoiding misclassification of historical or future data as the current location. For matching areas that partially fall within the trajectory range (e.g., the area edge is tangent to the trajectory volume space), the system adopts a "volume percentage" strategy: if more than 50% of the area voxels are within the trajectory range, it is considered valid; otherwise, it is considered invalid. For example, if the center coordinates of a matching area are (4, 2, 1.5), it falls within the trajectory bounding box (x = 1~5 meters, y = 1.7~2.3 meters, z = 0.6~2.4 meters) and is less than the radius from the center line, and the acquisition time is consistent with the trajectory prediction time, it is considered a valid matching area; if the center coordinates of another area exceed x_max, it is directly excluded. Through the above spatiotemporal consistency verification, the system can accurately filter out target locations highly correlated with the predicted trajectory, ensuring that the positioning results conform to both the target movement trend and the physical constraints of three-dimensional space.
[0143] 521. If not, then calculate the distance from the center point of the matching region to the range of the motion trajectory based on the center coordinates of the matching region;
[0144] Based on the center coordinates of the matching region, the distance from its center point to the range of the motion trajectory is calculated. Here, the center coordinates refer to the geometric center (x_center, y_center, z_center) obtained by calculating all voxels within the matching region, representing the most probable location of that region. The range of the motion trajectory in three-dimensional space is usually represented by a specific geometric shape, such as an axis-aligned bounding box (AABB, a cuboid region aligned with the coordinate axes) or a cylinder (a three-dimensional region with the trajectory centerline as its axis and a radius covering the range of human movement). For AABBs, distance calculation requires determining whether the center point is inside the bounding box: if inside, the distance is 0; if outside, the Euclidean distance from the center point to the nearest surface of the bounding box is calculated (i.e., the square root of the sum of the squares of the minimum deviation distances along the X, Y, and Z axes). For cylinders, the projection of the center point onto a plane perpendicular to the centerline must first be calculated. If the projection lies within the trajectory line segment, the distance is the perpendicular distance from that point to the centerline; if it exceeds this range, the straight-line distance to the nearest endpoint is calculated. For example, if the center coordinates of a matching region are (6, 2, 1.5), and the AABB of the trajectory range are x∈[1,5], y∈[1.7,2.3], z∈[0.6,2.4], then the point is 1 meter beyond the right boundary in the X-axis direction, while the Y and Z axes are within the range. The distance from the point to the trajectory range is the deviation of 1 meter in the X-axis direction.
[0145] 522. Based on the distance values, sort the matching regions to obtain the matching region with the smallest distance;
[0146] Based on the distance values of all matching regions that do not fall within the trajectory range, they are sorted in ascending order to obtain the matching region with the smallest distance. The sorting process employs efficient algorithms (such as quicksort or heapsort) to ensure real-time performance when processing a large number of candidate regions. Simultaneously, to avoid misjudgments caused by relying solely on spatial distance, the system combines the multimodal feature matching scores of the matching regions (such as visible light color similarity, infrared thermal feature matching degree, and posture feature consistency score) to weight and correct the distance values. For example, a region that is slightly farther away but whose thermal features perfectly match human body temperature (30~35℃) may have a higher overall priority than a region that is closer but whose thermal features are abnormal (such as approaching flame temperature). After sorting, the matching region with the smallest distance is considered the candidate region with the smallest deviation from the predicted motion path of the target and is most likely to be the true location.
[0147] 523. Based on the positioning information of the matching area with the smallest distance, the real-time position and direction of movement of the monitored target are obtained;
[0148] The system determines the real-time location and direction of movement of the monitored target based on the positioning information of the matching area with the smallest distance. The real-time location is directly obtained using the center coordinates (x_center, y_center, z_center) of that area, which are accurately calibrated using multimodal data fusion technology (integrating visible light, infrared, and depth data). The derivation of the direction of movement combines the historical trajectory data of the target before it disappeared: the position coordinates of the last few frames are retrieved (e.g., (x_prev, y_prev, z_prev)), and the vector difference between the current center coordinates and the historical positions is calculated (x_center - x_prev, y_center - y_prev, z_center - z_prev). The direction of this vector is the direction of movement of the target (e.g., vector (1, 0, 0) represents movement along the positive X-axis). In addition, the system will combine the traversable paths in the dynamic 3D environment model (e.g., corridor directions, staircase locations) to correct the direction of movement, ensuring that the direction conforms to the actual environmental constraints (e.g., it will not point towards walls or other obstacles). For example, if the center coordinates of the region with the smallest distance are (5.2, 2, 1.6) (adjacent to the right boundary of the trajectory range x=5), and the historical position is (4, 2, 1.5), then the motion direction vector is (1.2, 0, 0.1). After environmental constraint correction, it is determined that the target moves along the positive X-axis direction.
[0149] 524. If so, the spatial location of the monitored target is determined based on the matching area to obtain the real-time location and direction of movement of the monitored target.
[0150] The system directly locates the monitored target spatially based on the matching area. The real-time position is also based on the center coordinates of the matching area, which not only conforms to the spatial range of the predicted trajectory but is also verified through multimodal feature matching, ensuring high reliability. The determination of the motion direction combines the trajectory prediction results with position changes over multiple consecutive frames: if the predicted trajectory's motion trend at the current moment is along the negative Y-axis, and the center coordinates of the matching area decrease by 0.2 meters in the Y-axis direction compared to the previous frame (still within the allowable deviation range of the trajectory), then the motion direction is confirmed to be in the negative Y-axis direction; if position changes over multiple consecutive frames show the target moving at a stable speed along a specific direction (e.g., center coordinates of three consecutive frames are (2, 2, 1.5), (3, 2, 1.5), (4, 2, 1.5)), then the motion direction (here, the positive X-axis direction) is directly fitted using measured data, further improving the accuracy and real-time performance of the positioning.
[0151] The intelligent monitoring method described in this application integrates visible light, ultraviolet, infrared images, and environmental depth data through multi-source data fusion technology. Combined with dynamic 3D environment modeling, time series prediction algorithms, and multimodal feature matching mechanisms, it achieves high-precision real-time positioning and motion trajectory prediction of monitoring targets (rescuers and trapped personnel) in complex fire scenarios. This not only solves the technical problems of information loss and large positioning deviation in traditional single sensors under smoke obstruction and complex lighting conditions, but also significantly improves the system's dynamic perception of target movement trends and robustness in extreme environments through trajectory volume space expansion, obstacle screening, and spatiotemporal consistency verification of matching areas. It ensures that even when the target briefly leaves the monitoring screen or is affected by environmental interference, it can still be continuously tracked based on environmental constraints and historical features, providing reliable technical support for emergency rescue scenarios.
[0152] The embodiments of the intelligent monitoring method of this application have been described above. In the embodiments of this application, the monitoring method can be executed by an embedded device, which includes an electronic device. The hardware architecture of the electronic device of this application is described below:
[0153] Please see Figure 6 This is a schematic diagram of a hardware architecture of an electronic device in an embodiment of this application.
[0154] The electronic device includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on a program stored in Read-Only Memory (ROM) 602 or a program loaded from storage portion 408 into Random Access Memory (RAM) 603. The Random Access Memory (RAM) 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An Input / Output (I / O) interface 605 is also connected to the bus 604.
[0155] The following components are connected to the input / output (I / O) interface 605: an input section 606 including audio input devices, push-button switches, etc.; an output section 607 including displays, audio output devices, indicator lights, etc.; a storage section 608 including hard disks, etc.; and a communication section 609 including network interface cards such as LAN (Local Area Network) cards, modems, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.
[0156] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by the Central Processing Unit (CPU) 601, it performs the various functions defined in the present invention.
[0157] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.
[0159] Specifically, the electronic device of this embodiment includes a processor and a memory. The memory is coupled to one or more processors and is used to store computer program code. The computer program code includes computer instructions. One or more processors call the computer instructions to cause the electronic device to perform the method provided in the above embodiment.
[0160] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The storage medium carries one or more computer programs that, when executed by a processor of the electronic device, cause the electronic device to implement the methods provided in the above embodiments.
[0161] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0162] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0163] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0164] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for intelligent monitoring, applied to a monitoring system, characterized in that, The method comprises: receiving environmental depth data, visible light images, ultraviolet images and infrared images from a fire scene in a monitoring system; extracting features from the visible light images, the ultraviolet images and the infrared images to obtain human feature information and initial position region features of a monitoring target, the human feature information including color features, thermal features and posture features, and the monitoring target including rescue personnel and trapped personnel; in the case that the monitoring target disappears in a monitoring picture, constructing a dynamic three-dimensional environmental model of a disappearance region based on the environmental depth data and the initial position region features, the dynamic three-dimensional environmental model including an initial smoke-shielded region, a flame distribution region and the initial position region; extracting smoke spatial distribution features of the initial smoke-shielded region from the dynamic three-dimensional environmental model; based on the initial position region features, the human feature information and the smoke spatial distribution features, predicting a movement trajectory of the monitoring target in the smoke-shielded region by using a time series prediction algorithm; based on the human feature information, checking whether a matching region is within the range of the movement trajectory in the dynamic three-dimensional environmental model, the matching region including a region having similar features to the human feature information; if yes, positioning a spatial position of the monitoring target based on the matching region to obtain a real-time position and a movement direction of the monitoring target.
2. The method of claim 1, wherein, extracting features from the visible light images, the ultraviolet images and the infrared images to obtain human feature information and initial position region features of a monitoring target, specifically comprising: extracting features from the visible light images, the ultraviolet images and the infrared images to obtain image feature data; integrating the image feature data and the environmental depth data into a unified multi-modal image data stream by using a multi-modal data fusion algorithm, the multi-modal image data stream fusing visible light, ultraviolet, infrared and depth information; extracting human feature information and initial position region features of a monitoring target from the multi-modal image data stream.
3. The method of claim 1, wherein, in the case that the monitoring target disappears in a monitoring picture, constructing a dynamic three-dimensional environmental model of a disappearance region based on the environmental depth data and the initial position region features, specifically comprising: in the case that the monitoring target disappears in a monitoring picture, converting two-dimensional coordinates into a three-dimensional coordinate system based on the environmental depth data and the visible light images to generate an initial three-dimensional environmental model; labeling the initial position region in the initial three-dimensional environmental model to obtain a three-dimensional environmental model of the disappearance region of the monitoring target; labeling an initial smoke-shielded region in the three-dimensional environmental model by using the ultraviolet images and the environmental depth data; identifying a flame distribution region by using the infrared images to obtain a dynamic three-dimensional environmental model in which the initial smoke-shielded region and the flame distribution region are spatially superimposed.
4. The method of claim 1, wherein, based on the initial position region features, the human feature information and the smoke spatial distribution features, predicting a movement trajectory of the monitoring target in the smoke-shielded region by using a time series prediction algorithm, specifically comprising: Screening a passable path of the monitoring target based on the initial position area feature and the smoke spatial distribution feature; Matching the passable path with the human body feature information to obtain a suspected motion trend of the monitoring target; Predicting a suspected motion trajectory of the monitoring target in the dynamic three-dimensional environment model by using a time series prediction algorithm combined with the suspected motion trend; Screening the suspected motion trajectory in the dynamic three-dimensional environment model to obtain a reasonable motion trajectory of the monitoring target in a smoke shielding area.
5. The method of claim 4, wherein, Screening the suspected motion trajectory in the dynamic three-dimensional environment model to obtain a reasonable motion trajectory of the monitoring target in a smoke shielding area, specifically comprising: Extracting spatial coordinates of the suspected motion trajectory based on the suspected motion trajectory and the dynamic three-dimensional environment model; Comparing spatial coordinates of an obstacle position in the dynamic three-dimensional environment model with spatial coordinates of the suspected motion trajectory to obtain a trajectory point with overlap; Marking the trajectory point as an invalid point to obtain remaining points in the suspected motion trajectory; Re-determining a reasonable motion trajectory based on the remaining points.
6. The method of claim 1, wherein, Checking whether a matching area is within a range of the motion trajectory in the dynamic three-dimensional environment model based on the human body feature information, specifically comprising: Expanding the motion trajectory into a volumetric space with an area to obtain a motion trajectory range; Labeling the motion trajectory range in the dynamic three-dimensional environment model to obtain a labeled three-dimensional environment model; Screening a matching area with similar features in the labeled three-dimensional environment model based on the human body feature information; Determining whether the matching area is within the range of the motion trajectory by positioning information.
7. The method of claim 1, wherein, After the step of checking whether a matching area is within a range of the motion trajectory in the dynamic three-dimensional environment model based on the human body feature information, the method further comprises: If not, calculating a distance value from a center point of the matching area to the motion trajectory range based on a center coordinate of the matching area; According to the distance value, sorting the matching areas to obtain a matching area with the smallest distance; Obtaining a real-time position and a motion direction of the monitoring target based on positioning information of the matching area with the smallest distance.
8. A monitoring system, characterized by A monitoring system comprising one or more processors and a memory; the memory is coupled with the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, the one or more processors invoke the computer instructions to enable the monitoring system to perform the method of any one of claims 1-7.
9. An embedded device, characterized by When the embedded device runs on the monitoring system, the monitoring system is enabled to perform the method of any one of claims 1-7.
10. A computer readable storage medium storing computer instructions, characterized in that, When the computer instructions run on the monitoring system, the monitoring system is enabled to perform the method of any one of claims 1-7.
Citation Information
Patent Citations
Video flame detecting method based on multi-feature fusion technology
CN103116746A
Fire early warning method and system based on smoke image joint feature analysis
CN119181197A