Intelligent response method and system based on building multi-view video fusion

By integrating multi-view video data and BIM models within a building, and utilizing weighted smoothing and seamless integration of virtual and real technologies, the location of abnormal events can be accurately pinpointed, solving the problem of inaccurate location in existing technologies and improving response efficiency.

CN121353394BActive Publication Date: 2026-04-07HUBEI KAIMEI ENERGY TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, multiple cameras inside a building have independent perspectives, and the video footage lacks spatial correlation with the building's BIM model, making it impossible to accurately locate abnormal events and affecting response efficiency.

Method used

By acquiring surveillance video data collected by cameras in buildings, initial panoramic video data is generated based on a weighted smooth fusion strategy. Combined with the basic spatial parameters and supplementary features of the corresponding blind spots of the cameras, video data is fused using a seamless virtual-real connection mechanism. Abnormal events are accurately located through the perspective transformation matrix of BIM coordinates and pixel coordinates.

Benefits of technology

It achieves millimeter-level binding between video footage and BIM 3D coordinates, accurately locating the position of abnormal events in the building's 3D space and improving response efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353394B_ABST
    Figure CN121353394B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent response method and system based on building multi-view video fusion, it is related to video processing technical field, the method comprises: based on weighted smoothing fusion strategy, the monitoring video data of existing view overlap is fused, and initial panoramic video data is obtained;Based on the basic space parameters and supplementary features of the camera corresponding monitoring blind area, generate blind area video data;Based on virtual seamless connection mechanism, the initial panoramic video data and blind area video data are fused, and the panoramic video data of building is generated;Based on exception determination rule base, the panoramic video data is detected, and the pixel coordinates of abnormal site are determined;Based on perspective conversion matrix, the pixel coordinates of abnormal site are converted into the BIM coordinates of abnormal site, and the response route is planned in the BIM model of building.Through the above-mentioned mode, the position of abnormal event in building three-dimensional space can be accurately positioned, and the response efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to an intelligent response method and system based on multi-view video fusion of buildings. Background Technology

[0002] Multiple cameras inside the building have independent perspectives, and the video footage lacks spatial correlation with the building's BIM model. After an abnormal event (such as a person falling or equipment emitting smoke) occurs, only a planar image can be displayed, and it is impossible to locate the precise position of the event in the building's three-dimensional space. Relying on a single video feature, the false alarm rate for complex scenes (such as changes in light and shadow, crowd occlusion) is high, affecting response efficiency.

[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide an intelligent response method and system based on multi-view video fusion of buildings, which aims to solve the technical problem that traditional solutions in the prior art are unable to locate the precise location of abnormal events in the three-dimensional space of the building, thus affecting the response efficiency.

[0005] To achieve the above objectives, this application provides an intelligent response method based on multi-view video fusion of buildings, the method comprising:

[0006] Acquire surveillance video data collected by cameras in buildings, and fuse surveillance video data with overlapping viewpoints based on a weighted smooth fusion strategy to obtain initial panoramic video data;

[0007] Based on the basic spatial parameters and supplementary features of the monitoring blind spot corresponding to the camera, blind spot video data is generated.

[0008] Based on a seamless virtual-real integration mechanism, the initial panoramic video data and the blind spot video data are fused to generate panoramic video data of the building.

[0009] Based on the anomaly detection rule base, anomaly detection is performed on the panoramic video data to determine the pixel coordinates of the anomaly sites;

[0010] Based on the perspective transformation matrix between BIM coordinates and pixel coordinates, the pixel coordinates of the abnormal site are converted into the BIM coordinates of the abnormal site.

[0011] Based on the BIM coordinates of the anomaly sites, corresponding response routes are planned in the BIM model of the building.

[0012] In one embodiment, the step of acquiring surveillance video data collected by cameras in a building and fusing surveillance video data with overlapping viewpoints based on a weighted smoothing fusion strategy to obtain initial panoramic video data includes:

[0013] Acquire surveillance video data collected by cameras in the building, and construct a perspective transformation matrix between BIM coordinates and pixel coordinates based on the building's BIM model and the surveillance video data collected by the cameras in the building.

[0014] Surveillance video data with overlapping viewpoints are used as overlapping surveillance video data. Based on a multi-feature point fusion matching strategy, corresponding common feature points are extracted from the overlapping viewpoint scenes of the overlapping surveillance video data.

[0015] Based on the perspective transformation matrix, the pixel coordinates of the common feature points are converted into the BIM coordinates of the common feature points;

[0016] Based on the BIM coordinates and pixel coordinates of the common feature points, the coordinate offset of the overlapping viewpoint of the overlapping monitoring video data is determined.

[0017] Based on the coordinate offset of the overlapping viewpoints of the overlapping surveillance video data, the overlapping surveillance video data is smoothly stitched together to obtain initial panoramic video data.

[0018] In one embodiment, the step of constructing a perspective transformation matrix between BIM coordinates and pixel coordinates based on the BIM model of the building and surveillance video data collected by cameras in the building includes:

[0019] Obtain the BIM model of the building and select the building reference point in the BIM model;

[0020] Based on the building reference point, the BIM coordinates of the center point of the spatial calibration target are obtained. The spatial calibration target is installed directly in front of the camera. A QR code pattern is provided in the middle of the spatial calibration target. The QR code pattern stores camera identification information, installation angle information and installation height information. Multiple infrared marker points are set around the edge of the spatial calibration target at preset intervals.

[0021] The image of the spatial calibration target is extracted from the surveillance video data collected by the camera, and the pixel coordinates of the infrared marker points are extracted from the image of the spatial calibration target.

[0022] Based on the pixel coordinates of the infrared marker points, determine the pixel coordinates of the center point of the spatial calibration target;

[0023] Extract the corresponding QR code pattern from the image of the spatial calibration target. Based on the extracted QR code pattern, match the pixel coordinates of the center point of the spatial calibration target with the BIM coordinates of the center point of the spatial calibration target to determine the matching coordinate pair.

[0024] Based on the matched coordinate pairs, a perspective transformation matrix between BIM coordinates and pixel coordinates is constructed.

[0025] In one embodiment, the step of smoothly stitching the overlapping surveillance video data based on the coordinate offset of the overlapping viewpoints to obtain initial panoramic video data includes:

[0026] Align the overlapping surveillance video data based on the coordinate offset of the overlapping viewpoints of the overlapping surveillance video data.

[0027] Based on the aligned overlapping surveillance video data, update the BIM coordinates of the overlapping viewpoints, and based on the updated BIM coordinates of the overlapping viewpoints, calculate the spatial distance between the overlapping point in the overlapping viewpoints and the corresponding camera in the overlapping viewpoints.

[0028] Based on the spatial distances corresponding to the overlapping points in the overlapping viewpoints, corresponding smoothing weights are assigned to the overlapping points in the overlapping viewpoints.

[0029] Based on the smoothing weight of the overlapping points in the overlapping viewpoints, the fused pixel value corresponding to the overlapping points is calculated, and the fused pixel value includes at least the fused color value and the fused brightness value;

[0030] Based on the fused pixel values ​​corresponding to the overlapping sites, spliced ​​video data corresponding to the overlapping monitoring video data is generated;

[0031] Initial panoramic video data is generated based on non-overlapping surveillance video data and stitched video data corresponding to overlapping surveillance video data.

[0032] In one embodiment, the step of generating blind spot video data based on the basic spatial parameters and supplementary features of the blind spot corresponding to the camera includes:

[0033] The basic spatial parameters of the monitoring blind spot are extracted from the building's BIM model. The basic spatial parameters include at least dimensional parameters, spatial constraint parameters, and coordinate anchor points. The dimensional parameters include at least the elevator car dimensions, the number of steps at the corner of the stairwell, the step height at the corner of the stairwell, the step width at the corner of the stairwell, the corner dimensions of the equipment room, and the position coordinates of the fixed equipment in the corner of the equipment room. The spatial constraint parameters include at least the motion boundary and the position coordinates of the fixed components. The coordinate anchor points are set at the entrance or exit of the monitoring blind spot.

[0034] Obtain supplementary features of the monitoring blind spot, wherein the supplementary features include at least texture features, light features, and environmental noise features;

[0035] Based on the supplementary features and basic spatial parameters of the monitoring blind spot, an optimized virtual scene model of the monitoring blind spot is generated;

[0036] When a moving target exists in the monitoring blind spot, blind spot video data of the monitoring blind spot is generated based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot. The core features include at least motion features, type features and appearance features.

[0037] In one embodiment, the step of generating blind spot video data of the monitoring blind spot based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot includes:

[0038] Based on the type and appearance characteristics of the moving target, a virtual appearance model of the moving target is generated;

[0039] Based on the motion characteristics of the moving target, the trajectory of the moving target is predicted to obtain the predicted trajectory;

[0040] Based on the optimized virtual scene model of the monitoring blind spot and the virtual appearance model and predicted trajectory of the moving target, continuous virtual video frames are generated.

[0041] The viewing angle of the continuous virtual video frames is adjusted to the expected viewing angle of the monitoring blind spot, and the interaction features between the moving target and the monitoring blind spot are added to the continuous virtual video frames to form the blind spot video data of the monitoring blind spot.

[0042] In one embodiment, before the step of generating blind spot video data of the monitoring blind spot based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot, the method further includes:

[0043] Based on an improved YOLOv8 target detection model, the core features of the moving target are extracted from the relevant video data. The improved YOLOv8 target detection model includes an input layer, a backbone network, a neck network, and a head network. The input layer includes a lighting correction submodule (replacing the Mosaic enhancement module in the original model), a multi-scale anchor point adaptation module, and a BIM coordinate pre-association module. The backbone network includes an optimized C2f module (replacing the C2f module in the original model) and an optimized SPPF module (replacing the SPPF module in the original model). The head network includes a newly added feature extraction head and the classification and regression heads from the original model. The input layer preprocesses the input relevant video data, inputs the obtained feature map into the backbone network, extracts features from the feature map, inputs the obtained backbone features into the neck network, performs multi-scale feature fusion on the backbone features, inputs the obtained fused features into the head network, and outputs the core features corresponding to the fused features in parallel.

[0044] In one embodiment, the step of fusing the initial panoramic video data and the blind spot video data based on a seamless virtual-real integration mechanism to generate panoramic video data of the building includes:

[0045] Physical boundaries are extracted from the initial panoramic video data and the blind spot video data, and an initial boundary fusion region is delineated based on the physical boundaries.

[0046] Extract the physical boundary features and the corresponding virtual boundary features from the initial boundary fusion region;

[0047] Based on the pixel coordinates of the physical boundary features and the pixel coordinates of the virtual boundary features corresponding to the physical boundary features, the initial panoramic video data and the blind spot video data are aligned, and the boundary fusion region is determined. When there is a moving target at the physical boundary, a transition frame is added between the initial panoramic video data and the blind spot video data.

[0048] Based on the fusion weight, the pixels in the boundary fusion area are weighted and fused to obtain the panoramic video data of the building. The fusion weight is determined based on the distance between the pixel and the physical boundary.

[0049] In one embodiment, the step of planning the corresponding response route in the BIM model of the building based on the BIM coordinates of the anomaly site includes:

[0050] Key parameters of accessible passages are extracted from the BIM model of the building, emergency resources are associated with the accessible passages, and fixed obstacles are marked. The key parameters include at least geometric parameters, access attributes, and associated facilities.

[0051] Based on the anomaly type and response subject of the aforementioned anomaly site, route constraints are formulated;

[0052] Based on the route constraints, the key parameters of the passable passage, the emergency resources, and the fixed obstacles, a response route is planned.

[0053] Furthermore, to achieve the above objectives, this application also proposes an intelligent response system based on multi-view video fusion of buildings. The intelligent response system based on multi-view video fusion of buildings includes:

[0054] The video fusion module is used to acquire surveillance video data collected by cameras in buildings. Based on a weighted smooth fusion strategy, it fuses surveillance video data with overlapping viewpoints to obtain initial panoramic video data.

[0055] The video fusion module is also used to generate blind spot video data based on the basic spatial parameters and supplementary features of the corresponding blind spot of the camera.

[0056] The video fusion module is also used to fuse the initial panoramic video data and the blind spot video data based on a seamless virtual-real connection mechanism to generate panoramic video data of the building.

[0057] An anomaly response module is used to perform anomaly detection on the panoramic video data based on an anomaly judgment rule base and determine the pixel coordinates of the anomaly site.

[0058] The anomaly response module is also used to convert the pixel coordinates of the anomaly site into the BIM coordinates of the anomaly site based on the perspective transformation matrix between BIM coordinates and pixel coordinates.

[0059] The anomaly response module is also used to plan a corresponding response route in the BIM model of the building based on the BIM coordinates of the anomaly site.

[0060] Furthermore, to achieve the above objectives, this application also proposes an intelligent response device based on multi-view video fusion of buildings. The intelligent response device based on multi-view video fusion of buildings includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the intelligent response method based on multi-view video fusion of buildings as described above.

[0061] In addition, to achieve the above objectives, the present invention also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the intelligent response method based on multi-view video fusion of buildings as described above.

[0062] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the intelligent response method based on multi-view video fusion of buildings as described above.

[0063] This application provides an intelligent response method based on multi-view video fusion of buildings. It acquires surveillance video data collected by cameras within a building, and fuses surveillance video data with overlapping views using a weighted smoothing fusion strategy to obtain initial panoramic video data. Based on the basic spatial parameters and supplementary features of the camera's corresponding blind spot, it generates blind spot video data. Using a seamless virtual-real connection mechanism, it fuses the initial panoramic video data with the blind spot video data to generate panoramic video data of the building. Based on an anomaly detection rule base, it performs anomaly detection on the panoramic video data to determine the pixel coordinates of anomaly points. Using a perspective transformation matrix between BIM coordinates and pixel coordinates, it converts the pixel coordinates of the anomaly points into BIM coordinates. Based on the BIM coordinates of the anomaly points, it plans the corresponding response route in the building's BIM model. This application can achieve millimeter-level binding between video images and BIM 3D coordinates, significantly improving the spatial positioning accuracy of abnormal events in buildings. It can accurately locate the position of abnormal events in the building's 3D space, effectively improving response efficiency and solving the technical problem that traditional solutions struggle to accurately locate the precise position of abnormal events in the building's 3D space, thus affecting response efficiency. Attached Figure Description

[0064] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0065] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0066] Figure 1 This is a flowchart illustrating an embodiment of the intelligent response method based on multi-view video fusion of buildings according to this application;

[0067] Figure 2This is a flowchart illustrating Embodiment 2 of the intelligent response method based on multi-view video fusion of buildings in this application;

[0068] Figure 3 This is a schematic diagram of the model structure of the intelligent response method based on multi-view video fusion of buildings provided in Embodiment 2 of this application;

[0069] Figure 4 This is a schematic diagram of the module structure of an intelligent response system based on multi-view video fusion of buildings, as described in an embodiment of this application.

[0070] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent response method based on multi-view video fusion of buildings in the embodiments of this application.

[0071] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0072] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0073] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0074] The main solution of this application embodiment is as follows: First, acquire surveillance video data collected by cameras in a building. Second, based on a weighted smoothing fusion strategy, fuse surveillance video data with overlapping perspectives to obtain initial panoramic video data. Third, based on the basic spatial parameters and supplementary features of the corresponding blind spots of the cameras, generate blind spot video data. Fourth, based on a seamless virtual-real connection mechanism, fuse the initial panoramic video data and the blind spot video data to generate panoramic video data of the building. Fifth, based on an anomaly detection rule base, perform anomaly detection on the panoramic video data to determine the pixel coordinates of anomaly points. Sixth, based on the perspective transformation matrix between BIM coordinates and pixel coordinates, convert the pixel coordinates of the anomaly points into BIM coordinates of the anomaly points. Finally, based on the BIM coordinates of the anomaly points, plan the corresponding response route in the BIM model of the building.

[0075] This application provides a solution that achieves millimeter-level binding between video footage and BIM 3D coordinates, significantly improving the spatial positioning accuracy of abnormal events in buildings. It can accurately locate the position of abnormal events in the 3D space of the building, effectively improving response efficiency and solving the technical problem that traditional solutions have difficulty locating the precise position of abnormal events in the 3D space of the building, thus affecting response efficiency.

[0076] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as an intelligent response device based on multi-view video fusion of buildings. This embodiment does not specifically limit it in this way. The following uses an intelligent response device based on multi-view video fusion of buildings as an example to describe this embodiment and the following embodiments.

[0077] This application provides an intelligent response method based on multi-view video fusion of buildings, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the intelligent response method based on multi-view video fusion of buildings in this application.

[0078] In this embodiment, the intelligent response method based on multi-view video fusion of buildings includes steps S10~S60:

[0079] Step S10: Acquire surveillance video data collected by cameras in the building, and fuse surveillance video data with overlapping viewpoints based on a weighted smooth fusion strategy to obtain initial panoramic video data.

[0080] It should be noted that buildings are usually equipped with cameras, such as high-speed rail stations, factories, and shopping malls. The surveillance video data collected by adjacent cameras often have overlapping perspectives, such as at corridor corners or multi-directional cameras in the lobby. In this embodiment, the surveillance video data with overlapping perspectives is weighted and smoothed to be stitched together into surveillance video data with an overall perspective, i.e., the initial panoramic video data.

[0081] In one feasible implementation, step S10 may include steps S101 to S105:

[0082] Step S101: Obtain surveillance video data collected by cameras in the building; based on the BIM model of the building and the surveillance video data collected by cameras in the building, construct a perspective transformation matrix between BIM coordinates and pixel coordinates.

[0083] In one feasible implementation, step S101 may include: acquiring a BIM model of the building, selecting a building reference point in the BIM model; acquiring the BIM coordinates of the center point of a spatial calibration target based on the building reference point, wherein the spatial calibration target is installed directly in front of a camera, a QR code pattern is provided in the center of the spatial calibration target, the QR code pattern stores camera identification information, installation angle information, and installation height information, and multiple infrared marker points are arranged around the edge of the spatial calibration target at preset intervals; extracting an image of the spatial calibration target from the monitoring video data collected by the camera, and extracting the pixel coordinates of the infrared marker points from the image of the spatial calibration target; determining the pixel coordinates of the center point of the spatial calibration target based on the pixel coordinates of the infrared marker points; extracting the corresponding QR code pattern from the image of the spatial calibration target, matching the pixel coordinates of the center point of the spatial calibration target with the BIM coordinates of the center point of the spatial calibration target based on the extracted QR code pattern, and determining a matching coordinate pair; and constructing a perspective transformation matrix between the BIM coordinates and the pixel coordinates based on the matching coordinate pair.

[0084] It should be noted that this embodiment imports a detailed BIM model of the building (with millimeter-level accuracy, including the three-dimensional coordinates and dimensional parameters of components such as walls, beams, columns, elevators, and HVAC equipment). A target is fixed in front of each camera installation location (such as the ceiling of a corridor or the wall of an elevator lobby). A QR code is embedded in the center of the target to store the camera ID, initial installation angle, and initial installation height. Four infrared markers (wavelength 850nm, diameter 5mm, spacing 10cm) surround the target to ensure that the camera can capture images clearly.

[0085] It is understandable that a building reference point is selected in the BIM model, such as the bottom of the column at the southeast corner of the first floor. This embodiment does not impose specific limitations on this. Using the building reference point as the origin, the three-dimensional coordinates (X, Y, Z) of the center point of each spatial calibration target in the BIM coordinate system are measured, with the error controlled within ±5mm, forming a "camera-target-BIM coordinate" mapping table.

[0086] It should be understood that by capturing images of spatial calibration targets with a camera, obtaining the camera ID by recognizing a QR code, and extracting the pixel coordinates of four infrared marker points, the pixel coordinates of the center point of the spatial calibration target are calculated. Then, according to the camera ID extracted from the image of the spatial calibration target, a "camera-target-pixel coordinate" mapping table is formed. According to the "camera-target-BIM coordinate" mapping table and the "camera-target-pixel coordinate" mapping table, matching pixel coordinates and BIM coordinates are found as matching coordinate pairs. The perspective transformation matrix between BIM coordinates and pixel coordinates is solved by the least squares method to obtain the mapping relationship between BIM coordinates and pixel coordinates. Thus, real-time conversion of any pixel point in the video image to BIM 3D coordinates can be achieved. For example, when a user clicks on any area (such as the location where a person falls) in the video monitoring interface, the corresponding 3D location will be automatically highlighted in the BIM model and a pop-up window will display the coordinate information (e.g., [X:23.5m, Y:12.8m, Z:9.0m], 3rd floor east corridor, 2.8m from the east elevator), while also marking the surrounding related components (e.g., the nearby fire hydrant ID is 3F-XF-005).

[0087] Step S102: Take the surveillance video data with overlapping viewpoints as overlapping surveillance video data, and extract the corresponding common feature points from the overlapping viewpoint scenes of the overlapping surveillance video data based on the multi-feature point fusion matching strategy.

[0088] It should be noted that surveillance video data with overlapping perspectives are treated as overlapping surveillance video data. Common feature points in the videos are extracted using a multi-feature point fusion matching algorithm, prioritizing fixed building components such as column corners, wall signs, and equipment shell textures. These feature points all have clear three-dimensional coordinates in the BIM model.

[0089] Step S103: Based on the perspective transformation matrix, convert the pixel coordinates of the common feature points into the BIM coordinates of the common feature points;

[0090] Understandably, for each common feature point, its pixel coordinates in the video data of two adjacent cameras are obtained, and combined with the established perspective transformation matrix between pixel coordinates and BIM coordinates, the pixel coordinates of the common feature point are converted into the corresponding BIM coordinates.

[0091] Step S104: Based on the BIM coordinates and pixel coordinates of the common feature points, determine the coordinate offset of the overlapping viewpoint of the overlapping monitoring video data.

[0092] It is understandable that the overlapping view is the view of the overlapping area. Based on the BIM coordinates and pixel coordinates of common feature points, the offset of pixel coordinates in the overlapping view is calculated, that is, the coordinate offset. For example, relative to camera B, camera A is offset by 0.2m on the X-axis and 0.1m on the Y-axis.

[0093] Step S105: Based on the coordinate offset of the overlapping viewpoints of the overlapping surveillance video data, the overlapping surveillance video data is smoothly stitched together to obtain initial panoramic video data.

[0094] It should be noted that a weighted average fusion algorithm is used to smoothly stitch together the pixels of overlapping viewpoints. The weights are distributed according to the distance between the camera and the feature points to generate seamless overall video data.

[0095] In one feasible implementation, step S105 may include: aligning the overlapping surveillance video data based on the coordinate offset of the overlapping viewpoints; updating the BIM coordinates of the overlapping viewpoints based on the aligned overlapping surveillance video data; calculating the spatial distance between the overlapping points in the overlapping viewpoints and the corresponding cameras in the overlapping viewpoints based on the updated BIM coordinates; assigning corresponding smoothing weights to the overlapping points in the overlapping viewpoints based on the spatial distances; calculating the fused pixel values ​​corresponding to the overlapping points based on the smoothing weights, wherein the fused pixel values ​​include at least fused color values ​​and fused brightness values; generating stitched video data corresponding to the overlapping surveillance video data based on the fused pixel values ​​corresponding to the overlapping points; and generating initial panoramic video data based on the non-overlapping surveillance video data and the stitched video data corresponding to the overlapping surveillance video data.

[0096] It's important to note that video data is aligned based on the coordinate offset of the overlapping viewpoints. Because different cameras have different installation angles and positions, the pixel coordinates of the same fixed feature point of a building (such as the corner of a column) will differ in different videos. The offset allows these coordinates to be mapped to the same BIM 3D coordinate system, ensuring the same object is consistently positioned in different videos. Furthermore, directly stitching together pixels from overlapping viewpoints can result in double edges and image shifts due to viewing angle differences. Using coordinate offsets corrects pixel positions, ensuring precise alignment of the two camera images in the overlapping area. The overlapping point is the location where the images captured by the two cameras with overlapping viewpoints overlap.

[0097] Understandably, each overlapping point can be assigned a weight used in the stitching process, namely a smoothing weight. The smoothing weight is determined according to the spatial distance between the overlapping point and the corresponding two cameras in the overlapping viewpoint. The smoothing weight is the proportion of the pixel value of the camera's captured image in the fusion process, including the weight of the first camera and the weight of the second camera. The weights of the first and second cameras at each overlapping point are added together to 1. Generally speaking, the closer the spatial distance, the larger the corresponding smoothing weight; the farther the spatial distance, the smaller the corresponding smoothing weight. For example, if cameras A and B have overlapping views, the smoothing weight for camera A is the first camera weight, and the smoothing weight for camera B is the second camera weight. Calculate the spatial distance d1 between the first overlapping point and camera A, and the spatial distance d2 between the first overlapping point and camera B. If d1 > d2, then set the weight of the first camera at the first overlapping point to 0.6, and the weight of the second camera at the first overlapping point to 0.4.

[0098] It should be understood that the pixel values ​​of overlapping points in the surveillance video data from two overlapping cameras are weighted and fused together, based on the two pixel values ​​of the overlapping points in the overlapping surveillance video data, combined with the weights of the first and second cameras at the overlapping points, to obtain the fused pixel value. The fused pixel value includes at least a fused color value and a fused brightness value. In practice, the color value of the overlapping points can be fused first, followed by the brightness value. The overlapping surveillance video data is fused according to the fused pixel value, and the resulting stitched video data is the stitched video data. In addition to the overlapping surveillance video data, there is also video data without overlapping perspectives, i.e., non-overlapping surveillance video data. This non-overlapping surveillance video data is then stitched together with the stitched video data to generate seamless initial panoramic video data. Furthermore, the initial panoramic video data is bound to the corresponding BIM area, allowing users to access the initial panoramic video data of any area in the BIM model by clicking on that area.

[0099] Step S20: Generate blind spot video data based on the basic spatial parameters and supplementary features of the blind spot corresponding to the camera.

[0100] It should be noted that a surveillance blind spot is an area inside a building that cannot be captured by cameras. Basic spatial parameters refer to the relevant parameters of the surveillance blind spot in the BIM model, including at least dimensional parameters, spatial constraint parameters, and coordinate anchor points.

[0101] Additionally, it should be noted that since the BIM model only contains geometric structures, this embodiment supplements detailed features such as texture and lighting through on-site scanning to ensure the visual consistency of the virtual image. These supplemented features are called supplementary features.

[0102] It is understandable that by integrating BIM geometric parameters with features such as texture and lighting collected on-site, and using a 3D modeling engine (such as Unity or Unreal Engine) to construct a physically-reproduced virtual scene of the blind spot, virtual video data of the monitoring blind spot is generated, i.e., blind spot video data.

[0103] Step S30: Based on the seamless integration mechanism of virtual and real, the initial panoramic video data and the blind spot video data are fused to generate panoramic video data of the building;

[0104] In one feasible implementation, step S30 may include steps S301 to S304:

[0105] Step S301: Extract physical boundaries from the initial panoramic video data and the blind spot video data, and delineate the initial boundary fusion region based on the physical boundaries;

[0106] It should be noted that the physical boundaries between monitoring blind spots and non-monitoring blind spots are extracted from the BIM model. For example, the plane when the elevator door is closed, the intersection of the stair landing and the step. In the initial panoramic video data and blind spot video data, the initial boundary fusion area is delineated with the physical boundary as the center. The width is 50 pixels, covering the image on both sides of the boundary (30 pixels on the non-blind spot side and 20 pixels on the blind spot side).

[0107] Step S302: Extract the physical boundary features and the virtual boundary features corresponding to the physical boundary features from the initial boundary fusion region;

[0108] Understandably, in the initial boundary fusion region of the initial panoramic video data, fixed architectural features are extracted as physical boundary features, such as the edge of an elevator door frame and the texture of a stair handrail. Their pixel coordinates (u1, v1) and BIM coordinates are recorded. In the boundary fusion region of the blind spot video data, virtual features with the same BIM coordinates are extracted as corresponding virtual boundary features, such as virtual elevator door frames and virtual handrails. Their pixel coordinates (u2, v2) are recorded to ensure that the feature types of the two are consistent (e.g., if the actual door frame has a metal texture, the virtual door frame must also match the texture).

[0109] Step S303: Based on the pixel coordinates of the physical boundary features and the pixel coordinates of the virtual boundary features corresponding to the physical boundary features, align the initial panoramic video data and the blind spot video data, and determine the boundary fusion region;

[0110] Understandably, by accurately matching physical boundary features with virtual boundary features, pixel-level alignment of the images at the boundary between blind and non-blind zone videos is ensured, and the fusion area at the boundary is redefined, i.e., the boundary fusion area.

[0111] It should be understood that if there is a moving target at the boundary (such as a person entering an elevator), it is necessary to ensure that the position of the target in the last frame of the initial panoramic video data (e.g., half of the body in the elevator lobby) and the position in the first frame of the blind spot video data (e.g., half of the body inside the elevator car) are continuous in BIM coordinates (deviation ≤ 10cm). This can be achieved by fine-tuning the initial coordinates of the target in the blind spot video data. When there is a moving target at the physical boundary, transition frames are added between the initial panoramic video data and the blind spot video data. Specifically, three transition frames are inserted between the frame before the boundary of the initial panoramic video data (the last frame before the moving target enters the blind spot) and the frame after the boundary of the blind spot video data (the first frame after the moving target enters the blind spot). The first transition frame: 70% of the pixel value of the initial panoramic video data and 30% of the pixel value of the blind spot video data; the second transition frame: 50% of the pixel value of the initial panoramic video data and 50% of the pixel value of the blind spot video data; the third transition frame: 30% of the pixel value of the initial panoramic video data and 70% of the pixel value of the blind spot video data. The moving target in the transition frame needs to maintain motion continuity (such as a smooth arm swing posture), and the intermediate posture is generated by a frame interpolation algorithm.

[0112] Step S304: Based on the fusion weight, the pixels in the boundary fusion area are weighted and fused to obtain the panoramic video data of the building. The fusion weight is determined based on the distance between the pixel and the physical boundary.

[0113] Understandably, for pixels at the same BIM coordinates within the boundary fusion area, both those in the initial panoramic video data and those in the blind spot video data, corresponding fusion weights are assigned to perform weighted fusion of pixel values, ultimately yielding video data representing the complete view of the building—i.e., panoramic video data. The fusion weights are determined based on the distance between the pixel and the physical boundary. The closer to the boundary, the fusion weight of pixels in the initial panoramic video data linearly decreases from 1.0 to 0.5, while the fusion weight of pixels in the blind spot video data linearly increases from 0.0 to 0.5.

[0114] It should be understood that before fusion, the global average brightness (L_real) and color average (R_real, G_real, B_real) of the initial panoramic video data can be calculated. For blind spot video data, the light source intensity of the virtual scene model (e.g., L_virtual = L_real ± 5%) and texture RGB values ​​(e.g., R_virtual = R_real ± 3) are adjusted to ensure that the difference between the two is ≤ 5%. After fusion, histogram equalization is performed on the panoramic video data to eliminate brightness jumps between regions (e.g., reducing the brightness difference between the corridor and the elevator lobby from 20% to 5%).

[0115] Step S40: Based on the anomaly detection rule base, perform anomaly detection on the panoramic video data to determine the pixel coordinates of the anomaly sites;

[0116] It should be noted that the anomaly detection rule base is divided into 4 major categories and 12 subcategories, each containing clear anomaly definitions and detection dimensions. The first-level categories (majors) include abnormal personnel behavior, abnormal equipment status, environmental hazards, and spatial correlation anomalies. The second-level categories (subcategories) of abnormal personnel behavior include falling, lingering, running, and climbing. Core detection dimensions (feature indicators) include posture characteristics (skeletal angles) and motion parameters (speed, dwell time), for example: a person falling in a stairwell (abnormal skeletal angle). The second-level categories of abnormal equipment status include elevator jamming, pipe leaks, and electrical faults. Feature indicators include equipment shape (deformation, displacement) and operating parameters (vibration, temperature), for example: abnormal vibration of an air conditioner outdoor unit (displacement > 5cm), elevator jamming (car BIM coordinates change < 0.5m within 5 seconds, normal operating speed 1-2m / s). The second-level categories of environmental hazards include smoke, flames, water accumulation, and abnormal lighting. Feature indicators include environmental characteristics (color, texture), sensor data (concentration, temperature), and light and shadow parameters (brightness change rate), for example: smoke appearing in a machine room (RGB values ​​match smoke characteristics). The secondary categories of spatial association anomalies include area intrusion, multi-target conflict, and access control anomalies. The characteristic indicators include spatial coordinates (BIM coordinates) and target interaction relationships (distance, time difference). For example, personnel entering a restricted area (BIM coordinates exceed the permission range).

[0117] It is understandable that an anomaly detection rule base is used to detect anomalies in panoramic video data, determine whether anomalies exist, and if an anomaly occurs, obtain the pixel coordinates of the anomaly location.

[0118] Step S50: Based on the perspective transformation matrix between BIM coordinates and pixel coordinates, convert the pixel coordinates of the abnormal site into the BIM coordinates of the abnormal site.

[0119] Understandably, the perspective transformation matrix can be used to convert the pixel coordinates of abnormal sites into corresponding BIM coordinates, thereby accurately locating the position of abnormal events in the three-dimensional space of the building.

[0120] Step S60: Based on the BIM coordinates of the abnormal site, plan the corresponding response route in the BIM model of the building.

[0121] In one feasible implementation, step S60 may include steps S601 to S603:

[0122] Step S601: Extract key parameters of accessible passages from the BIM model of the building, associate emergency resources with the accessible passages, and mark fixed obstacles;

[0123] It should be noted that key parameters include at least geometric parameters, access attributes, and associated facilities. All accessible passageways (such as corridors, staircases, elevators, and fire exits) are extracted from the BIM model, and key parameters are recorded. For example, geometric parameters include: passageway width (e.g., 1.8m for a corridor, 2.4m for a fire exit), length (e.g., 50m for the east corridor on the 3rd floor), and BIM coordinate range (e.g., X: 15-30m, Y: 8-10m, Z: 9m); access attributes include: whether it is an emergency exit (fire exits are marked as priority access), whether it is accessible (e.g., whether there are ramps, whether elevators are wheelchair-friendly), and direction of travel (e.g., one-way / two-way stairs); associated facilities include: BIM coordinates and operating parameters of vertical access facilities within the passageway, such as elevators (rated load, operating speed, floors they stop at), and stairs (number of steps, slope).

[0124] Additionally, it should be noted that fixed obstacles that are impassable in the BIM model (walls, beams, columns, equipment rooms, large equipment) and reserved marks for dynamic obstacles (such as temporary construction areas and movable equipment) are marked and their BIM coordinate boundaries are stored.

[0125] It is understood that in this embodiment, the BIM coordinates of emergency resources within the building are extracted and associated with the passage information. For example, rescue equipment includes AEDs, fire hydrants, fire extinguishers, emergency lighting, and demolition tools; service facilities include elevators (especially fire elevators), stairwells, emergency broadcasts, and medical points. Subsequently, a mapping table of "resource type-BIM coordinates-passage ID" is established.

[0126] Step S602: Based on the anomaly type of the abnormal site and the response subject, formulate route constraints;

[0127] It should be noted that route constraints should be formulated based on the type of anomaly (such as fire, person falling, equipment failure) and the responding entity (maintenance personnel, firefighters, emergency medical personnel).

[0128] Understandably, if the anomaly type is a person falling (first aid) and the respondent is medical personnel, the route constraints are: prioritize accessible pathways (slope ≤ 5%), prioritize the use of medical elevators, and avoid narrow passages (width < 1.2m); if the anomaly type is a fire hazard and the respondent is firefighters, the route constraints are: mandatory use of fire exits, avoid smoke-affected areas (high temperature / smoke diffusion areas marked in the BIM model), and prohibit the use of ordinary elevators; if the anomaly type is equipment failure (maintenance) and the respondent is maintenance personnel, the route constraints are: prioritize the nearest passage, use freight elevators (if tools need to be carried), and avoid densely populated office areas (to reduce interference); if the anomaly type is elevator entrapment and the respondent is maintenance personnel, the route constraints are: prioritize direct access to the elevator machine room / car stopping floor, and the route must pass through the elevator emergency control panel location.

[0129] Step S603: Based on the route constraints, the key parameters of the passable passage, the emergency resources, and the fixed obstacles, plan a response route.

[0130] Understandably, the optimal route that meets the route constraints is calculated by taking the BIM coordinates of the response starting point (such as the current location of maintenance personnel or the fire control room) and the BIM coordinates of the abnormal location as input, combined with the key parameters of the passable passage, emergency resources and fixed obstacles, and this route is used as the final response route.

[0131] This embodiment provides an intelligent response method based on multi-view video fusion of buildings. It acquires surveillance video data collected by cameras within the building, and fuses surveillance video data with overlapping views using a weighted smoothing fusion strategy to obtain initial panoramic video data. Based on the basic spatial parameters and supplementary features of the camera's corresponding blind spots, blind spot video data is generated. Using a seamless virtual-real connection mechanism, the initial panoramic video data and the blind spot video data are fused to generate panoramic video data of the building. Based on an anomaly detection rule base, anomalies are detected in the panoramic video data to determine the pixel coordinates of anomaly points. Using a perspective transformation matrix between BIM coordinates and pixel coordinates, the pixel coordinates of the anomaly points are converted to BIM coordinates. Based on the BIM coordinates of the anomaly points, a corresponding response route is planned in the building's BIM model. This embodiment can achieve millimeter-level binding between video footage and BIM 3D coordinates, significantly improving the spatial positioning accuracy of abnormal events in buildings. It can accurately locate the position of abnormal events in the building's 3D space, effectively improving response efficiency.

[0132] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2Step S20 may include steps S201 to S204:

[0133] Step S201: Extract the basic spatial parameters of the monitoring blind spot from the building's BIM model;

[0134] It should be noted that the basic spatial parameters include at least dimensional parameters, spatial constraint parameters, and coordinate anchor points. Dimensional parameters should include at least the elevator car dimensions (e.g., 1.5m long × 1.8m wide × 2.5m high), the number of steps at each stairwell corner (e.g., 12 steps), the step height at each stairwell corner (e.g., 15cm), the step width at each stairwell corner (e.g., 30cm), the dimensions of the equipment room corner (e.g., 2m long × 1.5m wide × 3m high), and the coordinates of the fixed equipment in the equipment room corner. Spatial constraint parameters should include at least... This includes the coordinates of movement boundaries (e.g., elevators can only move up and down along the Z-axis with a maximum speed of 2m / s, stairwells can only move along the steps and are prohibited from passing through walls) and fixed components (e.g., control panels and handrails inside elevators, railings and wall markings in stairwells). Coordinate anchor points are set at the entrance or exit of the monitoring blind spot. Usually, more than three coordinate anchor points are set at the entrance / exit of the blind spot, such as the ground markings aligned with the elevator hall and car door, and the corner of the wall at the stair landing. Their BIM coordinates are known and serve as the spatial alignment reference between the subsequent virtual image and the actual video.

[0135] Step S202: Obtain supplementary features of the monitoring blind spot;

[0136] It should be noted that supplementary features include at least texture features, lighting features, and environmental noise features.

[0137] Understandably, a laser scanner is used to perform a 360° scan of the blind area, acquiring texture images of walls, floors, and fixed equipment, such as the metal interior wall texture of an elevator car and the marble floor texture of a stairwell, with a resolution ≥2K. These texture images are then used as texture features. Light sensors (sampling frequency 1Hz) are deployed within the blind area to continuously collect light intensity (e.g., 500 lux for elevator interior lighting, 800 lux for natural daylight in the stairwell, and 50 lux for emergency lights at night) and light source direction (e.g., top light for elevator ceiling lights and side light for stairwell wall lights) at different times of day for 24 hours. The collected light intensity and light source direction are then used as light features. The inherent environmental noise of the blind area is recorded as environmental noise features, such as the mechanical noise of the elevator and echoes in the stairwell, for audio synchronization in subsequent virtual scenes.

[0138] Step S203: Based on the supplementary features and basic spatial parameters of the monitoring blind spot, generate an optimized virtual scene model of the monitoring blind spot;

[0139] Understandably, using BIM coordinate anchor points as a reference, a 3D mesh of the monitoring blind spots is reconstructed in the engine according to basic spatial parameters (e.g., the hexahedral mesh of the elevator car, the mesh of the stair treads). Scanned texture features are then overlaid onto the corresponding mesh surfaces (e.g., the elevator interior wall texture is overlaid onto the hexahedron of the car, and the floor texture is overlaid onto the stair treads). Next, virtual light sources are set based on lighting characteristics (e.g., a top light source is added inside the elevator with a brightness of 500 lux, and a side light source is added to the stairwell with its brightness dynamically changing over time to simulate natural light / emergency lighting), ensuring that the lighting effects of the virtual scene are consistent with reality. Finally, environmental noise features are added to generate an optimized virtual scene model.

[0140] It should be understood that if there are no moving targets (people, equipment, etc.) in the blind spot, virtual images can be generated directly according to the optimized virtual scene model, thereby generating blind spot video data. If there are moving targets in the blind spot, the virtual images need to accurately reproduce the movement state and appearance of the moving targets in the blind spot. Therefore, it is necessary to collect complete data containing the movement state and appearance of the moving targets through other cameras, i.e., relevant video data, and extract the corresponding features, i.e., core features. The core features include at least movement features, type features, and appearance features.

[0141] Step S204: When a moving target exists in the monitoring blind spot, generate blind spot video data of the monitoring blind spot based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot.

[0142] It should be noted that type features are used to characterize the type of mobile target, such as: people, equipment, other objects (e.g., suitcases), and to refine subclasses. For example, people: adults / children / elderly; equipment: wheelchairs / toolboxes / cleaning equipment. Appearance features are used to characterize the appearance information of mobile targets. For example, for people: clothing color (RGB value), hairstyle (long / short hair), carried items (e.g., backpacks, document bags), height (converted from BIM coordinates, error ±5cm); for equipment: shape (cuboid / cylinder), color, and size (length, width, height, calculated based on BIM coordinates).

[0143] Additionally, it should be noted that the motion data of the target in the 3 seconds before entering the blind zone is recorded by adjacent cameras at a sampling frequency of 10Hz (once every 0.1 seconds) as motion features. This includes at least real-time BIM coordinates (obtaining the BIM coordinates of the moving target in the BIM coordinate system through a perspective transformation matrix and adding a timestamp), motion parameters (instantaneous velocity, direction of motion, and acceleration of the moving target), and motion state (whether it is uniform speed, whether it is turning, and whether it is stopped).

[0144] In one feasible implementation, step S204 may include: generating a virtual appearance model of the moving target based on its type and appearance features; predicting the trajectory of the moving target based on its motion features to obtain a predicted trajectory; generating continuous virtual video frames based on the optimized virtual scene model of the monitoring blind spot, the virtual appearance model of the moving target, and the predicted trajectory; adjusting the viewing angle of the continuous virtual video frames to the expected viewing angle of the monitoring blind spot, and adding interaction features between the moving target and the monitoring blind spot to the continuous virtual video frames to form blind spot video data of the monitoring blind spot.

[0145] It should be noted that traditional trajectory prediction does not consider the physical boundaries of the blind zone, which can easily lead to unreasonable movements such as passing through walls. This embodiment incorporates blind zone spatial constraints (e.g., the elevator can only move along the Z-axis) into the trajectory prediction model. Model input: Motion data (coordinates, velocity, direction, acceleration, etc.) of the moving target 3 seconds before entering the blind zone and blind zone spatial constraint parameters (determined according to the motion boundary, such as the elevator's Z-axis and maximum speed of 2 m / s); Constraint embedding: A spatial constraint loss function is added to the model output layer. When the predicted coordinates exceed the physical boundaries of the blind zone (e.g., outside the elevator car) or the motion parameters violate the constraints (e.g., the elevator's lateral velocity > 0), the loss function value increases sharply, forcing the model to correct the prediction results; Loss function formula: ,in, For trajectory prediction error, Penalties for violations of spatial constraints Weighting coefficient (e.g., 5); Predicted output: Real-time predicted trajectory of the moving target within the monitoring blind zone, outputting a BIM 3D coordinate every 0.1 seconds until the target leaves the blind zone (triggered by actual detection by the blind zone exit camera).

[0146] Understandably, based on the type and appearance characteristics of the moving target, a virtual appearance model consistent with the moving target is constructed in the optimized virtual scene model of the monitoring blind spot. For example, personnel modeling involves calling a 3D human model library, matching a basic model based on height and clothing color, adding details through texture mapping (e.g., texture of a red shirt, model of a black backpack), and binding skeletal animation (e.g., walking, standing postures, automatically switching based on the predicted trajectory speed). Equipment modeling involves matching or parametrically generating a 3D model from the model library based on the equipment's shape, size, and color (e.g., a cuboid toolbox: 0.5m long × 0.3m wide × 0.2m high, with a yellow shell), and adding physical properties (e.g., conforming to the ground when stationary, sliding along the trajectory when moving). The origin of the virtual appearance model of the moving target is bound to the coordinates of the predicted trajectory to ensure that the position of the virtual appearance model in the optimized virtual scene model is consistent with the predicted trajectory.

[0147] It should be understood that, based on optimized virtual scene models, virtual appearance models, and predicted trajectories, a continuous stream of virtual video frames is generated through the rendering engine to ensure that the frame rate matches that of the actual cameras. The viewing angle of the virtual video frames strictly simulates the expected viewing angle of the camera at the blind spot exit, i.e., the shooting angle of the camera (such as an elevator lobby camera) after the moving target leaves the blind spot, avoiding sudden changes in viewing angle when the moving target leaves the blind spot. During each frame rendering, the position (based on the predicted trajectory), posture (such as leg animation when a person walks, or tilt angle when equipment moves), and lighting effects (such as the shadow of the moving target under lights, matching the lighting characteristics of the blind spot) of the moving target are updated synchronously. Environmental interaction effects are added, i.e., the interaction features between the moving target and the monitoring blind spot, such as: the hand movements of a person pressing a button in an elevator (triggered based on the predicted trajectory when approaching the control panel), and slight friction marks between the equipment and the ground when moving (only briefly displayed in the virtual image to enhance realism).

[0148] Furthermore, in one feasible implementation, step S204 may include: extracting the core features of the moving target from the relevant video data of the moving target based on an improved YOLOv8 target detection model.

[0149] It should be noted that the reference Figure 3 The improved YOLOv8 object detection model includes an input layer, a backbone network, a neck network, and a head network (detection head). The input layer includes a lighting correction submodule (replacing the Mosaic enhancement module in the original model), a multi-scale anchor point adaptation module, and a BIM coordinate pre-association module. The backbone network includes an optimized C2f module (replacing the C2f module in the original model) and an optimized SPPF module (replacing the SPPF module in the original model). The head network includes a newly added feature extraction head, as well as the classification and regression heads from the original model. Based on the original YOLOv8 object detection model, the entire link structure of the input layer, backbone network, head network, and neck network is optimized to enhance the specificity and robustness of feature extraction.

[0150] The input layer is used to preprocess the relevant input video data and input the obtained feature maps into the backbone network. First, the data quality is corrected, and then spatial information is embedded. Structurally, it is divided into three serial sub-modules: a lighting correction sub-module (a lightweight version of Retinex-Net), a multi-scale anchor point adaptation module, and a BIM coordinate pre-association module. To address the variable lighting and shadows of buildings, the lighting correction submodule employs a lightweight structure of 3 convolutional layers + 1 decomposition and fusion layer to avoid complex calculations. Decomposition convolutional layer 1 (3×3 kernel, stride 1, padding=1, SiLU activation function) extracts basic image features. Decomposition convolutional layer 2 (3×3 kernel, stride 1, padding=1, with added normalization processing) separates the illumination component and the reflection component. The illumination adjustment layer uses a fully connected layer to output illumination adjustment coefficients (e.g., low-light scene coefficient = 1.5). The fusion layer (1×1 kernel, stride 1) reconstructs the corrected image (reflection component × adjustment coefficient + illumination component × 0.3). Notably, the convolutional kernels of the decomposition convolutional layers are all 3×3, which balances the receptive field and computational cost. Padding=1 (filling one layer of edge pixels) ensures consistent input and output dimensions. The SiLU activation function is used to enhance non-linear expression. The multi-scale anchor point adaptation module can adapt to the building target distribution without changing the network structure, only by re-clustering the anchor point size. It uses the K-Means algorithm (K=9) for anchor point clustering, and the final grouping corresponds one-to-one with the three output scales of the neck (large, medium, and small targets): large-scale output (80×80) adapts to small targets (equipment parts, fine smoke), medium-scale output (40×40) adapts to medium targets (people, air conditioner outdoor units), and small-scale output (20×20) adapts to large targets (flames, crowds). The BIM coordinate pre-association module can embed spatial information into the input features without increasing the amount of additional computation. It only modifies the feature label format. Specifically, it first maps the pixel coordinates of the input image features to BIM coordinates through a perspective transformation matrix (pre-stored in the model parameters), divides the spatial region according to the BIM coordinates (such as the east corridor on the 3rd floor), generates a 1×1×1 spatial label vector, and concatenates the spatial label vector with the preprocessed image features in the channel dimension to output the feature map.

[0151] The backbone network is used to extract features from the feature maps and inputs the obtained backbone features into the neck network. The backbone network contains 6 optimized C2f modules (C2f-ECA modules) and 1 optimized SPPF module (SPPF+ module). The optimized C2f modules embed the ECA (Efficient Channel Attention Module) mechanism in the residual branches of the original C2f modules. That is, after the output of the last Bottleneck of the branch, an ECA module (without fully connected layers, only 1D convolution layer) is added. The ECA module includes a channel-dimensional global average pooling layer, a 1D convolution layer, and a Sigmoid activation layer to output channel weights. Then, the main path features and the branch weighted features (branch features × channel weights output by the ECA module) are concatenated and processed by convolution (1×1) and standardization to output the corresponding features. The number of Bottlenecks is 8 to balance feature extraction and computation. The convolution kernel of the ECA is k=5, and the attention weights are dynamically adjusted in the range of [0.1, 1.0]. The optimized SPPF module adds a multi-scale fusion layer to the original SPPF (single-scale pooling). The overall structure includes: a global pooling layer for parallel execution of four pooling scales (1×1, 5×5, 9×9, 13×13), all using padding equal to the kernel size / 2 to ensure consistent output size; and a feature fusion layer that concatenates the four pooling results with the original input features channel-wise, followed by convolution (1×1) and normalization, and then weighted fusion (weights [0.4, 0.3, 0.2, 0.1], with the original input weight at 0.4 and pooling feature weights decreasing sequentially) to obtain the final backbone features. The optimized SPPF module can simultaneously capture local details (1×1 pooling) and global context (13×13 pooling), adapting to feature extraction of small (local) and large (global) targets within buildings. Furthermore, the last two convolutional layers of the backbone network (after the output of the optimized C2f module and before the input of the optimized SPPF module) employ depthwise separable convolutions.

[0152] The neck network is used to perform multi-scale feature fusion on the backbone features, and the obtained fused features are input into the head network. During upsampling fusion, a channel attention gate (Conv+Sigmoid) is added, and the weights are adjusted to adjust the contribution of shallow features (such as edge details). Downsampling uses a combination of convolution (3×3 kernel, stride 2) + max pooling (2×2) to retain more detailed features. During the fusion process, elements are added one by one (rather than concatenated) to reduce the number of channels and reduce the computational cost of the head network.

[0153] The head network is used to output the core features corresponding to the fused features in parallel. The head network has been upgraded from a dual-head output to a triple-head parallel output. The classification head is used to output the category and confidence score. Its structure consists of 3 convolutional layers + Focal Loss + CIoU optimization. The output dimensions are 320×320×12 (small scale), 160×160×12 (medium scale), and 80×80×12 (large scale), with each dimension corresponding to the confidence score of one class of target. The regression head is used to output the bounding box and BIM coordinates. Its structure consists of 3 convolutional layers + coordinate calibration branch. The loss calculation is SIoU Loss (considering orientation, overlap rate, and distance). The output dimensions are 320×320×7 for each scale (4 bounding box parameters + 3 BIM coordinate parameters). The feature extraction head is used to output new detailed features. Its structure consists of 3 layers of convolution + global average pooling. The loss calculation is Triplet Loss (anchor feature - positive sample feature distance < anchor feature - negative sample feature distance + margin = 0.5). The output dimension is: each detected target corresponds to a multi-dimensional vector, which includes appearance, shape, and detailed features (such as personnel clothing and abnormal parts of equipment).

[0154] This embodiment provides an intelligent response method based on multi-view video fusion of buildings. It parses the basic spatial parameters of monitoring blind spots from the building's BIM model; obtains supplementary features of the monitoring blind spots; and generates an optimized virtual scene model of the monitoring blind spots in the building's BIM model based on the supplementary features and basic spatial parameters. When a moving target exists in the monitoring blind spot, it generates blind spot video data based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot. This embodiment can achieve millimeter-level binding between video footage and BIM 3D coordinates, significantly improving the spatial positioning accuracy of abnormal events in buildings. It can accurately locate the position of abnormal events in the building's 3D space, effectively improving response efficiency.

[0155] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the intelligent response method based on multi-view video fusion of buildings in this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0156] This application also provides an intelligent response system based on multi-view video fusion of buildings, please refer to... Figure 4 The intelligent response system based on multi-view video fusion of buildings includes:

[0157] The video fusion module 10 is used to acquire surveillance video data collected by cameras in the building. Based on a weighted smooth fusion strategy, it fuses surveillance video data with overlapping viewpoints to obtain initial panoramic video data.

[0158] The video fusion module 10 is also used to generate blind spot video data based on the basic spatial parameters and supplementary features of the monitoring blind spot corresponding to the camera;

[0159] The video fusion module 10 is also used to fuse the initial panoramic video data and the blind spot video data based on a seamless virtual-real connection mechanism to generate panoramic video data of the building.

[0160] Anomaly response module 20 is used to perform anomaly detection on the panoramic video data based on anomaly judgment rule base and determine the pixel coordinates of the anomaly site;

[0161] The anomaly response module 20 is also used to convert the pixel coordinates of the anomaly site into the BIM coordinates of the anomaly site based on the perspective transformation matrix between BIM coordinates and pixel coordinates.

[0162] The anomaly response module 20 is also used to plan a corresponding response route in the BIM model of the building based on the BIM coordinates of the anomaly site.

[0163] In one feasible implementation, the video fusion module 10 is further configured to acquire surveillance video data collected by cameras in the building, and construct a perspective transformation matrix between BIM coordinates and pixel coordinates based on the BIM model of the building and the surveillance video data collected by cameras in the building.

[0164] Surveillance video data with overlapping viewpoints are used as overlapping surveillance video data. Based on a multi-feature point fusion matching strategy, corresponding common feature points are extracted from the overlapping viewpoint scenes of the overlapping surveillance video data.

[0165] Based on the perspective transformation matrix, the pixel coordinates of the common feature points are converted into the BIM coordinates of the common feature points;

[0166] Based on the BIM coordinates and pixel coordinates of the common feature points, the coordinate offset of the overlapping viewpoint of the overlapping monitoring video data is determined.

[0167] Based on the coordinate offset of the overlapping viewpoints of the overlapping surveillance video data, the overlapping surveillance video data is smoothly stitched together to obtain initial panoramic video data.

[0168] In one feasible implementation, the video fusion module 10 is further configured to acquire the BIM model of the building and select building reference points in the BIM model;

[0169] Based on the building reference point, the BIM coordinates of the center point of the spatial calibration target are obtained. The spatial calibration target is installed directly in front of the camera. A QR code pattern is provided in the middle of the spatial calibration target. The QR code pattern stores camera identification information, installation angle information and installation height information. Multiple infrared marker points are set around the edge of the spatial calibration target at preset intervals.

[0170] The image of the spatial calibration target is extracted from the surveillance video data collected by the camera, and the pixel coordinates of the infrared marker points are extracted from the image of the spatial calibration target.

[0171] Based on the pixel coordinates of the infrared marker points, determine the pixel coordinates of the center point of the spatial calibration target;

[0172] Extract the corresponding QR code pattern from the image of the spatial calibration target. Based on the extracted QR code pattern, match the pixel coordinates of the center point of the spatial calibration target with the BIM coordinates of the center point of the spatial calibration target to determine the matching coordinate pair.

[0173] Based on the matched coordinate pairs, a perspective transformation matrix between BIM coordinates and pixel coordinates is constructed.

[0174] In one feasible implementation, the video fusion module 10 is further used to align the overlapping monitoring video data based on the coordinate offset of the overlapping viewpoints of the overlapping monitoring video data.

[0175] Based on the aligned overlapping surveillance video data, update the BIM coordinates of the overlapping viewpoints, and based on the updated BIM coordinates of the overlapping viewpoints, calculate the spatial distance between the overlapping point in the overlapping viewpoints and the corresponding camera in the overlapping viewpoints.

[0176] Based on the spatial distances corresponding to the overlapping points in the overlapping viewpoints, corresponding smoothing weights are assigned to the overlapping points in the overlapping viewpoints.

[0177] Based on the smoothing weight of the overlapping points in the overlapping viewpoints, the fused pixel value corresponding to the overlapping points is calculated, and the fused pixel value includes at least the fused color value and the fused brightness value;

[0178] Based on the fused pixel values ​​corresponding to the overlapping sites, spliced ​​video data corresponding to the overlapping monitoring video data is generated;

[0179] Initial panoramic video data is generated based on non-overlapping surveillance video data and stitched video data corresponding to overlapping surveillance video data.

[0180] In one feasible implementation, the video fusion module 10 is further configured to parse the basic spatial parameters of the monitoring blind spot from the building's BIM model. The basic spatial parameters include at least size parameters, spatial constraint parameters, and coordinate anchor points. The size parameters include at least the elevator car size, the number of steps at the stairwell corner, the step height at the stairwell corner, the step width at the stairwell corner, the corner size of the equipment room, and the position coordinates of the fixed equipment in the corner of the equipment room. The spatial constraint parameters include at least the motion boundary and the position coordinates of the fixed components. The coordinate anchor points are set at the entrance or exit of the monitoring blind spot.

[0181] Obtain supplementary features of the monitoring blind spot, wherein the supplementary features include at least texture features, light features, and environmental noise features;

[0182] Based on the supplementary features and basic spatial parameters of the monitoring blind spot, an optimized virtual scene model of the monitoring blind spot is generated;

[0183] When a moving target exists in the monitoring blind spot, blind spot video data of the monitoring blind spot is generated based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot. The core features include at least motion features, type features and appearance features.

[0184] In one feasible implementation, the video fusion module 10 is further configured to generate a virtual appearance model of the moving target based on the type and appearance features of the moving target;

[0185] Based on the motion characteristics of the moving target, the trajectory of the moving target is predicted to obtain the predicted trajectory;

[0186] Based on the optimized virtual scene model of the monitoring blind spot and the virtual appearance model and predicted trajectory of the moving target, continuous virtual video frames are generated.

[0187] The viewing angle of the continuous virtual video frames is adjusted to the expected viewing angle of the monitoring blind spot, and the interaction features between the moving target and the monitoring blind spot are added to the continuous virtual video frames to form the blind spot video data of the monitoring blind spot.

[0188] In one feasible implementation, the video fusion module 10 is further configured to extract the core features of the moving target from the relevant video data of the moving target based on an improved YOLOv8 target detection model. The improved YOLOv8 target detection model includes an input layer, a backbone network, a neck network, and a head network. The input layer includes a lighting correction submodule (replacing the Mosaic enhancement module in the original model), a multi-scale anchor point adaptation module, and a BIM coordinate pre-association module. The backbone network includes an optimized C2f module (replacing the C2f module in the original model) and an optimized SPPF module (replacing the SPPF module in the original model). The head network includes a newly added feature extraction head and the classification and regression heads from the original model. The input layer preprocesses the input relevant video data and inputs the obtained feature map into the backbone network. The backbone network extracts features from the feature map and inputs the obtained backbone features into the neck network. The neck network performs multi-scale feature fusion on the backbone features and inputs the obtained fused features into the head network. The head network outputs the core features corresponding to the fused features in parallel.

[0189] In one feasible implementation, the anomaly response module 20 is further configured to extract physical boundaries from the initial panoramic video data and the blind spot video data, and delineate an initial boundary fusion region based on the physical boundaries;

[0190] Extract the physical boundary features and the corresponding virtual boundary features from the initial boundary fusion region;

[0191] Based on the pixel coordinates of the physical boundary features and the pixel coordinates of the virtual boundary features corresponding to the physical boundary features, the initial panoramic video data and the blind spot video data are aligned, and the boundary fusion region is determined. When there is a moving target at the physical boundary, a transition frame is added between the initial panoramic video data and the blind spot video data.

[0192] Based on the fusion weight, the pixels in the boundary fusion area are weighted and fused to obtain the panoramic video data of the building. The fusion weight is determined based on the distance between the pixel and the physical boundary.

[0193] In one feasible implementation, the anomaly response module 20 is further configured to extract key parameters of accessible passages from the BIM model of the building, associate emergency resources with the accessible passages, and mark fixed obstacles. The key parameters include at least geometric parameters, access attributes, and associated facilities.

[0194] Based on the anomaly type and response subject of the aforementioned anomaly sites, route constraints are formulated;

[0195] Based on the route constraints, the key parameters of the passable passage, the emergency resources, and the fixed obstacles, a response route is planned.

[0196] The intelligent response system based on multi-view video fusion of buildings provided in this application, employing the intelligent response method based on multi-view video fusion of buildings in the above embodiments, can solve the technical problem that traditional solutions struggle to pinpoint the precise location of abnormal events in the three-dimensional space of a building, thus affecting response efficiency. Compared with the prior art, the beneficial effects of the intelligent response system based on multi-view video fusion of buildings provided in this application are the same as those of the intelligent response method based on multi-view video fusion of buildings provided in the above embodiments, and other technical features of the intelligent response system based on multi-view video fusion of buildings are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0197] This application provides an intelligent response device based on multi-view video fusion of buildings. The intelligent response device based on multi-view video fusion of buildings includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the intelligent response method based on multi-view video fusion of buildings in the above embodiment 1.

[0198] The following is for reference. Figure 5 This document illustrates a structural schematic diagram of an intelligent response device suitable for implementing the building multi-view video fusion embodiments of this application. The intelligent response device based on building multi-view video fusion in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The intelligent response device based on multi-view video fusion of buildings shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0199] like Figure 5As shown, the intelligent response device based on building multi-view video fusion may include a processing system 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage system 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the intelligent response device based on building multi-view video fusion. The processing system 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input systems 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output systems 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage systems 1003 including, for example, magnetic tapes, hard disks, etc.; and communication systems 1009. Communication system 1009 allows the intelligent response device based on building multi-view video fusion to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows an intelligent response device based on building multi-view video fusion with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0200] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication system, or installed from storage system 1003, or installed from ROM 1002. When the computer program is executed by processing system 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0201] The intelligent response device based on multi-view video fusion of buildings provided in this application, employing the intelligent response method based on multi-view video fusion of buildings in the above embodiments, can solve the technical problem that traditional solutions struggle to pinpoint the precise location of abnormal events in the three-dimensional space of a building, thus affecting response efficiency. Compared with the prior art, the beneficial effects of the intelligent response device based on multi-view video fusion of buildings provided in this application are the same as those of the intelligent response method based on multi-view video fusion of buildings provided in the above embodiments, and other technical features in this intelligent response device based on multi-view video fusion of buildings are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0202] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0203] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0204] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the intelligent response method based on multi-view video fusion of buildings in the above embodiments.

[0205] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0206] The aforementioned computer-readable storage medium may be included in a smart response device based on building multi-view video fusion; or it may exist independently and not be assembled into a smart response device based on building multi-view video fusion.

[0207] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by an intelligent response device based on multi-view video fusion of buildings, the intelligent response device performs the following actions: acquires surveillance video data collected by cameras in the building; fuses surveillance video data with overlapping views based on a weighted smoothing fusion strategy to obtain initial panoramic video data; generates blind spot video data based on the basic spatial parameters and supplementary features of the corresponding blind spots of the cameras; fuses the initial panoramic video data and the blind spot video data based on a seamless virtual-real connection mechanism to generate panoramic video data of the building; performs anomaly detection on the panoramic video data based on an anomaly judgment rule base to determine the pixel coordinates of the anomaly points; converts the pixel coordinates of the anomaly points to BIM coordinates of the anomaly points based on the perspective transformation matrix between BIM coordinates and pixel coordinates; and plans the corresponding response route in the BIM model of the building based on the BIM coordinates of the anomaly points.

[0208] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0209] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0210] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0211] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described intelligent response method based on multi-view video fusion of buildings. This solves the technical problem that traditional solutions struggle to pinpoint the precise location of abnormal events in the three-dimensional space of a building, thus affecting response efficiency. Compared to existing technologies, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the intelligent response method based on multi-view video fusion of buildings provided in the above embodiments, and will not be elaborated upon here.

[0212] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the intelligent response method based on multi-view video fusion of buildings as described above.

[0213] The computer program product provided in this application can solve the technical problem that traditional solutions struggle to pinpoint the precise location of abnormal events in the three-dimensional space of a building, thus affecting response efficiency. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the intelligent response method based on multi-view video fusion of buildings provided in the above embodiments, and will not be elaborated upon here.

[0214] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A smart response method based on multi-view video fusion of buildings, characterized in that, The method includes: Acquire surveillance video data collected by cameras in buildings, and fuse surveillance video data with overlapping viewpoints based on a weighted smooth fusion strategy to obtain initial panoramic video data; Based on the basic spatial parameters and supplementary features of the monitoring blind zone corresponding to the camera, video data of the blind zone is generated. The supplementary features include at least texture features, light features and environmental noise features. The texture features are the texture images of the walls, ground and fixed equipment in the monitoring blind zone. The light features are the light intensity and light source direction of the monitoring blind zone at different times. The environmental noise features are the inherent environmental noise of the monitoring blind zone. Based on a seamless virtual-real integration mechanism, the initial panoramic video data and the blind spot video data are fused to generate panoramic video data of the building. Based on the anomaly detection rule base, anomaly detection is performed on the panoramic video data to determine the pixel coordinates of the anomaly sites; Based on the perspective transformation matrix between BIM coordinates and pixel coordinates, the pixel coordinates of the abnormal site are converted into the BIM coordinates of the abnormal site. Based on the BIM coordinates of the anomaly sites, corresponding response routes are planned in the BIM model of the building.

2. The method as described in claim 1, characterized in that, The steps of acquiring surveillance video data collected by cameras in buildings and fusing surveillance video data with overlapping viewpoints based on a weighted smoothing fusion strategy to obtain initial panoramic video data include: Acquire surveillance video data collected by cameras in the building, and construct a perspective transformation matrix between BIM coordinates and pixel coordinates based on the building's BIM model and the surveillance video data collected by the cameras in the building. Surveillance video data with overlapping viewpoints are used as overlapping surveillance video data. Based on a multi-feature point fusion matching strategy, corresponding common feature points are extracted from the overlapping viewpoint scenes of the overlapping surveillance video data. Based on the perspective transformation matrix, the pixel coordinates of the common feature points are converted into the BIM coordinates of the common feature points; Based on the BIM coordinates and pixel coordinates of the common feature points, the coordinate offset of the overlapping viewpoint of the overlapping monitoring video data is determined. Based on the coordinate offset of the overlapping viewpoints of the overlapping surveillance video data, the overlapping surveillance video data is smoothly stitched together to obtain initial panoramic video data.

3. The method as described in claim 2, characterized in that, The steps for constructing the perspective transformation matrix between BIM coordinates and pixel coordinates based on the building's BIM model and surveillance video data collected by cameras in the building include: Obtain the BIM model of the building and select the building reference point in the BIM model; Based on the building reference point, the BIM coordinates of the center point of the spatial calibration target are obtained. The spatial calibration target is installed directly in front of the camera. A QR code pattern is provided in the middle of the spatial calibration target. The QR code pattern stores camera identification information, installation angle information and installation height information. Multiple infrared marker points are set around the edge of the spatial calibration target at preset intervals. The image of the spatial calibration target is extracted from the surveillance video data collected by the camera, and the pixel coordinates of the infrared marker points are extracted from the image of the spatial calibration target. Based on the pixel coordinates of the infrared marker points, determine the pixel coordinates of the center point of the spatial calibration target; Extract the corresponding QR code pattern from the image of the spatial calibration target. Based on the extracted QR code pattern, match the pixel coordinates of the center point of the spatial calibration target with the BIM coordinates of the center point of the spatial calibration target to determine the matching coordinate pair. Based on the matched coordinate pairs, a perspective transformation matrix between BIM coordinates and pixel coordinates is constructed.

4. The method as described in claim 2, characterized in that, The step of smoothly stitching together the overlapping surveillance video data based on the coordinate offset of the overlapping viewpoints to obtain the initial panoramic video data includes: Align the overlapping surveillance video data based on the coordinate offset of the overlapping viewpoints of the overlapping surveillance video data. Based on the aligned overlapping surveillance video data, update the BIM coordinates of the overlapping viewpoints, and based on the updated BIM coordinates of the overlapping viewpoints, calculate the spatial distance between the overlapping point in the overlapping viewpoints and the corresponding camera in the overlapping viewpoints. Based on the spatial distances corresponding to the overlapping points in the overlapping viewpoints, corresponding smoothing weights are assigned to the overlapping points in the overlapping viewpoints. Based on the smoothing weight of the overlapping points in the overlapping viewpoints, the fused pixel value corresponding to the overlapping points is calculated, and the fused pixel value includes at least the fused color value and the fused brightness value; Based on the fused pixel values ​​corresponding to the overlapping sites, spliced ​​video data corresponding to the overlapping monitoring video data is generated; Initial panoramic video data is generated based on non-overlapping surveillance video data and stitched video data corresponding to overlapping surveillance video data.

5. The method as described in claim 1, characterized in that, The step of generating blind spot video data based on the basic spatial parameters and supplementary features of the blind spot corresponding to the camera includes: The basic spatial parameters of the monitoring blind spot are extracted from the building's BIM model. The basic spatial parameters include at least dimensional parameters, spatial constraint parameters, and coordinate anchor points. The dimensional parameters include at least the elevator car dimensions, the number of steps at the corner of the stairwell, the step height at the corner of the stairwell, the step width at the corner of the stairwell, the corner dimensions of the equipment room, and the position coordinates of the fixed equipment in the corner of the equipment room. The spatial constraint parameters include at least the motion boundary and the position coordinates of the fixed components. The coordinate anchor points are set at the entrance or exit of the monitoring blind spot. Obtain supplementary features of the monitoring blind spot, wherein the supplementary features include at least texture features, light features, and environmental noise features; Based on the supplementary features and basic spatial parameters of the monitoring blind spot, an optimized virtual scene model of the monitoring blind spot is generated; When a moving target exists in the monitoring blind spot, blind spot video data of the monitoring blind spot is generated based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot. The core features include at least motion features, type features and appearance features.

6. The method as described in claim 5, characterized in that, The step of generating blind spot video data for the monitoring blind spot based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot includes: Based on the type and appearance characteristics of the moving target, a virtual appearance model of the moving target is generated; Based on the motion characteristics of the moving target, the trajectory of the moving target is predicted to obtain the predicted trajectory; Based on the optimized virtual scene model of the monitoring blind spot and the virtual appearance model and predicted trajectory of the moving target, continuous virtual video frames are generated. The viewing angle of the continuous virtual video frames is adjusted to the expected viewing angle of the monitoring blind spot, and the interaction features between the moving target and the monitoring blind spot are added to the continuous virtual video frames to form the blind spot video data of the monitoring blind spot.

7. The method as described in claim 5, characterized in that, Before the step of generating blind spot video data for the monitoring blind spot based on the core features of the moving target and the optimized virtual scene model of the monitoring blind spot, the following steps are also included: Based on an improved YOLOv8 target detection model, the core features of the moving target are extracted from the relevant video data. The improved YOLOv8 target detection model includes an input layer, a backbone network, a neck network, and a head network. The input layer includes a lighting correction submodule (replacing the Mosaic enhancement module in the original model), a multi-scale anchor point adaptation module, and a BIM coordinate pre-association module. The backbone network includes an optimized C2f module (replacing the C2f module in the original model) and an optimized SPPF module (replacing the SPPF module in the original model). The head network includes a newly added feature extraction head and the classification and regression heads from the original model. The input layer is used to preprocess the input relevant video data, obtaining... The feature map is input into the backbone network, which is used to extract features from the feature map. The obtained backbone features are input into the neck network, which is used to perform multi-scale feature fusion on the backbone features. The obtained fused features are input into the head network, which is used to output the core features corresponding to the fused features in parallel. The optimized C2f module adds an ECA module after the last Bottleneck layer of the residual branch of the original C2f module. The ECA module includes a channel-dimensional global average pooling layer, a one-dimensional convolutional layer, and a Sigmoid activation layer. The channel weights output by the ECA module are used to weight the features output by the residual branch. The optimized SPPF module adds a feature fusion layer to the original SPPF module.

8. The method as described in claim 1, characterized in that, The step of fusing the initial panoramic video data and the blind spot video data based on the virtual-real seamless integration mechanism to generate panoramic video data of the building includes: Physical boundaries are extracted from the initial panoramic video data and the blind spot video data, and an initial boundary fusion region is delineated based on the physical boundaries. Extract the physical boundary features and the corresponding virtual boundary features from the initial boundary fusion region; Based on the pixel coordinates of the physical boundary features and the pixel coordinates of the virtual boundary features corresponding to the physical boundary features, the initial panoramic video data and the blind spot video data are aligned, and the boundary fusion region is determined. When there is a moving target at the physical boundary, a transition frame is added between the initial panoramic video data and the blind spot video data. Based on the fusion weight, the pixels in the boundary fusion area are weighted and fused to obtain the panoramic video data of the building. The fusion weight is determined based on the distance between the pixel and the physical boundary.

9. The method as described in claim 1, characterized in that, The step of planning the corresponding response route in the BIM model of the building based on the BIM coordinates of the anomaly sites includes: Key parameters of accessible passages are extracted from the BIM model of the building, emergency resources are associated with the accessible passages, and fixed obstacles are marked. The key parameters include at least geometric parameters, access attributes, and associated facilities. Based on the anomaly type and response subject of the aforementioned anomaly sites, route constraints are formulated; Based on the route constraints, the key parameters of the passable passage, the emergency resources, and the fixed obstacles, a response route is planned.

10. An intelligent response system based on multi-view video fusion of buildings, characterized in that, The system includes: The video fusion module is used to acquire surveillance video data collected by cameras in buildings. Based on a weighted smooth fusion strategy, it fuses surveillance video data with overlapping viewpoints to obtain initial panoramic video data. The video fusion module is also used to generate blind spot video data based on the basic spatial parameters and supplementary features of the monitoring blind spot corresponding to the camera. The supplementary features include at least texture features, light features and environmental noise features. The texture features are the texture images of the wall, ground and fixed equipment in the monitoring blind spot. The light features are the light intensity and light source direction of the monitoring blind spot at different times. The environmental noise features are the inherent environmental noise of the monitoring blind spot. The video fusion module is also used to fuse the initial panoramic video data and the blind spot video data based on a seamless virtual-real connection mechanism to generate panoramic video data of the building. An anomaly response module is used to perform anomaly detection on the panoramic video data based on an anomaly judgment rule base and determine the pixel coordinates of the anomaly site. The anomaly response module is also used to convert the pixel coordinates of the anomaly site into the BIM coordinates of the anomaly site based on the perspective transformation matrix between BIM coordinates and pixel coordinates. The anomaly response module is also used to plan a corresponding response route in the BIM model of the building based on the BIM coordinates of the anomaly site.

Citation Information

Patent Citations

  • Abnormality positioning method and system based on panoramic video fusion

    CN119625161A

  • Digital twin and panoramic camera mapping method and system based on BIM model

    CN120783011A