Machine vision intelligent identification method and system based on digital building
By integrating sensor data and camera data in the building area for real-time environmental perception, combined with image enhancement, abnormality detection and network optimization, the accuracy and reliability problems of traditional building identification methods in complex environments are solved, and efficient and intelligent building monitoring and risk management are achieved.
Patent Information
- Application Number
- CN202510143333.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional building identification methods rely on manual inspection and manual annotation, resulting in low accuracy and reliability of identification results, and it is difficult to achieve real-time and accurate dynamic detection in complex environments.
By obtaining sensor data and camera data in the building area, the environmental state analysis and dynamic environmental perception are integrated, and the dynamic environment perception map is generated to realize real-time environmental perception of the building area. Combining image adaptive dynamic enhancement, abnormal event detection and risk event identification, automated identification of potential hazards in building areas. Optimize the transmission and storage of video data using network topological analysis and adaptive video compression.
Real-time environmental perception of the building area is realized, the accuracy and reliability of identification results are improved, manual intervention is reduced, monitoring efficiency and response speed is improved, and video data is efficient and stable transmission is ensured.
Smart Images

Figure CN120107782A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video coding technology, and in particular to a machine vision intelligent recognition method and system based on digital architecture. Background Art
[0002] With the rapid development of information technology, the encoding, decoding, compression and decompression technologies of digital video signals play a vital role in multimedia applications, video surveillance, communications and intelligent buildings. As a core component of building intelligence, intelligent monitoring has gradually become an important technical means to improve building management efficiency, ensure safety and provide value-added services. However, traditional building recognition methods usually rely on manual inspections and manual labeling. Especially in the identification and monitoring of building objects in complex environments, the workload of manual inspections is huge and easily interfered by human factors, resulting in low accuracy and reliability of recognition results. In addition, many traditional building detection methods rely on fixed sensors or camera equipment, and their data acquisition methods are single and limited by environmental conditions, making it difficult to achieve real-time and accurate dynamic detection. For example, traditional two-dimensional image recognition technology often has problems such as insufficient recognition accuracy and severe background noise interference when facing complex scenes and changing environments. The defects of these traditional methods make the work efficiency of aspects such as building structure monitoring, hazard source identification, and material management low in complex construction processes, and cannot effectively cope with the multi-dimensional data requirements involved in large-scale construction projects. Summary of the invention
[0003] Based on this, it is necessary for the present invention to provide a machine vision intelligent recognition method and system based on digital buildings to solve at least one of the above technical problems.
[0004] To achieve the above purpose, a machine vision intelligent recognition method based on digital buildings includes the following steps:
[0005] Step S1: acquiring sensor data and camera data of the building area, and performing building area environmental status analysis based on the sensor data of the building area, thereby obtaining building area environmental status data; performing dynamic environmental perception fusion based on the building area environmental status data and the camera data of the building area, thereby obtaining a dynamic environmental perception map of the building area;
[0006] Step S2: performing image adaptive dynamic enhancement on the camera data of the building area according to the dynamic environment perception map of the building area, so as to obtain an enhanced video frame image set of the building area;
[0007] Step S3: performing abnormal event detection on the building area enhanced video frame image set to obtain an abnormal event video frame image set, and performing risk event identification on the abnormal event video frame image set to obtain a risk event video frame image set;
[0008] Step S4: acquiring building area network transmission data, and performing building area network topology analysis on the building area network transmission data, thereby obtaining a building area network topology model; performing network status encoding based on the building area network topology model, thereby obtaining network status encoding data;
[0009] Step S5: performing video feature encoding data on the risk event video frame image set to obtain video feature encoding data, and performing adaptive video compression according to the network status encoding data and the video feature encoding data to obtain building area event video compression data;
[0010] Step S6: Allocate data transmission resources for the compressed video data of the building area event through the building area network topology model, so as to obtain the building area compressed video transmission strategy, and upload it to the building area risk management platform to execute the building area risk event warning task.
[0011] The present invention can capture the dynamic changes of the building environment in real time by fusing the sensor data and camera data of the building area, and realizes the comprehensive perception of the real-time environment of the building area through environmental state analysis and the generation of dynamic environmental perception map. This process not only improves the accuracy of environmental perception, but also helps to eliminate noise interference in complex scenes, and improves the clarity and availability of images. Based on this enhanced video image set, high-quality input can be provided for subsequent abnormal event detection, ensuring the timely identification of potential dangers in the building area. Through the intelligent detection of abnormal events and the identification of risk events, the key moments with potential safety hazards can be automatically extracted from the video stream, further reducing manual intervention, and improving monitoring efficiency and response speed. This process provides strong technical support for real-time risk management and security, and can accurately lock each frame of abnormal video data, thereby improving the accuracy of safety warning. Using the network topology analysis method, the overall situation of the internal network of the building area can be accurately grasped, the demand and pressure of the network bandwidth can be effectively predicted, and the transmission path and resource allocation of the video data can be optimized. Network status coding can adaptively compress the video data based on the current network conditions, avoid transmission delays or screen freezes caused by insufficient bandwidth, and ensure efficient and stable transmission of video data in the building area. Adaptive video compression technology not only ensures the improvement of transmission efficiency, but also avoids excessive compression of video data, ensuring that the quality of video images is not lost. Through the synergy of these steps, the present invention realizes an efficient video data transmission strategy for the building area, which can be dynamically adjusted according to the real-time network conditions to ensure that when a safety risk occurs in the building area, the video data of the risk event can be transmitted to the risk management platform in a timely and accurate manner. Through this intelligent technical process, risk events in the building area can be accurately identified and quickly fed back to the risk management platform, providing decision makers with real-time and reliable data support, thereby achieving efficient and safe supervision during construction and management. It not only enhances the perception of the building environment, but also optimizes the processing and transmission methods of video data, thereby effectively improving the safety monitoring level and management efficiency of the building area, and has important practical application value and broad market prospects.
[0012] Optionally, step S1 specifically includes:
[0013] Step S11: acquiring building area sensor data and building area camera data, and performing data preprocessing on the building area sensor data and the building area camera data respectively, so as to obtain the building area sensor data to be analyzed and the building area camera data to be analyzed;
[0014] Step S12: performing sensor data time series analysis on the sensor data of the building area to be analyzed, thereby obtaining time series sensor data of the building area;
[0015] Step S13: deriving the building area environment state according to the building area time series sensing data, thereby obtaining the building area environment state data;
[0016] Step S14: decoding the camera video stream of the camera data of the building area to be analyzed, thereby obtaining the camera video of the building area, and dividing the camera video of the building area into time window frames, thereby obtaining a continuous video frame image set;
[0017] Step S15: Perform multimodal dynamic environment perception fusion on the building area environmental status data and the continuous video frame image set, so as to obtain a dynamic environment perception map of the building area.
[0018] The present invention can extract high-precision, structured time-series sensor data through preprocessing and time series analysis of sensor data and camera data in the building area, and deduce the environmental state of the building area based on this. This analysis method effectively overcomes the limitation of a single data source that responds slowly to environmental changes, enables the monitoring system to quickly identify changes in environmental states, and improves the sensitivity and accuracy of perception. The decoding of the camera video stream and the time window frame division can construct a continuous video frame image set, providing a reliable basis for subsequent multimodal fusion. Through time window division, the monitoring system can realize segmented processing of video data, which helps to reduce data redundancy and improve the real-time and efficiency of video analysis. Multimodal dynamic environment perception fusion further integrates the advantages of sensor and camera data, and uses the correlation between environmental state data and video images to achieve deep perception of the dynamic environment of the building area. Compared with traditional single-mode perception technology, this method can more comprehensively reflect the actual situation of the building area, especially in complex scenes, it can provide more accurate environmental perception results. This fusion method not only improves the adaptability of the monitoring system in a dynamic environment, but also lays a solid foundation for subsequent tasks such as anomaly detection and event analysis.
[0019] Optionally, step S13 is specifically:
[0020] Step S131: normalizing the time series sensing data of the building area to obtain normalized time series sensing data;
[0021] Step S132: performing long-short term memory time series feature classification on the normalized time series sensor data, thereby obtaining time series periodic sensor data and time series non-periodic sensor data;
[0022] Step S133: performing Bayesian network modeling of building area sensor data according to the time-series periodic sensor data, thereby obtaining a building area normal environment state derivation model, and performing periodic environment state derivation through the building area normal environment state derivation model, thereby obtaining periodic environment state data;
[0023] Step S134: performing principal component analysis on the time series non-periodic sensor data to obtain non-periodic reduced dimension sensor data;
[0024] Step S135: performing joint distribution modeling of building area sensor data on the non-periodic reduced dimension sensor data, thereby obtaining a derivation model of abnormal environmental state of the building area, and performing non-periodic environmental state derivation through the derivation model of abnormal environmental state of the building area, thereby obtaining non-periodic environmental state data;
[0025] Step S136: Time-series merge the periodic environmental status data and the non-periodic environmental status data to obtain building area environmental status data.
[0026] The present invention constructs a comprehensive derivation method for normal and abnormal environmental states through in-depth analysis and modeling of building area sensor data, which can significantly improve the adaptability and perception ability of the monitoring system to the dynamic environment. The normalization processing of sensor data can effectively eliminate the interference caused by differences in sensor type, measurement unit and data range, and provide a data basis with higher consistency and reliability for subsequent analysis. The normalized time series sensor data is further classified by long-term and short-term memory time series features to achieve efficient separation of periodic and non-periodic data, so that more accurate analysis strategies can be formulated for data with different characteristics. The Bayesian network modeling of periodic sensor data provides a tool for probabilistic reasoning for the derivation of normal environmental states, and can generate a reliable normal environmental state derivation model in combination with the periodic characteristics of the time dimension. This model can quickly identify conventional behavior patterns in building areas and significantly improve the system's judgment efficiency on normal operating states. On the other hand, non-periodic sensor data is reduced in dimension through principal component analysis, which removes redundant information in the data, reduces computational complexity, and retains the main features that have a decisive influence on environmental state changes. On this basis, the abnormal environmental state derivation model is constructed through joint distribution modeling, which can effectively respond to emergencies and abnormal phenomena in the building area, quickly identify potential risks and provide timely feedback. The time series merging of periodic and non-periodic environmental state data can integrate the characteristics of normal and abnormal environmental states, and provide complete and accurate environmental state information for the monitoring system. This method not only improves the system's global perception of complex environmental changes, but also dynamically adjusts the monitoring strategy, significantly improving the real-time and robustness of the monitoring system. The time series merging of periodic and non-periodic environmental state data can integrate the characteristics of normal and abnormal environmental states, and provide complete and accurate environmental state information for the monitoring system. This method not only improves the system's global perception of complex environmental changes, but also dynamically adjusts the monitoring strategy, significantly improving the real-time and robustness of the monitoring system.
[0027] Optionally, step S15 is specifically:
[0028] Step S151: aligning the building area environment state data and the continuous video frame image set, thereby obtaining the environment state aligned data and the continuous video frame aligned image set;
[0029] Step S152: extracting image visual features from the continuous video frame aligned image set to obtain video frame visual feature data, and integrating scene visual dynamic features based on the video frame visual feature data to obtain scene visual dynamic feature data;
[0030] Step S153: integrating the scene environment dynamic features according to the environment state alignment data, thereby obtaining the scene environment dynamic feature data;
[0031] Step S154: performing scene dynamic feature fusion based on the scene visual dynamic feature data and the scene environment dynamic feature data, thereby obtaining a dynamic environment perception map of the building area.
[0032] The present invention comprehensively improves the monitoring system's perception of dynamic environments through data alignment, feature extraction and fusion methods, and solves the shortcomings of traditional monitoring in multimodal data processing and dynamic scene analysis. By aligning the data of environmental state data and video frame image sets, the interference caused by acquisition time differences and format inconsistencies between data sources is eliminated, providing a highly consistent data basis for subsequent analysis. This alignment method not only ensures synchronization in the spatiotemporal dimension, but also effectively enhances the correlation between environmental data and video data, thereby laying a solid foundation for multimodal fusion. Visual feature extraction of continuous video frames can efficiently capture key information in the image, such as edges, textures, and motion trajectories, and extract visual feature data that reflects changes in video content. This feature extraction provides an important input basis for subsequent scene dynamic feature analysis, and through the integration of scene visual dynamic features, comprehensively presents the dynamic changes of visual scenes in building areas. This method significantly improves the depth and breadth of video analysis, enabling the monitoring system to more accurately identify key visual information and support rapid response and decision-making of events. Based on the environmental state alignment data, through the integration of scene environment dynamic features, the system can comprehensively analyze the dynamic changes of environmental data, such as fluctuations in temperature, humidity, and light, so as to comprehensively depict the spatiotemporal evolution characteristics of environmental states. This feature integration not only improves the precision of environmental perception, but also provides the ability to predict the potential impact of environmental changes. Ultimately, by deeply integrating visual dynamic features with environmental dynamic features, the generated dynamic environmental perception map effectively makes up for the shortcomings of a single data source in dynamic perception, and provides the monitoring system with multi-dimensional and all-round dynamic perception capabilities.
[0033] Optionally, step S2 specifically includes:
[0034] Step S21: extracting the environment state and scene dynamics according to the dynamic environment perception map of the building area, thereby obtaining the perception map environment state data and scene dynamics data;
[0035] Step S22: performing environmental change statistics based on the perception map environmental state data to obtain regional environmental change data, and performing image quality impact factor calculation on the regional environmental change data and the continuous video frame aligned image set to obtain an image quality impact factor set;
[0036] Step S23: estimating the scene image enhancement strength according to the scene dynamic data and the image quality influencing factor set, thereby obtaining a scene image enhancement strength template;
[0037] Step S24: performing image scene matching on the scene image enhancement strength template and the continuous video frame aligned image set, thereby obtaining an image set enhancement strength strategy;
[0038] Step S25: performing image adaptive dynamic enhancement on the continuous video frame aligned image set according to the image set enhancement strength strategy, so as to obtain the building area enhanced video frame image set.
[0039] The present invention effectively solves the problem of image quality degradation in traditional video surveillance systems in complex dynamic scenes through dynamic environment perception and adaptive image enhancement technology. Environmental state extraction and scene dynamic extraction can accurately capture the core features of environmental conditions and scene changes in building areas, and provide real-time and dynamic background information support for image processing. This feature significantly improves the system's ability to cope with complex environmental changes, especially under conditions where light, weather or obstructions change greatly, and can accurately reflect the real situation of the building area. Through environmental change statistics and calculation of image quality influencing factors, the system can quantify the impact of environmental changes on image quality, providing a scientific basis for subsequent image enhancement. This process deeply integrates environmental data with image processing, which not only improves the accuracy of the enhancement strategy, but also ensures the pertinence and effectiveness of image quality adjustment. Enhancement strength estimation is performed based on scene dynamic data and image quality influencing factor sets, and the enhancement strategy is further optimized, so that it can be flexibly adjusted according to the needs of different scenes, avoiding the problem of over-processing or under-processing that may be caused by traditional enhancement methods. The implementation of image scene matching enables the enhancement strength strategy to be accurately applied to specific scenes, ensuring that each frame of video image has the best clarity and detail performance after processing. This targeted processing method significantly improves the visual effect of video surveillance, especially in capturing details in key scenes and identifying important targets. Through adaptive dynamic enhancement technology, continuous video frames can maintain high-quality output in different dynamic environments, solving the problem of traditional video quality degradation under conditions such as low light, high noise or fast movement, and providing important support for the real-time and reliability of the monitoring system.
[0040] Optionally, step S3 specifically includes:
[0041] Step S31: performing visual feature statistics on the building area enhanced video frame image set, thereby obtaining high-frequency visual feature data and low-frequency visual feature data;
[0042] Step S32: performing scene pattern recognition based on the low-frequency visual feature data to obtain normal scene pattern data;
[0043] Step S33: performing video stream motion pattern capture on the high-frequency visual feature data to obtain video stream motion pattern data, and performing motion pattern classification on the video stream motion pattern data to obtain artificial motion pattern data and natural motion pattern data;
[0044] Step S34: performing pattern temporal association according to the human action pattern data and the normal scene pattern data, thereby obtaining scene-human action interaction event data, and performing low-frequency interaction event recognition on the scene-human action event pattern data, thereby obtaining abnormal human interaction event data;
[0045] Step S35: performing pattern temporal association on the natural action pattern data and the normal scene pattern data, thereby obtaining scene-natural action interaction event data, and performing continuous interaction event recognition on the scene-natural action interaction event data, thereby obtaining abnormal natural interaction event data;
[0046] Step S36: performing time-series merging on the abnormal human interaction event data and the abnormal natural interaction event data to obtain abnormal event time-series data, and performing abnormal event time-series video frame extraction on the building area enhanced video frame image set according to the abnormal event time-series data to obtain an abnormal event video frame image set;
[0047] Step S37: performing risk event identification on the abnormal event video frame image set, thereby obtaining a risk event video frame image set.
[0048] The present invention effectively improves the abnormal event detection capability and risk assessment accuracy of the monitoring system through multi-level video analysis and interactive event mining technology. The statistics and distinction of high-frequency and low-frequency visual features help to focus on both detailed features and overall changes, ensuring comprehensive capture and multi-angle analysis of scene dynamics. Through scene pattern recognition, the system can construct a normal scene pattern, provide a basic reference framework, and lay a stable standard for abnormal detection and event analysis. At the same time, video stream action pattern capture and classification achieve accurate separation of human actions and natural actions, effectively avoid noise interference, and improve the recognition ability of specific behaviors. In interactive event analysis, by temporally correlating human action patterns with normal scene patterns, the system can automatically identify potential abnormal human interaction events, such as actions that do not conform to conventional behavior patterns or actions that are not coordinated with the environment, thereby achieving early warning. Similarly, interactive event analysis of natural action patterns can quickly capture abnormal natural events, such as natural phenomena or physical changes that do not conform to environmental characteristics, further enhancing the system's adaptability and dynamic response capabilities in complex scenes. The process of temporal merging integrates a variety of abnormal event data, constructs a unified abnormal event temporal information flow, and provides efficient data input for subsequent analysis. The extraction of abnormal event video frames and event risk assessment further deepen the intelligent capabilities of the monitoring system, which can quickly screen and locate key video frame images related to high-risk events, and provide intuitive risk basis for security management and decision-making. This processing method not only reduces the time cost of manual screening, but also greatly improves the accuracy and efficiency of event analysis, enabling the monitoring system to respond to potential threats in a timely manner and scientifically evaluate them, providing a solid technical guarantee for improving public safety and emergency response capabilities. The comprehensive application of these steps has achieved full process coverage from dynamic monitoring to abnormal identification to risk assessment, providing important support for the innovative development of the field of intelligent monitoring.
[0049] Optionally, step S4 is specifically:
[0050] Step S41: acquiring building area network transmission data, and performing data preprocessing on the building area network transmission data, thereby obtaining the building area network transmission data to be analyzed;
[0051] Step S42: performing transmission node connection frequency statistics on the network transmission data of the building area to be analyzed, thereby obtaining network transmission node connection frequency data, and performing node connection relationship identification based on the network transmission node connection frequency data, thereby obtaining network transmission node connection relationship data;
[0052] Step S43: Modeling the building area network topology structure according to the network transmission node connection relationship data, thereby obtaining a building area network topology structure model;
[0053] Step S44: evaluating the connection path network status according to the building area network topology model, thereby obtaining path network status data, and allocating connection path weights based on the path network status data, thereby obtaining connection path weight data;
[0054] Step S45: encoding the connection path network status according to the connection path weight data and the path network status data, thereby obtaining network status encoding data.
[0055] The present invention achieves a deep insight into the network connection characteristics and topological structure through multi-level analysis and modeling of the network transmission data in the building area, and greatly improves the efficiency of network management and optimization. Data preprocessing ensures the standardization and accuracy of the input data, laying a reliable foundation for subsequent analysis. In the node connection frequency statistics and connection relationship identification links, the system can quickly capture the interaction behavior between network nodes, build complete node connection relationship data, and clearly show the actual situation of network communication, thereby providing high-quality data support for topological modeling. The implementation of network topological structure modeling successfully constructs the overall picture of the network architecture of the building area, allowing the system to have an intuitive understanding of node distribution, connection paths and overall network structure. This model provides a reliable framework for subsequent network status evaluation, and also provides an important basis for network optimization and fault location. In the network path status evaluation, based on the analysis of real-time path status data, the system can accurately reflect the actual performance of each connection path, such as delay, bandwidth and stability. Combined with path weight allocation, the system further optimizes the resource allocation strategy, improves the communication efficiency of the key path through the reasonable allocation of weights, and reduces the load risk of the overall network. The introduction of network status coding converts complex network status information into coded data that is easy to understand and process, providing a powerful tool for the intelligent management of monitoring systems. This coding method can not only quickly transmit network health status, but also provide efficient data support for subsequent analysis, prediction and early warning. Overall, these steps, from data acquisition to modeling, evaluation and coding, constitute a complete network analysis and optimization process, which significantly improves the monitoring system's ability to manage network resources and respond to faults, and provides strong technical support for intelligent monitoring and dynamic network optimization.
[0056] Optionally, step S5 specifically includes:
[0057] Step S51: extracting visual features of the risk event video frame image set, thereby obtaining visual feature data of the video frame image set;
[0058] Step S52: performing spatiotemporal correlation modeling based on the visual feature data of the video frame image set, thereby obtaining a spatiotemporal evolution model of the visual feature;
[0059] Step S53: compressing the high-dimensional features of the visual feature spatiotemporal evolution model to obtain a low-dimensional visual feature vector;
[0060] Step S54: encoding the video visual features according to the low-dimensional visual feature vector, thereby obtaining video feature encoding data;
[0061] Step S55: Adaptively compress the video according to the network status encoding data and the video feature encoding data, thereby obtaining building area event video compression data.
[0062] The present invention realizes efficient video processing and compression through multi-level analysis and optimization of risk event video frame image sets, and provides technical support for intelligent monitoring systems. The image visual feature extraction link can capture key information in video frames and extract high-quality visual feature data, laying the foundation for subsequent modeling. In the spatiotemporal correlation modeling, a spatiotemporal evolution model is generated by analyzing the temporal and spatial correlation of visual features. This model can not only reflect the changing trend of dynamic events, but also provide a panoramic perspective for event identification and analysis. High-dimensional feature compression significantly reduces the complexity of data processing and improves the system's computing efficiency by mapping complex high-dimensional data into low-dimensional vectors, while retaining the key feature information of the event. The introduction of video visual feature coding converts low-dimensional feature vectors into standardized coded data, which facilitates the system to efficiently process data during data transmission and storage, and provides structured input for subsequent video compression. In the adaptive video compression stage, combined with network status coded data and video feature coded data, the system can dynamically adjust the compression strategy according to the current network transmission conditions, maximize the balance between video quality and transmission efficiency, and ensure that event videos can be transmitted quickly and efficiently in a low-bandwidth environment.
[0063] Optionally, step S55 is specifically:
[0064] Step S551: calculating the video frame information volume of the risk event video frame image set, thereby obtaining risk event video frame information volume data, and dividing the risk key frames according to the risk event video frame information volume data, thereby obtaining risk key frame division data;
[0065] Step S552: selecting the ratio of key frames to non-key frames for the video feature coding data based on the risk key frame classification data, thereby obtaining a video frame ratio strategy;
[0066] Step S553: analyzing the compression strength dynamic adjustment strategy according to the network status coding data and the visual feature spatiotemporal evolution model, thereby obtaining the compression strength dynamic adjustment strategy;
[0067] Step S554: integrating the adaptive video compression strategy based on the video frame ratio strategy and the compression strength dynamic adjustment strategy, thereby obtaining an adaptive video compression strategy;
[0068] Step S555: Adaptively compress the video feature encoding data according to the adaptive video compression strategy and the network status encoding data, thereby obtaining building area event video compression data.
[0069] The present invention realizes an efficient adaptive video compression strategy through accurate processing and optimization of risk event video frames, and improves the transmission and storage efficiency of event monitoring videos under different network environments. First, the calculation of the amount of video frame information can evaluate the key information of each frame of video, so as to identify the most representative and important video frames through risk key frame division, which provides a basis for subsequent video compression and ensures that important information is not lost during the compression process. Based on the selection of the use ratio of key frames and non-key frames, the storage and transmission efficiency of video frames can be optimized by reasonably allocating the ratio of key frames and non-key frames. In the strategy analysis of dynamically adjusting the compression strength, the network status coding data and the spatiotemporal evolution model of visual features are combined, and the compression strength can be dynamically adjusted according to the current network conditions and changes in video content, so as to effectively balance the compression effect and video quality, and ensure the clarity and transmission stability of the video under different network environments. By integrating the video frame ratio strategy with the compression strength dynamic adjustment strategy, an adaptive video compression strategy is finally formed, which can flexibly respond to different network bandwidth conditions and video content changes, optimize the compression process in real time, and improve the intelligent level of video transmission. Video compression is performed based on the adaptive video compression strategy and network condition encoding data, which not only reduces the size of the video file, but also ensures the smoothness and quality of the video during transmission, making it possible to efficiently transmit video data of building area events in a low-bandwidth environment.
[0070] Optionally, the present specification further provides a machine vision intelligent recognition system based on digital buildings, which is used to execute the machine vision intelligent recognition method based on digital buildings as described above. The machine vision intelligent recognition system based on digital buildings includes:
[0071] The dynamic environment perception fusion module is used to obtain sensor data and camera data of the building area, and perform building area environmental status analysis based on the sensor data of the building area, so as to obtain building area environmental status data; perform dynamic environment perception fusion based on the building area environmental status data and the camera data of the building area, so as to obtain a dynamic environment perception map of the building area;
[0072] An image adaptive dynamic enhancement module is used to perform image adaptive dynamic enhancement on the camera data of the building area according to the dynamic environment perception map of the building area, so as to obtain an enhanced video frame image set of the building area;
[0073] An abnormal event detection module is used to perform abnormal event detection on the building area enhanced video frame image set to obtain an abnormal event video frame image set, and perform risk event identification on the abnormal event video frame image set to obtain a risk event video frame image set;
[0074] A network status encoding module is used to obtain building area network transmission data, and perform building area network topology structure analysis on the building area network transmission data, so as to obtain a building area network topology structure model; perform network status encoding based on the building area network topology structure model, so as to obtain network status encoding data;
[0075] A video compression module is used to perform video feature encoding data on the risk event video frame image set, thereby obtaining video feature encoding data, and to perform adaptive video compression according to the network status encoding data and the video feature encoding data, thereby obtaining building area event video compression data;
[0076] The data transmission resource allocation module is used to allocate data transmission resources for the compressed video data of the building area event through the building area network topology model, so as to obtain the building area compressed video transmission strategy and upload it to the building area risk management platform to perform the building area risk event warning task.
[0077] The digital building-based machine vision intelligent recognition system of the present invention can implement any digital building-based machine vision intelligent recognition method of the present invention, and is used to combine the operation and signal transmission medium between various modules to complete the digital building-based machine vision intelligent recognition method. The internal modules of the system cooperate with each other, thereby improving the risk identification efficiency and response speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments thereof made with reference to the following drawings:
[0079] Figure 1 It is a schematic diagram of the steps of the machine vision intelligent recognition method based on digital buildings of the present invention;
[0080] Figure 2 Detailed step flow diagram of step S1 in the present invention;
[0081] Figure 3 Detailed step flow diagram of step S2 in the present invention;
[0082] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0083] The technical method of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.
[0084] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0085] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0086] To achieve this, please refer to Figures 1 to 3 The present invention provides a machine vision intelligent recognition method based on digital buildings, the method comprising the following steps:
[0087] Step S1: acquiring sensor data and camera data of the building area, and performing building area environmental status analysis based on the sensor data of the building area, thereby obtaining building area environmental status data; performing dynamic environmental perception fusion based on the building area environmental status data and the camera data of the building area, thereby obtaining a dynamic environmental perception map of the building area;
[0088] In this embodiment, environmental status data is obtained through various sensors (such as temperature, humidity, motion, smoke, etc.) arranged in the building area. And combined with the video data collected by the camera, the environmental data is used to analyze the real-time status of the building area. For example, by analyzing the temperature and humidity sensor data, it is determined whether there is a fire hazard or abnormal temperature change. Then, through dynamic perception fusion technology, combined with the camera data for processing, a dynamic environmental perception map of the building area is generated. This map integrates multi-source data, can reflect environmental changes in real time, and provide accurate background information for subsequent video enhancement and abnormal event detection.
[0089] Step S2: performing image adaptive dynamic enhancement on the camera data of the building area according to the dynamic environment perception map of the building area, so as to obtain an enhanced video frame image set of the building area;
[0090] In this embodiment, based on the dynamic environment perception map of the building area, the camera data is adaptively and dynamically enhanced. Specifically, the image brightness, contrast, clarity and other parameters are automatically adjusted according to different environmental conditions (such as lighting changes, smoke or haze, etc.). For example, in a low-light environment, the visibility of the video frame is improved through the brightness enhancement algorithm to ensure that even at night or in dim conditions, the key areas can still be clearly presented, avoiding monitoring blind spots caused by insufficient ambient light.
[0091] Step S3: performing abnormal event detection on the building area enhanced video frame image set to obtain an abnormal event video frame image set, and performing risk event identification on the abnormal event video frame image set to obtain a risk event video frame image set;
[0092] In this embodiment, abnormal event detection is performed on the enhanced video frame image set. A deep learning target detection model (such as YOLO or Faster R-CNN) is used to identify abnormal events in the video, such as intrusion, falling objects, or fire. Risk assessment is performed based on the recognition results, and the degree of danger of the detected events is assessed according to the preset risk analysis model, and a video frame image set of high-risk events is generated. This process ensures that potential emergencies can be detected quickly and provides effective support for subsequent responses.
[0093] Step S4: acquiring building area network transmission data, and performing building area network topology analysis on the building area network transmission data, thereby obtaining a building area network topology model; performing network status encoding based on the building area network topology model, thereby obtaining network status encoding data;
[0094] In this embodiment, the network transmission data in the building area is obtained through the management cloud platform of the building area, and the topological structure analysis of the network nodes is performed. Specifically, by analyzing the network transmission data (such as bandwidth utilization, transmission delay, packet loss rate, etc.), a topological structure model of the regional network is constructed. According to the model, the connection relationship between each node can be identified, and it is clear which nodes have the most critical transmission and which nodes have potential bottlenecks. Subsequently, the network status is encoded according to the topological structure data to generate network status encoding data, which provides a reference for subsequent video compression and data transmission optimization.
[0095] Step S5: performing video feature encoding data on the risk event video frame image set to obtain video feature encoding data, and performing adaptive video compression according to the network status encoding data and the video feature encoding data to obtain building area event video compression data;
[0096] In this embodiment, video feature encoding is performed on the risk event video frame image set. By adopting a video compression algorithm (such as H.264 or HEVC), the event video data is compressed into a low-dimensional feature representation to reduce the bandwidth requirement for data transmission. During the compression process, the compression strategy is dynamically adjusted according to the network status of the encoded data. For example, if the network bandwidth is sufficient, the video quality can be maintained at a high standard; conversely, when the network bandwidth is limited, the compression ratio is automatically increased to reduce transmission delays and bandwidth occupancy, ensuring smooth video transmission.
[0097] Step S6: Allocate data transmission resources for the compressed video data of the building area event through the building area network topology model, so as to obtain the building area compressed video transmission strategy, and upload it to the building area risk management platform to execute the building area risk event warning task.
[0098] In this embodiment, the building area network topology model is used to allocate data transmission resources for the compressed video data. Based on the current network conditions (such as bandwidth, node load, etc.), the most suitable transmission path is intelligently selected. For example, if some nodes are under high load, the path with lower load will be preferentially selected for video transmission. At the same time, the transmission strategy is adjusted according to the real-time bandwidth situation when necessary to ensure that the video data can be transmitted stably and efficiently when the network load changes. Finally, the video data is uploaded to the monitoring management platform for real-time viewing and early warning by management personnel.
[0099] Optionally, step S1 specifically includes:
[0100] Step S11: acquiring building area sensor data and building area camera data, and performing data preprocessing on the building area sensor data and the building area camera data respectively, so as to obtain the building area sensor data to be analyzed and the building area camera data to be analyzed;
[0101] In this embodiment, environmental status data is obtained through various sensors (such as temperature, humidity, motion, smoke, etc.) arranged in the building area; and video data of the building area is obtained through cameras deployed in the building area. After the sensor data and video data of the building area are obtained, data preprocessing is first performed. Sensor data preprocessing includes denoising, calibration, and outlier detection to ensure that the data collected by the sensor is accurate and valid. For example, for a temperature sensor, it is necessary to remove erroneous readings affected by environmental interference and convert them into a unified unit. Video data preprocessing includes frame rate adjustment, image denoising, and enhancement processing to ensure the accuracy of subsequent analysis. For example, image denoising technology can reduce picture blur caused by camera noise or environmental factors, thereby ensuring the clarity of the video. The preprocessed sensor data and video data will be saved as data to be analyzed, respectively, for use in subsequent steps.
[0102] Step S12: performing sensor data time series analysis on the sensor data of the building area to be analyzed, thereby obtaining time series sensor data of the building area;
[0103] In this embodiment, for the sensor data of the building area to be analyzed, time series analysis technology is used, and methods such as ARIMA model or LSTM network are used for analysis to identify the periodic changes and trends of the data. For example, for gas sensors, time series analysis can help detect the gradual change or sudden drastic fluctuation of gas concentration, and further determine whether there is a risk of a dangerous source. This analysis helps to track the environmental changes in the building area in real time and identify potential risks in advance.
[0104] Step S13: deriving the building area environment state according to the building area time series sensing data, thereby obtaining the building area environment state data;
[0105] In this embodiment, environmental status data is further derived based on the building area sensor data obtained by time series analysis. For example, by analyzing data such as temperature, humidity, and gas concentration, it is determined whether there is a danger such as fire or toxic gas leakage. Assuming that in the building area sensor, the temperature and gas concentration are abnormally increased at the same time, such environmental status data will indicate that there may be a fire risk, and the system will generate an alarm accordingly to promote further prevention or response measures.
[0106] Step S14: decoding the camera video stream of the camera data of the building area to be analyzed, thereby obtaining the camera video of the building area, and dividing the camera video of the building area into time window frames, thereby obtaining a continuous video frame image set;
[0107] In this embodiment, for the camera data to be analyzed, the video stream is decoded, the video file is decompressed and the frame images therein are extracted. These frame images are segmented by time window technology, that is, the video stream is divided into multiple time periods according to time, ensuring that the images in each time period can focus on reflecting the specific situation of the time period. For example, the video stream of the building area is divided into a time window of 30 seconds, which can ensure that sufficient visual information is captured during image processing to avoid missing important events.
[0108] Step S15: Perform multimodal dynamic environment perception fusion on the building area environmental status data and the continuous video frame image set, so as to obtain a dynamic environment perception map of the building area.
[0109] In this embodiment, the environmental status data of the building area is fused with a continuous set of video frame images to generate a dynamic environmental perception map. Through multimodal data fusion technology, the environmental information (such as temperature, humidity, gas concentration, etc.) provided by the sensor is combined with the video data to display the environmental changes in the building area in real time. For example, when there are signs of fire, the sensor shows an abnormal increase in temperature, and the video frame may show flames or thick smoke. The dynamic environmental perception map obtained after fusion will be able to accurately reflect the environmental changes in the area and help monitoring personnel make quick response decisions.
[0110] Optionally, step S13 is specifically:
[0111] Step S131: normalizing the time series sensing data of the building area to obtain normalized time series sensing data;
[0112] In this embodiment, the time-series sensor data of the building area will be normalized to ensure the comparability of various sensor data and eliminate the influence of different dimensions. Specifically, for sensor data such as temperature, humidity, gas concentration, etc., their numerical ranges may vary greatly, so these data are normalized to between 0 and 1, using the Min-Max normalization method to make subsequent analysis more accurate. This process helps to eliminate data inconsistency problems caused by different sensor types and ensure that each type of data can play an equal role in the analysis process.
[0113] Step S132: performing long-short term memory time series feature classification on the normalized time series sensor data, thereby obtaining time series periodic sensor data and time series non-periodic sensor data;
[0114] In this embodiment, the normalized time series sensor data will be classified by long short-term memory network (LSTM) for time series feature classification. LSTM can effectively capture the long-term dependencies in time series data and identify the periodic characteristics of sensor data by training the network. For example, the data of the temperature sensor shows an obvious day and night variation pattern, which is periodic data; while the gas concentration data is affected by emergencies and does not have a fixed periodicity, which is non-periodic data. Through the analysis of LSTM, these data can be divided into time series periodic sensor data and time series non-periodic sensor data, providing a clear classification basis for subsequent modeling.
[0115] Step S133: performing Bayesian network modeling of building area sensor data according to the time-series periodic sensor data, thereby obtaining a building area normal environment state derivation model, and performing periodic environment state derivation through the building area normal environment state derivation model, thereby obtaining periodic environment state data;
[0116] In this embodiment, a Bayesian network is used to model the normal environmental state of the building area for time-series periodic sensor data. The Bayesian network can derive the causal relationship of the environmental state by modeling the probabilistic relationship of the data. For example, by analyzing periodic data such as temperature and humidity, the Bayesian network can derive the relationship between these data and then establish a normal environmental state derivation model. This model can be used to derive periodic environmental states. When the temperature data shows regular changes, the corresponding environmental state data can be derived according to the model to help predict the regular environmental changes in the area.
[0117] Step S134: performing principal component analysis on the time series non-periodic sensor data to obtain non-periodic reduced dimension sensor data;
[0118] In this embodiment, the time-series non-periodic sensor data will be processed by principal component analysis (PCA) to reduce the dimension to extract the main features of the data. PCA technology converts multidimensional data into lower-dimensional data through linear transformation, thereby retaining the maximum variance of the data. For example, for gas concentration and pollutant sensor data, PCA can remove redundant features, retain key components related to events, reduce the complexity of data processing, and make subsequent modeling more efficient. The reduced-dimensional data can be more easily used for model training and analysis.
[0119] Step S135: performing joint distribution modeling of building area sensor data on the non-periodic reduced dimension sensor data, thereby obtaining a derivation model of abnormal environmental state of the building area, and performing non-periodic environmental state derivation through the derivation model of abnormal environmental state of the building area, thereby obtaining non-periodic environmental state data;
[0120] In this embodiment, the sensor data after non-periodic dimensionality reduction will be used for joint distribution modeling of sensor data in the building area. Through joint distribution modeling, the correlation and dependency between different types of sensor data can be captured, so as to obtain an abnormal environmental state derivation model. After using principal component analysis (PCA) for dimensionality reduction, the reduced dimensionality data obtained will be processed by the joint distribution model. The joint distribution model can establish the correlation between different types of sensor data through statistical methods. For example, a sudden increase in gas concentration may be accompanied by fluctuations in temperature or humidity. In this process, a Gaussian mixture model (GMM) or other related statistical modeling methods are used to capture the relationship between each data feature, and a model is established to derive the environmental state by fitting the joint probability distribution of these data. For non-periodic sensor data such as gas concentration and temperature changes, when the gas concentration increases abnormally, the joint distribution model can analyze its relationship with temperature changes in combination with historical data, and then derive possible environmental abnormal states, such as safety risks such as fire or leakage. This model helps to further improve the accuracy of environmental state derivation, making environmental monitoring more accurate and sensitive, and able to feedback potential safety hazards in real time.
[0121] Step S136: Time-series merge the periodic environmental status data and the non-periodic environmental status data to obtain building area environmental status data.
[0122] In this embodiment, the periodic environmental status data and the non-periodic environmental status data will be merged in time series to form complete building area environmental status data. This merging process can integrate periodic and non-periodic data to form a more comprehensive environmental status assessment. For example, in the merging process, periodic data (such as temperature) and non-periodic data (such as sudden changes in gas concentration) can be time-aligned according to the characteristics of the time series, and finally a complete set of environmental status data is obtained, which provides accurate data support for subsequent early warning, monitoring decisions and response measures.
[0123] Optionally, step S15 is specifically:
[0124] Step S151: performing data alignment on the building area environment state data and the continuous video frame image set, thereby obtaining the environment state aligned data and the continuous video frame aligned image set;
[0125] In this embodiment, the environmental state data and the continuous video frame image set need to be accurately aligned in time and space. This is usually achieved through timestamps or event markers. According to the sampling frequency of the sensor and the camera, a suitable time window is selected to match the environmental state data with the video frame image set. For example, if the sampling frequency of the sensor data is once per second, and the camera collects 30 frames of images per second, the environmental data can be aligned with the corresponding video frames through interpolation or sampling synchronization, thereby ensuring that the data of the two can correspond correctly, and obtaining environmental state alignment data and continuous video frame alignment image sets.
[0126] Step S152: extracting image visual features from the continuous video frame aligned image set to obtain video frame visual feature data, and integrating scene visual dynamic features based on the video frame visual feature data to obtain scene visual dynamic feature data;
[0127] In this embodiment, by extracting the visual features of the video frame images, a convolutional neural network (CNN) can be used to extract the visual information of each frame image, such as features such as edges, colors, textures, and shapes. Based on these feature data, the visual dynamic features of the scene are further integrated. Specifically, a deep learning model such as ResNet or VGG network is first used for feature extraction, and then the visual features of continuous video frames are integrated through timing analysis or self-attention mechanism to form a comprehensive scene visual dynamic feature dataset, which can capture the dynamic features of scene changes in the video, such as pedestrians walking, object movement, etc.
[0128] Step S153: integrating the scene environment dynamic features according to the environment state alignment data, thereby obtaining the scene environment dynamic feature data;
[0129] In the present embodiment, data smoothing techniques (such as moving average or exponential weighted smoothing) can be used to remove noise and maintain the continuity and stability of the data. Subsequently, the processed environmental state data are aligned with the video frame data in the time window to ensure that the data of different sensors match the video frames of the corresponding time period one by one. For example, if the sensor provides environmental data every 5 minutes, and the video stream collects one frame of image per second, then the environmental data needs to be interpolated so that it can be consistent with the video frame time point. In this way, the environmental dynamic features related to the video content can be accurately extracted, such as when abnormal behavior occurs in a certain area in the video, the environmental changes (such as sudden temperature changes, increased humidity, etc.) of the time period are obtained in combination with the sensor data, and then the scene environment dynamic feature data is obtained.
[0130] Step S154: performing scene dynamic feature fusion based on the scene visual dynamic feature data and the scene environment dynamic feature data, thereby obtaining a dynamic environment perception map of the building area.
[0131] In this embodiment, deep learning methods such as convolutional neural networks (CNN) and recurrent neural networks (RNN) are used to extract visual dynamic features from video frames, such as the movement trajectory of characters, the appearance and disappearance of objects, and other dynamic information. At the same time, data from environmental sensors (such as temperature and humidity changes) are used for feature extraction. Next, feature fusion technology is used to combine these information from visual and environmental data. For example, a weighted average method can be used to assign different weights to visual features and environmental features according to their respective importance, or a long short-term memory (LSTM) network can be used to capture the temporal relationship between the two, thereby forming comprehensive dynamic feature data of the scene. The result of this fusion is a multimodal dynamic perception map that can simultaneously consider environmental changes and visual content, and enhance the perception of complex events in the building area. For example, when a person is detected entering an area in the video and the ambient temperature suddenly rises, the fused dynamic environmental perception map can mark the abnormal risks that may exist in the area and assist in subsequent early warning and analysis.
[0132] Optionally, step S2 specifically includes:
[0133] Step S21: extracting the environment state and scene dynamics according to the dynamic environment perception map of the building area, thereby obtaining the perception map environment state data and scene dynamics data;
[0134] In this embodiment, the environmental state and scene dynamics are extracted by analyzing the dynamic environmental perception map of the building area. Specifically, the multimodal data in the perception map, such as temperature and humidity sensors, noise sensors, and visual information of video frames, are used to extract environmental state data, such as temperature changes and humidity fluctuations in the current area. At the same time, continuous video frame analysis is used to extract scene dynamic data, such as object movement trajectories, the flow of people or vehicles in the scene, etc. Through image processing technology (such as background subtraction and optical flow method), combined with environmental sensor data, activity patterns and environmental changes in the building area can be accurately identified, providing a basis for subsequent data analysis.
[0135] Step S22: performing environmental change statistics based on the perception map environmental state data to obtain regional environmental change data, and performing image quality impact factor calculation on the regional environmental change data and the continuous video frame aligned image set to obtain an image quality impact factor set;
[0136] In this embodiment, environmental change statistics are performed based on the perception map environmental state data, and statistical analysis is performed on the environmental data. For example, the trend of environmental changes is evaluated by calculating statistics such as the standard deviation and mean of the environmental data. Regional environmental change data may include temperature change rate, humidity fluctuation range, changes in air pressure, etc. Next, these change data are aligned with continuous video frame images, and the impact of different environmental changes on image quality is evaluated through an image quality analysis algorithm (for example, by calculating image contrast, noise ratio, image clarity, etc.). The result will generate a set of image quality influencing factors, which can help evaluate the specific impact of environmental changes on video quality and provide a basis for image enhancement.
[0137] Step S23: estimating the scene image enhancement strength according to the scene dynamic data and the image quality influencing factor set, thereby obtaining a scene image enhancement strength template;
[0138] In this embodiment, the scene image enhancement intensity is estimated based on the scene dynamic data and the image quality influencing factor set. By using the motion information in the scene dynamic data (such as the movement of objects in the scene, the behavior patterns of people, etc.) and the environmental change data, a mathematical model (such as a regression model or a neural network) is established to calculate the image enhancement intensity required under different environmental conditions. For example, in a low-light environment, the brightness enhancement intensity is increased; in a dynamic scene, the clarity and contrast of the image are enhanced. Finally, based on these estimation results, a scene image enhancement intensity template is formed to indicate what intensity of image enhancement should be performed in which scenes and under what environmental conditions.
[0139] Step S24: performing image scene matching on the scene image enhancement strength template and the continuous video frame aligned image set, thereby obtaining an image set enhancement strength strategy;
[0140] In this embodiment, the scene image enhancement intensity template is matched with the continuous video frame alignment image set to perform image scene matching, and the video frame is matched with the enhancement intensity template through image processing algorithms (such as image segmentation, feature point matching, etc.), and the image enhancement strategy is dynamically adjusted according to the different characteristics of the scene. For example, in a crowded scene, it is necessary to enhance contrast and edge sharpening, while in a static environment, more attention is paid to brightness and color enhancement. Through the matching process, it can be ensured that the enhancement intensity of each frame of the image is consistent with the current scene and environmental changes, ensuring the optimization of the image enhancement effect.
[0141] Step S25: performing image adaptive dynamic enhancement on the continuous video frame aligned image set according to the image set enhancement strength strategy, so as to obtain the building area enhanced video frame image set.
[0142] In this embodiment, according to the image set enhancement strength strategy, the continuous video frame image set is adaptively and dynamically enhanced. Adaptive algorithms (such as adaptive histogram equalization, local contrast enhancement, etc.) are used to enhance each frame of the image to ensure that the image can be most appropriately enhanced under different environmental conditions. For example, in low-light environments, increase exposure and brightness; in daytime conditions with strong sunlight, adjust contrast and color balance. In this way, an enhanced video frame image set with high visual clarity and more accurate information extraction can be provided for the building area, effectively improving the quality and efficiency of video surveillance.
[0143] Optionally, step S3 specifically includes:
[0144] Step S31: performing visual feature statistics on the building area enhanced video frame image set, thereby obtaining high-frequency visual feature data and low-frequency visual feature data;
[0145] In this embodiment, visual feature statistics are performed on the enhanced video frame image set of the building area. Specifically, a convolutional neural network (CNN) is used to extract features from the video frame images, and the image information is divided into high-frequency and low-frequency features. High-frequency visual features mainly include rapidly changing details, such as the edges and textures of moving objects, while low-frequency visual features represent slower changing parts, such as color changes and lighting of the scene background. High-frequency features can be extracted by methods such as Fourier transform, and low-frequency features are extracted by averaging or Gaussian filtering. In the feature statistics process, the changes of each pixel in the video frame image are processed to obtain a series of high-frequency and low-frequency feature data, which provide basic data for subsequent analysis.
[0146] Step S32: performing scene pattern recognition based on the low-frequency visual feature data to obtain normal scene pattern data;
[0147] In this embodiment, scene pattern recognition is performed based on low-frequency visual feature data. By analyzing the low-frequency data, a clustering algorithm or a time series model (such as a hidden Markov model) is used to perform pattern recognition on the scene. The purpose of scene pattern recognition is to distinguish common scenes (such as office areas, corridors, parking lots, etc.) from non-abnormal activities in building areas. Normal scene pattern data includes typical background changes, environmental stability, etc. At this time, based on factors such as the stability of the scene and time changes, the patterns that usually appear in the environment are identified as normal patterns, such as traffic flow, people entering and exiting, etc. This process helps to provide a benchmark for subsequent anomaly detection.
[0148] Step S33: Capture the video streaming action mode on the high-frequency visual feature data to obtain the video streaming action mode data, and classify the video streaming action mode data to obtain the human action mode data and the natural action mode data;
[0149] In this embodiment, the high-frequency visual feature data is used to capture the video stream motion pattern, and the optical flow method, feature matching and target detection methods are used to identify the rapidly changing parts in the video frame and capture the motion pattern in the video. The video stream motion pattern data contains the motion trajectory and changes of dynamic objects, such as the movement of people and the movement of vehicles. Furthermore, through the motion pattern classification algorithm, such as support vector machine (SVM) or neural network, the actions in the video are divided into artificial motion patterns and natural motion patterns. Artificial motion patterns include specific behaviors of people, such as walking, running, fighting, etc.; natural motion patterns include natural changes in the environment, such as wind blowing, trees swaying, etc. This classification process can effectively distinguish dynamic changes caused by human factors and natural factors.
[0150] Step S34: performing pattern temporal association according to the human action pattern data and the normal scene pattern data, thereby obtaining scene-human action interaction event data, and performing low-frequency interaction event recognition on the scene-human action event pattern data, thereby obtaining abnormal human interaction event data;
[0151] In this embodiment, the time series association of patterns is performed based on the human action pattern data and the normal scene pattern data, and the relationship between the two in the time series is analyzed. The human action pattern is associated with the scene pattern through a time series pattern matching algorithm (such as dynamic time warping DTW), and the scene-human action interaction event data is extracted. If a person performs abnormal behavior (such as fast running, pausing, etc.) in a building area, and the behavior appears in a specific normal scene, it may constitute a potential safety risk. Low-frequency interaction event recognition is further performed on these event pattern data, and by analyzing the long-term changes in the low-frequency feature data, it is identified whether there is abnormal behavior, and finally the abnormal human interaction event data is obtained.
[0152] Step S35: performing pattern temporal association on the natural action pattern data and the normal scene pattern data, thereby obtaining scene-natural action interaction event data, and performing continuous interaction event recognition on the scene-natural action interaction event data, thereby obtaining abnormal natural interaction event data;
[0153] In this embodiment, the natural action pattern data and the normal scene pattern data are temporally associated to analyze the relationship between the natural action and the scene. By temporally associating the scene and the natural action, the scene-natural action interaction event data is identified. The natural action pattern generally refers to unconscious or naturally occurring phenomena, such as the free fall of an object or the shaking of a tree. When these natural actions interact with the normal scene, accidents or safety hazards may occur. Through the continuous interaction event recognition algorithm, continuous interaction behaviors (such as fire, etc.) in the natural action mode, that is, abnormal behaviors that exist, can be identified, thereby obtaining abnormal natural interaction event data.
[0154] Step S36: performing time-series merging on the abnormal human interaction event data and the abnormal natural interaction event data to obtain abnormal event time-series data, and performing abnormal event time-series video frame extraction on the building area enhanced video frame image set according to the abnormal event time-series data to obtain an abnormal event video frame image set;
[0155] In this embodiment, the abnormal human interaction event data and the abnormal natural interaction event data are merged in time series, and the time and feature information of the abnormal events are combined to merge these data into a complete abnormal event time series data. By analyzing the time series, it is possible to identify which time periods have intersections and associations of multiple abnormal events. On this basis, the relevant abnormal event video frame image sets are further extracted from the building area enhanced video frame image set to ensure that all image information related to the abnormal events can be captured. These abnormal event video frame image sets contain those video frames that are related to abnormal behaviors in time and features, providing detailed image data support for subsequent risk assessment.
[0156] Step S37: performing risk event identification on the abnormal event video frame image set, thereby obtaining a risk event video frame image set.
[0157] In this embodiment, risk event identification is performed on the abnormal event video frame image set. Risk analysis is performed on each abnormal event in combination with historical data and event analysis. By matching abnormal behavior with environmental changes and background patterns, it is evaluated whether these abnormal events may lead to security risks or emergencies. By using deep learning models or traditional risk assessment algorithms, combined with the dynamic behavior displayed in the video frame images, each abnormal event is classified into a risk level. Finally, a risk event video frame image set is generated, which includes those event frames that are judged to have a higher risk during the assessment process. These video frame image sets provide security personnel or emergency response teams with timely risk information, helping them to respond quickly.
[0158] Optionally, step S4 is specifically:
[0159] Step S41: acquiring building area network transmission data, and performing data preprocessing on the building area network transmission data, thereby obtaining the building area network transmission data to be analyzed;
[0160] In this embodiment, the network transmission data in the building area is obtained by collecting the data exchange records, data packet flow, transmission delay and other information between each transmission node in the building area cloud platform. Then, these network transmission data are preprocessed, including removing redundant information, data cleaning and normalization processing. Specifically, the transmission data can be analyzed in time series to remove abnormal fluctuations and invalid data, so as to obtain the building area network transmission data to be analyzed, which provides a basis for further network analysis and modeling. During preprocessing, signal processing techniques such as smoothing or filtering algorithms can be used to ensure data quality and reduce noise interference.
[0161] Step S42: performing transmission node connection frequency statistics on the network transmission data of the building area to be analyzed, thereby obtaining network transmission node connection frequency data, and performing node connection relationship identification based on the network transmission node connection frequency data, thereby obtaining network transmission node connection relationship data;
[0162] In this embodiment, the transmission node connection frequency statistics are performed on the building area network transmission data to be analyzed. Specifically, the connection frequency between each transmission node (such as a router, switch, terminal device, etc.) and other nodes is counted. This can be done by counting the number of data packets sent or the communication duration between each node to obtain the network transmission node connection frequency data. Through further analysis of these data, the node connection relationship is identified. For example, a graph theory algorithm (such as the shortest path algorithm, K-means clustering, etc.) is used to identify which nodes frequently communicate, and then the node connection relationship data is constructed to represent the connection density and connection strength between each node, which helps to identify the key transmission nodes in the network.
[0163] Step S43: Modeling the building area network topology structure according to the network transmission node connection relationship data, thereby obtaining a building area network topology structure model;
[0164] In this embodiment, a topological structure model of the building area network is constructed based on the obtained network transmission node connection relationship data. This model graphically displays each network node and its connection relationship, reflecting the structural characteristics of the network. In specific implementation, an adjacency matrix or adjacency table can be used to represent the node connection of the network, and a complete topological structure model can be constructed through a network topology optimization algorithm (such as a minimum spanning tree, a Dijkstra algorithm, etc.). This process ensures that all nodes of the network and their connection paths are accurately represented, and reveals the overall architecture of the network, laying the foundation for subsequent path network evaluation and optimization.
[0165] Step S44: evaluating the connection path network status according to the building area network topology model, thereby obtaining path network status data, and allocating connection path weights based on the path network status data, thereby obtaining connection path weight data;
[0166] In this embodiment, the network status of the connection path is evaluated based on the building area network topology model. In this step, the performance indicators such as bandwidth, delay, packet loss rate, etc. of each connection path are evaluated to obtain the path network status data. In order to achieve this goal, network performance monitoring tools (such as ping command, traceroute, SNMP protocol, etc.) can be used to monitor each path in real time to obtain the transmission quality of each connection path. During the evaluation process, it is necessary to assign weights according to the transmission performance of each path, focusing on those key paths with large data flow or sensitive to delay. Finally, the obtained path network status data helps to understand the performance bottlenecks of each connection path and provide a basis for weight allocation. Based on these network status data, the connection paths are weighted, and the better paths have higher weights and the worse paths have lower weights. For example, bandwidth and delay are key influencing factors, and paths with high bandwidth and low delay will obtain higher weights.
[0167] Step S45: encoding the connection path network status according to the connection path weight data and the path network status data, thereby obtaining network status encoding data.
[0168] In this embodiment, the connection path network status is encoded based on the obtained connection path weight data and path network status data. Specifically, in combination with the network topology and path status, an encoding method (such as Huffman coding, difference coding, etc.) is used to convert the network status of each connection path into encoded data. This encoding process can effectively compress and digitize the performance information of the network, facilitating subsequent data transmission and processing. By encoding different path conditions, the transmission quality information in the network can be made more efficient during transmission or storage, providing efficient data support for network management. Ultimately, through this encoded data, the network status can be quickly judged, thereby providing a basis for subsequent network optimization and troubleshooting.
[0169] Optionally, step S5 specifically includes:
[0170] Step S51: extracting visual features of the risk event video frame image set, thereby obtaining visual feature data of the video frame image set;
[0171] In this embodiment, the image visual features are extracted from the risk event video frame image set, and each frame of the image is processed using a deep learning model such as a convolutional neural network (CNN). The CNN model performs convolution and pooling operations on the image to extract high-order visual features such as edges, textures, shapes, and colors from the image. For example, in a security monitoring scene, CNN can extract the contour features of people, objects, vehicles, etc. in the building area. These visual features can effectively describe the specific circumstances of the event.
[0172] Step S52: performing spatiotemporal correlation modeling based on the visual feature data of the video frame image set, thereby obtaining a spatiotemporal evolution model of the visual feature;
[0173] In this embodiment, spatiotemporal correlation modeling is performed based on the visual feature data of the video frame image set, a time series model such as a long short-term memory network (LSTM) is used to process the time series data in the video, and a spatial convolutional network (SCN) is combined to analyze the spatial relationship of the image to establish a spatiotemporal correlation model. In this way, the temporal changes between adjacent frames in the video frame and the mutual relationship between objects in space can be captured, thereby establishing a spatiotemporal evolution model of visual features. For example, for a dynamic monitoring scene, LSTM can identify the position changes of objects or people within a specific time period, and SCN can analyze the spatial relationship of these objects to form a complete spatiotemporal evolution model.
[0174] Step S53: compressing the high-dimensional features of the visual feature spatiotemporal evolution model to obtain a low-dimensional visual feature vector;
[0175] In this embodiment, principal component analysis (PCA) or autoencoder is used to compress the high-dimensional features of the spatiotemporal evolution model, compressing it from high-dimensional space to low-dimensional space, and retaining the key information in the data. For example, PCA removes redundant features and retains only the main change information, reducing the spatiotemporal features from hundreds of dimensions to dozens of dimensions, and generating low-dimensional visual feature vectors. This process can not only reduce the space for data storage, but also accelerate subsequent data processing and analysis.
[0176] Step S54: encoding the video visual features according to the low-dimensional visual feature vector, thereby obtaining video feature encoding data;
[0177] In this embodiment, the video visual feature encoding is performed based on the low-dimensional visual feature vector, and the feature vector is encoded using a quantization encoding method to convert it into a digital format for efficient storage. For example, by encoding the low-dimensional feature using Huffman coding, the feature data can be compressed in a binary format, so that the data occupies less space during transmission and storage, thereby improving data transmission efficiency.
[0178] Step S55: Adaptively compress the video according to the network status encoding data and the video feature encoding data, thereby obtaining building area event video compression data.
[0179] In this embodiment, adaptive video compression is performed based on the network status coded data and the video feature coded data, and the video compression parameters are adjusted in combination with factors such as network bandwidth and delay to ensure the stability and smoothness of the video data during transmission. For example, when the network bandwidth is low, the video compression ratio is automatically increased, and the video resolution or frame rate is reduced to reduce the amount of data transmission; when the bandwidth is wide, the video quality is improved, and a high image resolution and frame rate are maintained, thereby achieving dynamic adaptive video compression.
[0180] Optionally, step S55 is specifically:
[0181] Step S551: calculating the video frame information volume of the risk event video frame image set, thereby obtaining risk event video frame information volume data, and dividing the risk key frames according to the risk event video frame information volume data, thereby obtaining risk key frame division data;
[0182] In this embodiment, the information content of each frame is quantified by calculating the video frame information content of the risk event video frame image set, and the entropy-based calculation method is used to quantify the information content of each frame. Specifically, the entropy value is used to calculate the image complexity of each frame, and the video frames are divided into two types of frames with high information content and low information content according to the entropy value. For example, frames with high information content contain important motion information or significant changes, while frames with low information content are frames with stable or unchanged backgrounds. Based on these data, risk key frame division is performed, and frames containing high information content are marked as key frames, and frames with low information content are classified as non-key frames.
[0183] Step S552: selecting the ratio of key frames to non-key frames for the video feature coding data based on the risk key frame classification data, thereby obtaining a video frame ratio strategy;
[0184] In this embodiment, based on the risk key frame classification data, the usage ratio of key frames and non-key frames is selected. For example, in the video encoding process, a strategy with a higher ratio of key frames is selected to ensure that the main information of the video can be preserved with higher quality, while non-key frames can use lower quality or higher compression ratio. This selection process can be adjusted by setting a ratio threshold, and the specific ratio can be dynamically adjusted according to the complexity of the video content and the urgency of the event.
[0185] Step S553: analyzing the compression strength dynamic adjustment strategy according to the network status coding data and the visual feature spatiotemporal evolution model, thereby obtaining the compression strength dynamic adjustment strategy;
[0186] In this embodiment, the influence of factors such as network bandwidth and delay on the video compression strength is analyzed based on the network condition coding data and the spatiotemporal evolution model of visual features. For example, when the network condition is poor, a stronger compression strategy is adopted to reduce the transmission load, while when the network is relatively smooth, the compression strength can be reduced to retain more visual details. Further analysis is performed through the spatiotemporal evolution model to identify the timing characteristics of important scenes in the video, ensuring that the visual quality of key areas is not affected by compression.
[0187] Step S554: integrating the adaptive video compression strategy based on the video frame ratio strategy and the compression strength dynamic adjustment strategy, thereby obtaining an adaptive video compression strategy;
[0188] In this embodiment, an adaptive video compression strategy is integrated based on the video frame ratio strategy and the compression strength dynamic adjustment strategy. During the strategy integration process, the compression ratio and key frame interval of the video frame are adjusted by analyzing the key frame distribution of the video and the network status. For example, when the network status is poor, the number of key frames can be reduced and the compression rate of non-key frames can be increased. Conversely, the frequency and quality of key frames can be increased.
[0189] Step S555: Adaptively compress the video feature encoding data according to the adaptive video compression strategy and the network status encoding data, thereby obtaining building area event video compression data.
[0190] In this embodiment, adaptive video compression is performed based on the adaptive video compression strategy and the network status encoding data. According to the integrated compression strategy, the video feature encoding data is actually processed, and the compression parameters in the video encoding, such as bit rate, frame rate, resolution, etc., are adjusted to ensure that the transmission efficiency of the video data is maximized when the network bandwidth changes, while ensuring the integrity and smoothness of key visual information. For example, under high bandwidth, a lower compression rate is selected to retain the detailed information of the video, while under low bandwidth, the compression ratio is increased to maintain smooth video playback.
[0191] Optionally, the present specification further provides a machine vision intelligent recognition system based on digital buildings, which is used to execute the machine vision intelligent recognition method based on digital buildings as described above. The machine vision intelligent recognition system based on digital buildings includes:
[0192] The dynamic environment perception fusion module is used to obtain sensor data and camera data of the building area, and perform building area environmental status analysis based on the sensor data of the building area, so as to obtain building area environmental status data; perform dynamic environment perception fusion based on the building area environmental status data and the camera data of the building area, so as to obtain a dynamic environment perception map of the building area;
[0193] An image adaptive dynamic enhancement module is used to perform image adaptive dynamic enhancement on the camera data of the building area according to the dynamic environment perception map of the building area, so as to obtain an enhanced video frame image set of the building area;
[0194] An abnormal event detection module is used to perform abnormal event detection on the building area enhanced video frame image set to obtain an abnormal event video frame image set, and perform risk event identification on the abnormal event video frame image set to obtain a risk event video frame image set;
[0195] A network status encoding module is used to obtain building area network transmission data, and perform building area network topology structure analysis on the building area network transmission data, so as to obtain a building area network topology structure model; perform network status encoding based on the building area network topology structure model, so as to obtain network status encoding data;
[0196] A video compression module is used to perform video feature encoding data on the risk event video frame image set, thereby obtaining video feature encoding data, and to perform adaptive video compression according to the network status encoding data and the video feature encoding data, thereby obtaining building area event video compression data;
[0197] The data transmission resource allocation module is used to allocate data transmission resources for the compressed video data of the building area event through the building area network topology model, so as to obtain the building area compressed video transmission strategy and upload it to the building area risk management platform to perform the building area risk event warning task.
[0198] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.
[0199] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A machine vision intelligent recognition method based on digital buildings, characterized in that: The following steps are involved: Step S1: acquiring building area sensor data and building area camera data, and performing building area environmental status analysis based on the building area sensor data, thereby obtaining building area environmental status data; Dynamic environment perception fusion is performed based on the building area environmental status data and the building area camera data to obtain a dynamic environment perception map of the building area; Step S2: performing image adaptive dynamic enhancement on the camera data of the building area according to the dynamic environment perception map of the building area, so as to obtain an enhanced video frame image set of the building area; Step S3: performing abnormal event detection on the building area enhanced video frame image set to obtain an abnormal event video frame image set, and performing risk event identification on the abnormal event video frame image set to obtain a risk event video frame image set; Step S4: acquiring building area network transmission data, and performing building area network topology analysis on the building area network transmission data, thereby obtaining a building area network topology model; Performing network status coding based on the building area network topology structure model to obtain network status coding data; Step S5: performing video feature encoding data on the risk event video frame image set to obtain video feature encoding data, and performing adaptive video compression according to the network status encoding data and the video feature encoding data to obtain building area event video compression data; Step S6: Allocate data transmission resources for the compressed video data of the building area event through the building area network topology model, so as to obtain the building area compressed video transmission strategy, and upload it to the building area risk management platform to execute the building area risk event warning task.
2. According to claim 1, the method for intelligent machine vision recognition based on digital buildings is characterized in that: Step S1 is specifically as follows: Step S11: acquiring building area sensor data and building area camera data, and performing data preprocessing on the building area sensor data and the building area camera data respectively, so as to obtain the building area sensor data to be analyzed and the building area camera data to be analyzed; Step S12: performing sensor data time series analysis on the sensor data of the building area to be analyzed, thereby obtaining time series sensor data of the building area; Step S13: deriving the building area environment state according to the building area time series sensing data, thereby obtaining the building area environment state data; Step S14: decoding the camera video stream of the camera data of the building area to be analyzed, thereby obtaining the camera video of the building area, and dividing the camera video of the building area into time window frames, thereby obtaining a continuous video frame image set; Step S15: Perform multimodal dynamic environment perception fusion on the building area environmental status data and the continuous video frame image set, so as to obtain a dynamic environment perception map of the building area.
3. The method for intelligent machine vision recognition based on digital buildings according to claim 2 is characterized in that: Step S13 is specifically as follows: Step S131: normalizing the time series sensing data of the building area to obtain normalized time series sensing data; Step S132: performing long-short term memory time series feature classification on the normalized time series sensor data, thereby obtaining time series periodic sensor data and time series non-periodic sensor data; Step S133: performing Bayesian network modeling of building area sensor data according to the time-series periodic sensor data, thereby obtaining a building area normal environment state derivation model, and performing periodic environment state derivation through the building area normal environment state derivation model, thereby obtaining periodic environment state data; Step S134: performing principal component analysis on the time series non-periodic sensor data to obtain non-periodic reduced dimension sensor data; Step S135: performing joint distribution modeling of building area sensor data on the non-periodic reduced dimension sensor data, thereby obtaining a derivation model of abnormal environmental state of the building area, and performing non-periodic environmental state derivation through the derivation model of abnormal environmental state of the building area, thereby obtaining non-periodic environmental state data; Step S136: Time-series merge the periodic environmental status data and the non-periodic environmental status data to obtain building area environmental status data.
4. The method for intelligent machine vision recognition based on digital buildings according to claim 2 is characterized in that: Step S15 is specifically as follows: Step S151: aligning the building area environment state data and the continuous video frame image set, thereby obtaining the environment state aligned data and the continuous video frame aligned image set; Step S152: extracting image visual features from the continuous video frame aligned image set to obtain video frame visual feature data, and integrating scene visual dynamic features based on the video frame visual feature data to obtain scene visual dynamic feature data; Step S153: integrating the scene environment dynamic features according to the environment state alignment data, thereby obtaining the scene environment dynamic feature data; Step S154: performing scene dynamic feature fusion based on the scene visual dynamic feature data and the scene environment dynamic feature data, thereby obtaining a dynamic environment perception map of the building area.
5. The method for intelligent machine vision recognition based on digital buildings according to claim 1 is characterized in that: Step S2 is specifically as follows: Step S21: extracting the environment state and scene dynamics according to the dynamic environment perception map of the building area, thereby obtaining the perception map environment state data and scene dynamics data; Step S22: performing environmental change statistics based on the perception map environmental state data to obtain regional environmental change data, and performing image quality impact factor calculation on the regional environmental change data and the continuous video frame aligned image set to obtain an image quality impact factor set; Step S23: estimating the scene image enhancement strength according to the scene dynamic data and the image quality influencing factor set, thereby obtaining a scene image enhancement strength template; Step S24: performing image scene matching on the scene image enhancement strength template and the continuous video frame aligned image set, thereby obtaining an image set enhancement strength strategy; Step S25: performing image adaptive dynamic enhancement on the continuous video frame aligned image set according to the image set enhancement strength strategy, so as to obtain the building area enhanced video frame image set.
6. The method for intelligent machine vision recognition based on digital buildings according to claim 1 is characterized in that: Step S3 is specifically as follows: Step S31: performing visual feature statistics on the building area enhanced video frame image set, thereby obtaining high-frequency visual feature data and low-frequency visual feature data; Step S32: performing scene pattern recognition based on the low-frequency visual feature data to obtain normal scene pattern data; Step S33: performing video stream motion pattern capture on the high-frequency visual feature data to obtain video stream motion pattern data, and performing motion pattern classification on the video stream motion pattern data to obtain artificial motion pattern data and natural motion pattern data; Step S34: performing pattern temporal association according to the human action pattern data and the normal scene pattern data, thereby obtaining scene-human action interaction event data, and performing low-frequency interaction event recognition on the scene-human action event pattern data, thereby obtaining abnormal human interaction event data; Step S35: performing pattern temporal association on the natural action pattern data and the normal scene pattern data, thereby obtaining scene-natural action interaction event data, and performing continuous interaction event recognition on the scene-natural action interaction event data, thereby obtaining abnormal natural interaction event data; Step S36: performing time-series merging on the abnormal human interaction event data and the abnormal natural interaction event data to obtain abnormal event time-series data, and performing abnormal event time-series video frame extraction on the building area enhanced video frame image set according to the abnormal event time-series data to obtain an abnormal event video frame image set; Step S37: performing risk event identification on the abnormal event video frame image set, thereby obtaining a risk event video frame image set.
7. The method for intelligent machine vision recognition based on digital buildings according to claim 1 is characterized in that: Step S4 is specifically as follows: Step S41: acquiring building area network transmission data, and performing data preprocessing on the building area network transmission data, thereby obtaining the building area network transmission data to be analyzed; Step S42: performing transmission node connection frequency statistics on the network transmission data of the building area to be analyzed, thereby obtaining network transmission node connection frequency data, and performing node connection relationship identification based on the network transmission node connection frequency data, thereby obtaining network transmission node connection relationship data; Step S43: Modeling the building area network topology structure according to the network transmission node connection relationship data, thereby obtaining a building area network topology structure model; Step S44: evaluating the connection path network status according to the building area network topology model, thereby obtaining path network status data, and allocating connection path weights based on the path network status data, thereby obtaining connection path weight data; Step S45: encoding the connection path network status according to the connection path weight data and the path network status data, thereby obtaining network status encoding data.
8. The method for intelligent machine vision recognition based on digital buildings according to claim 1 is characterized in that: Step S5 is specifically as follows: Step S51: extracting visual features of the risk event video frame image set, thereby obtaining visual feature data of the video frame image set; Step S52: performing spatiotemporal correlation modeling based on the visual feature data of the video frame image set, thereby obtaining a spatiotemporal evolution model of the visual features; Step S53: compressing the high-dimensional features of the visual feature spatiotemporal evolution model to obtain a low-dimensional visual feature vector; Step S54: encoding the video visual features according to the low-dimensional visual feature vector, thereby obtaining video feature encoding data; Step S55: Adaptively compress the video according to the network status encoding data and the video feature encoding data, thereby obtaining building area event video compression data.
9. The method for intelligent machine vision recognition based on digital buildings according to claim 8 is characterized in that: Step S55 is specifically as follows: Step S551: calculating the video frame information volume of the risk event video frame image set, thereby obtaining risk event video frame information volume data, and dividing the risk key frames according to the risk event video frame information volume data, thereby obtaining risk key frame division data; Step S552: selecting the ratio of key frames to non-key frames for the video feature coding data based on the risk key frame classification data, thereby obtaining a video frame ratio strategy; Step S553: analyzing the compression strength dynamic adjustment strategy according to the network status coding data and the visual feature spatiotemporal evolution model, thereby obtaining the compression strength dynamic adjustment strategy; Step S554: integrating the adaptive video compression strategy based on the video frame ratio strategy and the compression strength dynamic adjustment strategy, thereby obtaining an adaptive video compression strategy; Step S555: Adaptively compress the video feature encoding data according to the adaptive video compression strategy and the network status encoding data, thereby obtaining building area event video compression data.
10. A machine vision intelligent recognition system based on digital buildings, characterized in that: For executing the method for intelligent machine vision recognition based on digital buildings as claimed in claim 1, the machine vision intelligent recognition system based on digital buildings comprises: The dynamic environment perception fusion module is used to obtain sensor data and camera data of the building area, and perform building area environmental status analysis based on the sensor data of the building area, so as to obtain building area environmental status data; perform dynamic environment perception fusion based on the building area environmental status data and the camera data of the building area, so as to obtain a dynamic environment perception map of the building area; An image adaptive dynamic enhancement module is used to perform image adaptive dynamic enhancement on the camera data of the building area according to the dynamic environment perception map of the building area, so as to obtain an enhanced video frame image set of the building area; The abnormal event detection module is used to perform abnormal event detection on the building area enhanced video frame image set to obtain an abnormal event video frame image set, and perform risk event identification on the abnormal event video frame image set to obtain a risk event video frame image set; A network status encoding module is used to obtain building area network transmission data, and perform building area network topology analysis on the building area network transmission data, thereby obtaining a building area network topology model; perform network status encoding based on the building area network topology model, thereby obtaining network status encoding data; A video compression module is used to perform video feature encoding data on a risk event video frame image set, thereby obtaining video feature encoding data, and to perform adaptive video compression according to the network status encoding data and the video feature encoding data, thereby obtaining building area event video compression data; The data transmission resource allocation module is used to allocate data transmission resources for the compressed video data of the building area event through the building area network topology model, so as to obtain the building area compressed video transmission strategy and upload it to the building area risk management platform to perform the building area risk event warning task.