Video management method based on comprehensive operation of science park
By building a multi-dimensional data acquisition network and multi-modal deep learning model in the science and technology park, the information island problem is solved, and the intelligent scheduling and optimization of park resources is realized, and operational efficiency and service quality are improved.
Patent Information
- Application Number
- CN202510104552.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional science and technology parks have information islands in terms of comprehensive operation and management, which makes managers unable to control the park's operation dynamics in an overall manner, affecting management efficiency and decision-making quality.
A video management method based on multimodal deep learning is adopted to dynamically display and optimize resource scheduling by building a multi-dimensional data acquisition network, integrating and analyzing multi-dimensional data, performing multi-modal scene perception analysis, early warning and linkage response, and generating digital twin models.
It realizes comprehensive sharing and linkage of information, improves the intelligence level of resource scheduling, dynamically adjusts the resource allocation of parks, and improves operational efficiency and service quality.
Smart Images

Figure CN120031193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of park management, and in particular to a video management method based on the comprehensive operation of a science and technology park. Background Art
[0002] As a gathering place for high-tech R&D enterprises, science and technology parks have unique innovation vitality and development potential. With the high-quality development of urban economy, science and technology parks are not only the core gathering places for industrial factor resources, but also play an important role in promoting scientific and technological innovation, optimizing industrial structure, and promoting employment. However, although many parks have made remarkable achievements in the construction of physical facilities, they still face many challenges in the comprehensive operation and management of the parks.
[0003] In existing technologies, traditional science and technology parks often focus on infrastructure construction while neglecting the improvement of user experience and services. There are information islands between the systems in the park, and information cannot be shared, resulting in park managers being unable to globally control the dynamic status of park operations, affecting management efficiency and decision-making quality. Summary of the invention
[0004] In view of the deficiencies of the prior art, the present invention provides a video management method based on the comprehensive operation of a science and technology park to solve the problems raised in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] In a first aspect, an embodiment of the present invention provides a video management method based on comprehensive operation of a science and technology park, comprising the following steps:
[0007] S1. Build a multi-dimensional data collection network: deploy cameras, environmental sensors, and equipment status sensors in the science and technology park to collect multi-dimensional information of the park, and efficiently transmit data to the park server through wireless communication protocols to form a comprehensive data collection network, thereby obtaining multi-dimensional data in the park;
[0008] S2. Multi-dimensional data integration and building a middle platform: extract, clean, convert and store the multi-dimensional data obtained from the data collection network, so as to build a unified data middle platform. Through the integration of the data middle platform, global data is formed;
[0009] S3. Multimodal scene perception analysis of global data: Use multimodal deep learning models to jointly model video and sensor data, dynamically identify the operating status of the park, and form scene analysis results;
[0010] S4. Early warning of abnormal events through scenario analysis results: Develop multi-level early warning rules and event response strategies based on the historical data of the park, incorporate scenario analysis results into the rules and strategies, detect abnormal events through pattern matching, and obtain early warning information;
[0011] S5. Cross-system linkage response based on early warning information: Using early warning information as a trigger event, it automatically collaborates with security, energy and property systems and generates linkage response data;
[0012] S6. Use response data for dynamic display of digital twins: Use the Unity3D engine and combine multi-dimensional data to generate a digital twin model of the park. By synchronizing the virtual park with the real environment, the equipment operation status, abnormal events, and personnel flow information are dynamically displayed in the twin model;
[0013] S7. Optimize resource scheduling strategies based on dynamic display: adjust the allocation of energy, space, and equipment within the park based on dynamic display information, and adjust the park operation model.
[0014] To further optimize the technical solution, in step S3, the multimodal deep learning model includes:
[0015] Video feature extraction and transformation;
[0016] Sensor feature extraction and transformation;
[0017] Multimodal joint representation generation;
[0018] Scene state prediction;
[0019] The video features and sensor features are obtained through collection and integration in step S1 and step S2, and:
[0020] V t : Video data feature matrix, representing the frame features extracted at time t, size N v ×d v , where N v is the number of video frames, d v is the feature dimension of each frame, obtained by video feature extraction;
[0021] S t : Sensor data matrix, representing the multi-sensor data collected at time t, with a size of N s ×d s , where N s is the number of sensors, d s It is the feature vector dimension of the multidimensional data recorded by each sensor in one time step, indicating the dimension of the data collected from the sensor;
[0022] W v , Ws : Weight matrices, used for linear transformation of video and sensor features respectively;
[0023] F t : Joint feature representation, representing the fusion result of video and sensor data, size is d f ;
[0024] α t , β t : Dynamic weight factor, used to control the contribution ratio of video and sensor data, which satisfies α t +β t =1;
[0025] A v,t , A s,t : Attention weight matrix, representing the importance of video and sensor features respectively;
[0026] f(·): activation function.
[0027] To further optimize the technical solution, the formula model of the video feature extraction and transformation process in step S3 is:
[0028] V t =f(A v,t ·V t ·W v );
[0029] Among them, A v,t The temporal and spatial relationships of video frames are captured through the spatiotemporal attention mechanism calculation:
[0030]
[0031] It is the matrix transposition operator, which means the transposition operation of a matrix or vector. Softmax(·) is the soft maximization function.
[0032] To further optimize the technical solution, the formula model of the sensor feature extraction and transformation process in step S3 is:
[0033] S′ t =f(A s,t ·S t ·W s );
[0034] Among them, A s,t The correlation between multiple sensors is captured through the temporal attention mechanism calculation:
[0035]
[0036] To further optimize the technical solution, the formula model of the multimodal joint representation generation process in step S3 is:
[0037] F t =α t ·V′ t +β t ·S′ t ;
[0038] The dynamic weight factor is adjusted dynamically according to the current environment changes to meet the following requirements:
[0039]
[0040] Here, σ(·) is a normalization operation used to calculate the importance of the mode.
[0041] To further optimize the technical solution, in the process of scene state prediction in step S3:
[0042] Use the union representation F t Input to the classification or regression module to achieve scene state prediction:
[0043]
[0044] where g(·) is the output function of scene state prediction, and its specific form depends on the task.
[0045] To further optimize the technical solution, in step S6, the digital twin model of the park consists of the following parts:
[0046] Physical entity modeling: Use the Unity3D engine to build a virtual model of the park, including buildings, equipment, roads, and sensors;
[0047] Data acquisition and synchronization: real-time synchronization of sensor and video data with corresponding entities in the virtual model, updating the state of the virtual environment;
[0048] Spatiotemporal dynamic simulation: Use real-time data to adjust the state of objects in the virtual scene, simulate future scenes through prediction models, and conduct intelligent management;
[0049] Prediction and decision optimization: Provide park equipment scheduling and resource allocation based on dynamic simulation results;
[0050] The following symbols are defined:
[0051] X t : The state vector of the digital twin virtual scene contains the state information of each entity in the park, including buildings, equipment, and personnel, and its size is N e ×d e , where N e is the number of entities, de It is the state information dimension of each entity;
[0052] F t : The output of the scenario prediction model, which represents the state prediction of the park at the future time t+1, with a size of N e ×d e ;
[0053] P t : Scenario optimization decision results to guide park resource scheduling;
[0054] A t : The reward value for scene optimization, used as the value function in Q-learning.
[0055] To further optimize the technical solution, in step S6, when the digital twin model of the park is used:
[0056] Input data and feature extraction: In the park, the current state of the park is obtained through real-time sensors and video data, that is, data S is obtained through steps S1 and S2. t and V t ;
[0057] Dynamic scene update: Using spatiotemporal convolutional neural network (ST-CNN) to update the virtual state of the park t Perform real-time updates to simulate dynamic changes in real-world entities such as equipment and buildings;
[0058] Future state prediction: LSTM long short-term memory network model predicts the future state of the park based on historical data t , and provide a basis for the next step of resource scheduling and optimization decision-making;
[0059] Intelligent decision-making and optimization: Reinforcement learning algorithms adjust the use of campus resources at each time step to maximize the efficiency of campus operations and optimize decisions based on feedback rewards.
[0060] To further optimize the technical solution, in step S6, the digital twin model of the park includes:
[0061] Multimodal data fusion: The video and sensor data are weightedly fused to generate the dynamic virtual state of the park. The fusion formula is:
[0062] X t =α t ·V t +β t ·S t ;
[0063] Among them, α t and β tIt is a dynamic weight coefficient that is automatically adjusted according to the environmental conditions of the current time step, satisfying α t +β t =1;
[0064] Spatiotemporal dynamic simulation and state update: The spatiotemporal convolutional neural network (ST-CNN) is used to perform spatiotemporal dynamic simulation of the park virtual scene. The state change of each entity is affected by the state of its surrounding environment, and the state is updated through convolution operations:
[0065] X t+1 =ST-CNN(X t , V t , S t );
[0066] Among them, the spatiotemporal convolutional neural network (ST-CNN) model captures the dependencies between spatial and temporal dimensions through convolution operations, simulating how each park entity changes under the current environmental conditions;
[0067] Scenario prediction and dynamic adjustment: Based on the LSTM long short-term memory network, the future state of the park is predicted. The LSTM long short-term memory network can process long-term series data and generate state predictions at future moments:
[0068] F t =LSTM(X t , V t , S t );
[0069] Among them, F t is the prediction result, which indicates the state of the park in the future time step;
[0070] Scenario optimization decision and resource scheduling: Through the reinforcement learning algorithm, based on the virtual scenario prediction results, the park resource scheduling is automatically optimized. The specific Q-learning update formula is:
[0071]
[0072] Among them, Q(s t , a t ) is the state s t Take action a t The expected reward value, α is the learning rate, γ is the discount factor, r t is the current reward, α′ is the possible subsequent action;
[0073] To further optimize the technical solution, in step S7, the process of adjusting the energy, space, and equipment allocation in the park includes:
[0074] Energy dispatch;
[0075] Space resource management;
[0076] Equipment scheduling.
[0077] In a second aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, the steps of a video management method based on the comprehensive operation of a science and technology park as described in the first aspect of the present invention are implemented.
[0078] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, the steps of a video management method based on the comprehensive operation of a science and technology park as described in the first aspect of the present invention are implemented.
[0079] Compared with the prior art, the present invention provides a video management method based on the comprehensive operation of a science and technology park, which has the following beneficial effects:
[0080] This video management method based on the comprehensive operation of the science and technology park, through a multi-system linkage framework based on multimodal deep learning, can realize the comprehensive sharing and linkage of information in the park and improve the intelligent level of resource scheduling. While solving the problem of information islands, this method also enables the park's resources to be dynamically adjusted according to real-time needs through innovative collaborative decision-making and adaptive optimization mechanisms, especially in terms of space, equipment, energy distribution, etc. in the park. The system can perform intelligent scheduling and optimization based on data such as customer distribution, equipment operating status and environmental changes, thereby significantly improving the park's operating efficiency and service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0082] Figure 1 A schematic diagram of the structure of a video management method based on comprehensive operation of a science and technology park proposed by the present invention;
[0083] Figure 2 A schematic diagram of the structure of a multimodal deep learning model of a video management method based on comprehensive operation of a science and technology park proposed in the present invention;
[0084] Figure 3 This is a schematic diagram of the digital twin model process of a science and technology park based on a video management method for comprehensive operation of a science and technology park proposed in the present invention. DETAILED DESCRIPTION
[0085] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0086] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0087] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.
[0088] Embodiment 1:
[0089] Reference Figure 1 to Figure 3 , which is the first embodiment of the present invention, and provides a video management method based on the comprehensive operation of a science and technology park, comprising the following steps:
[0090] S1. Build a multi-dimensional data collection network: deploy cameras, environmental sensors, and equipment status sensors in the science and technology park to collect multi-dimensional information of the park, and efficiently transmit data to the park server through wireless communication protocols to form a comprehensive data collection network, thereby obtaining multi-dimensional data in the park;
[0091] Video collection covers the main public areas of the park (such as entrances, parking lots, meeting areas, etc.), and sensors are deployed in key locations (such as equipment rooms, air conditioning and ventilation systems) to ensure the comprehensiveness and real-time nature of the data. At the same time, through wireless communication protocols (such as LoRa or NB-IoT), data is efficiently transmitted to the park server to form an efficient data collection network. This step is the foundation of the entire system and provides underlying support for subsequent data integration, analysis and application.
[0092] S2. Multi-dimensional data integration and building a middle platform: extract, clean, convert and store the multi-dimensional data obtained from the data collection network, so as to build a unified data middle platform. Through the integration of the data middle platform, global data is formed;
[0093] Data integration uses Alibaba's data middle platform architecture to extract, clean, transform and store multi-source data through the ETL (Extract, Transform, Load) technical process. Video data is indexed by time axis and linked with environmental sensor data to support efficient query and management of multi-dimensional data;
[0094] The key modules of the data center platform include data storage (using the distributed storage system HDFS), data processing (using the Spark big data processing framework) and data service interface (providing a unified access portal through the REST API). This center platform connects the security, energy, property and other systems in the park, solves the problem of information islands, and lays the foundation for subsequent data analysis and linkage response.
[0095] S3. Multimodal scene perception analysis of global data: Use multimodal deep learning models to jointly model video and sensor data, dynamically identify the operating status of the park, and form scene analysis results;
[0096] The multimodal deep learning model includes:
[0097] Video feature extraction and transformation:
[0098] V′ t =f(A v,t ·V t ·W v );
[0099] The calculated result of this formula is V′ t Represents the video frame feature vector after spatiotemporal attention mechanism and weighting, which can capture important temporal and spatial patterns in video data for further analysis and decision making;
[0100] Among them, A v,t The temporal and spatial relationships of video frames are captured through the spatiotemporal attention mechanism calculation:
[0101]
[0102] Is the matrix transpose operator, which means the transpose operation of a matrix or vector. If V t is an m×n matrix, It will be an n×m matrix, and softmax(·) is the soft maximization function.
[0103] Sensor feature extraction and transformation:
[0104] S′ t =f(A s,t ·S t ·W s );
[0105] Calculation result S′ t Represents the sensor data features after temporal attention mechanism and weighting, which can capture the correlation between different sensors and further help generate more accurate resource management decisions;
[0106] Among them, A s,t The correlation between multiple sensors is captured through the temporal attention mechanism calculation:
[0107]
[0108] Multimodal Joint Representation Generation:
[0109] F t =α t ·V′ t +β t .S′ t ;
[0110] The dynamic weight factor is adjusted dynamically according to the current environment changes to meet the following requirements:
[0111]
[0112] Among them, σ(·) is a normalization operation used to calculate the importance of the mode;
[0113] Scene state prediction: using joint representation F t Input to the classification or regression module to achieve scene state prediction:
[0114]
[0115] where g(·) is the output function of scene state prediction, and its specific form depends on the task.
[0116] The video features and sensor features are obtained through collection and integration in step S1 and step S2, and:
[0117] V t : Video data feature matrix, representing the frame features extracted at time t, size N v ×d v , where N v is the number of video frames, d v is the feature dimension of each frame, obtained by video feature extraction;
[0118] S t : Sensor data matrix, representing the multi-sensor data collected at time t, with a size of N s ×d s , where N s is the number of sensors, d sIt is the feature vector dimension of the multidimensional data recorded by each sensor in one time step, indicating the dimension of the data collected from the sensor;
[0119] W v , W s : Weight matrices, used for linear transformation of video and sensor features respectively;
[0120] F t : Joint feature representation, representing the fusion result of video and sensor data, size is d f ;
[0121] α t , β t : Dynamic weight factor, used to control the contribution ratio of video and sensor data, which satisfies α t +β t =1;
[0122] A v,t , A s,t : Attention weight matrix, representing the importance of video and sensor features respectively;
[0123] f(·): activation function.
[0124] S4. Early warning of abnormal events through scenario analysis results: Develop multi-level early warning rules and event response strategies based on the historical data of the park, incorporate scenario analysis results into the rules and strategies, detect abnormal events through pattern matching, and obtain early warning information;
[0125] Use the rule engine technology Apache Drools to develop multi-level warning rules and event response strategies. Input dynamic data from video and environmental sensors into the rule engine to detect abnormal events such as fire smoke, equipment overload or illegal intrusion through pattern matching.
[0126] There are multiple levels of warnings:
[0127] Low-level warning: If abnormal equipment operating parameters are detected, the operation and maintenance personnel will be reminded to check;
[0128] Advanced warning: If an illegal intrusion is found, the security department will be notified in real time via SMS, App and monitoring platform, and video recording of the relevant area will be started;
[0129] By continuously optimizing the rule base, the accuracy and timeliness of event detection can be improved.
[0130] S5. Cross-system linkage response based on early warning information: Through the event-driven architecture, a multi-system linkage framework is built, with early warning information as the trigger event, automatically collaborating with security, energy and property systems, and generating linkage response data;
[0131] The multi-system linkage framework includes:
[0132] Video surveillance system: collects park video data in real time to monitor personnel, equipment, areas, etc.
[0133] Environmental sensor system: monitors environmental factors such as temperature and humidity, power load, and airflow within the park.
[0134] Equipment management system: manage the operating status of equipment in the park (such as air conditioning, lighting, power equipment, etc.).
[0135] Park dispatching system: By integrating data from all systems, it performs tasks such as energy dispatching, equipment dispatching, and space dispatching;
[0136] Construct a formula model of a multi-system linkage framework, covering two parts: collaborative decision-making and adaptive optimization. There are N systems (such as video surveillance, sensors, equipment management, etc.), each with its own status and decision-making mechanism. The framework performs data fusion and optimization through the following formula:
[0137] x t =[x 1 , x 2 , …, x N ]: represents the state vector of each system at time, where x i Indicates the status of the i-th system (such as equipment status, environmental parameters, etc.);
[0138] u t =[u 1 ,u 2 , ..., u N ]: indicates the decision vector of each system at time t, such as equipment scheduling decision, environmental control decision, etc.;
[0139] C t : Represents the collaborative constraints between systems, including data sharing and interdependence between systems;
[0140] rt: represents the optimized system resource allocation decision (such as energy allocation, space scheduling, etc.);
[0141] e t : Indicates the environmental impact factor of each system, reflecting the impact of environmental conditions on the system (such as the impact of temperature changes on air conditioning, etc.);
[0142] For the state of each system, weighted fusion is performed to obtain a global state representation. The fusion model is:
[0143]
[0144] Among them, ω iis the weight of the i-th system, indicating the contribution of the decision of the system at the current moment to the global state. i It will be adjusted dynamically based on the importance of the system and the real-time load.
[0145] By integrating the states of each system, the framework can generate a global decision vector u t , which contains the scheduling decisions of all systems. The collaborative decision formula is:
[0146] u t =f(x t , C t );
[0147] Among them, function f(·) is a decision function, which takes into account the mutual influence between systems (such as the relationship between space load and energy demand, the linkage between video surveillance and equipment scheduling, etc.). Function f(·) is an adaptive learning model that can adjust the decision rules based on real-time data;
[0148] The system needs to continuously adjust the decision-making process through an adaptive optimization mechanism, and its formula model is:
[0149]
[0150] Among them, r i represents the resource allocation (such as energy, space, etc.) of the ith system, c i is the corresponding cost factor (such as energy consumption cost, equipment maintenance cost, etc.). The goal is to minimize the overall cost by optimizing the resource allocation of each system.
[0151] Adaptive optimization can be achieved through dynamic adjustment based on reinforcement learning. The Q-learning algorithm is used for resource scheduling optimization. The scheduling decision of each system will receive a corresponding reward value during the optimization process and will be updated through the following formula:
[0152]
[0153] Among them, Q(x t ,u t ) is the state x t Next take action u t The expected reward, α is the learning rate, γ is the discount factor, r t is the current reward;
[0154] The model is in use:
[0155] Data input: Generate the state vector xt of each system through real-time collected video monitoring data, sensor data and equipment status information;
[0156] Decision output: According to the state vector of multiple systems and the coordination constraint C t , calculate the comprehensive decision vector u of each system t ;
[0157] Optimization and adjustment: Based on the objective function of system resource allocation and the adaptive optimization model, adjust the resource allocation of each system t , to minimize costs and maximize resource utilization;
[0158] Feedback loop: The system continuously adjusts its state based on the optimized decisions and performs long-term optimization through reinforcement learning to continuously improve the system's collaborative efficiency.
[0159] S6. Use response data for dynamic display of digital twins: Use the Unity3D engine and combine multi-dimensional data to generate a digital twin model of the park. By synchronizing the virtual park with the real environment, the equipment operation status, abnormal events, and personnel flow information are dynamically displayed in the twin model;
[0160] The digital twin model of the park consists of the following parts:
[0161] Physical entity modeling: Use the Unity3D engine to build a virtual model of the park, including buildings, equipment, roads, and sensors;
[0162] Data acquisition and synchronization: real-time synchronization of sensor and video data with corresponding entities in the virtual model, updating the state of the virtual environment;
[0163] Spatiotemporal dynamic simulation: Use real-time data to adjust the state of objects in the virtual scene, simulate future scenes through prediction models, and conduct intelligent management;
[0164] Prediction and decision optimization: Provide park equipment scheduling and resource allocation based on dynamic simulation results;
[0165] The following symbols are defined:
[0166] X t : The state vector of the digital twin virtual scene contains the state information of each entity in the park, including buildings, equipment, and personnel, and its size is N e ×d e , where N e is the number of entities, d e It is the state information dimension of each entity;
[0167] F t : The output of the scenario prediction model, which represents the state prediction of the park at the future time t+1, with a size of N e ×d e ;
[0168] P t: Scenario optimization decision results to guide park resource scheduling;
[0169] A t : The reward value of the scene optimization, used for the value function in Q-learning;
[0170] When using the digital twin model of the park:
[0171] Input data and feature extraction: In the park, the current state of the park is obtained through real-time sensors and video data, that is, data S is obtained through steps S1 and S2. t and V t ;
[0172] Dynamic scene update: Using spatiotemporal convolutional neural network (ST-CNN) to update the virtual state of the park t Perform real-time updates to simulate dynamic changes in real-world entities such as equipment and buildings;
[0173] Future state prediction: LSTM long short-term memory network model predicts the future state of the park based on historical data t , and provide a basis for the next step of resource scheduling and optimization decision-making;
[0174] Intelligent decision-making and optimization: Reinforcement learning algorithms adjust the use of park resources at each time step to maximize the efficiency of park operations and optimize decisions based on feedback rewards;
[0175] The digital twin model of the park includes:
[0176] Multimodal data fusion: The video and sensor data are weightedly fused to generate the dynamic virtual state of the park. The fusion formula is:
[0177] X t =α t ·V t +β t ·S t ;
[0178] Among them, α t and β t It is a dynamic weight coefficient that is automatically adjusted according to the environmental conditions of the current time step, satisfying α t +β t =1;
[0179] Spatiotemporal dynamic simulation and state update: The spatiotemporal convolutional neural network (ST-CNN) is used to perform spatiotemporal dynamic simulation of the park virtual scene. The state change of each entity is affected by the state of its surrounding environment, and the state is updated through convolution operations:
[0180] X t+1 =ST-CNN(X t, V t , S t );
[0181] Among them, the spatiotemporal convolutional neural network (ST-CNN) model captures the dependencies between spatial and temporal dimensions through convolution operations, simulating how each park entity changes under the current environmental conditions;
[0182] Scenario prediction and dynamic adjustment: Based on the LSTM long short-term memory network, the future state of the park is predicted. The LSTM long short-term memory network can process long-term series data and generate state predictions at future moments:
[0183] F t =LSTM(X t , V t , S t );
[0184] Among them, F t is the prediction result, which indicates the state of the park in the future time step;
[0185] Scenario optimization decision and resource scheduling: Through the reinforcement learning algorithm, based on the virtual scenario prediction results, the park resource scheduling is automatically optimized. The specific Q-learning update formula is:
[0186]
[0187] Among them, Q(s t , a t ) is the state s t Take action a t The expected reward value, α is the learning rate, γ is the discount factor, r t is the current reward, and a′ is the possible subsequent action.
[0188] S7. Optimize resource scheduling strategy based on dynamic display: adjust the allocation of energy, space, and equipment within the park according to the dynamic display information, and adjust the park operation mode;
[0189] In step S7, the process of adjusting the energy, space, and equipment allocation in the park includes:
[0190] Energy dispatch: During peak hours, the energy demand in the park is usually large, especially in areas where charging piles and other equipment are frequently used. By displaying the real-time power consumption in the charging pile area through the digital twin model, the system can adjust the allocation of power resources through reinforcement learning strategies. During high-load periods, the system may increase the power allocation in the charging pile area, while during low-load periods, the system will reduce power allocation and optimize energy consumption of other equipment;
[0191] Space resource management: In the park, the use of public spaces will affect the overall resource allocation. For example, when the usage rate of a conference room is high, the digital twin model will display the "overcrowding" status of the area in real time and further adjust the allocation of resources; when certain areas are not effectively used, the opening hours and environmental control of these areas can be dynamically adjusted to avoid energy waste;
[0192] Equipment scheduling: Through equipment status sensors and video monitoring, the digital twin model will display the status of the equipment in real time, such as the low energy consumption or operating efficiency of some equipment. The reinforcement learning model can predict the possible status of the equipment in the future by analyzing historical data and current data, and dynamically schedule equipment resources based on this. For example, if an air-conditioning device consumes more energy at night, the system can adjust the opening and closing time of the device to reduce ineffective operation and save energy.
[0193] Embodiment 2:
[0194] This embodiment also provides a computer device, which is suitable for a video management method based on the comprehensive operation of a science and technology park, including a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement a video management method based on the comprehensive operation of a science and technology park as proposed in the above embodiment.
[0195] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, a video management method based on the comprehensive operation of a science and technology park as proposed in the above embodiment is implemented.
[0196] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0197] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0198] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0199] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0200] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0201] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A video management method based on the comprehensive operation of a science and technology park, characterized in that: The following steps are involved: S1. Build a multi-dimensional data collection network: deploy cameras, environmental sensors, and equipment status sensors in the science and technology park to collect multi-dimensional information of the park, and efficiently transmit data to the park server through wireless communication protocols to form a comprehensive data collection network, thereby obtaining multi-dimensional data in the park; S2. Multi-dimensional data integration and building a middle platform: extract, clean, convert and store the multi-dimensional data obtained from the data collection network, so as to build a unified data middle platform. Through the integration of the data middle platform, global data is formed; S3. Multimodal scene perception analysis of global data: Use multimodal deep learning models to jointly model video and sensor data, dynamically identify the operating status of the park, and form scene analysis results; S4. Early warning of abnormal events through scenario analysis results: Develop multi-level early warning rules and event response strategies based on the historical data of the park, incorporate scenario analysis results into the rules and strategies, detect abnormal events through pattern matching, and obtain early warning information; S5. Cross-system linkage response based on early warning information: Using early warning information as a trigger event, it automatically collaborates with security, energy and property systems and generates linkage response data; S6. Use response data for dynamic display of digital twins: Use the Unity3D engine and combine multi-dimensional data to generate a digital twin model of the park. By synchronizing the virtual park with the real environment, the equipment operation status, abnormal events, and personnel flow information are dynamically displayed in the twin model; S7. Optimize resource scheduling strategies based on dynamic display: adjust the allocation of energy, space, and equipment within the park based on dynamic display information, and adjust the park operation model.
2. A video management method based on comprehensive operation of a science and technology park according to claim 1, characterized in that: In step S3, the multimodal deep learning model includes: Video feature extraction and transformation; Sensor feature extraction and transformation; Multimodal joint representation generation; Scene state prediction; The video features and sensor features are obtained through collection and integration in step S1 and step S2, and: : Video data feature matrix, representing the time The extracted frame features are of size ,in is the number of video frames, is the feature dimension of each frame, obtained by video feature extraction; : Sensor data matrix, representing the time The multi-sensor data collected is of size ,in is the number of sensors, It is the feature vector dimension of the multidimensional data recorded by each sensor in one time step, indicating the dimension of the data collected from the sensor; : Weight matrices, used for linear transformation of video and sensor features respectively; : Joint feature representation, representing the fusion result of video and sensor data, size is ; : Dynamic weight factor, used to control the contribution ratio of video and sensor data, which satisfies ; : Attention weight matrix, representing the importance of video and sensor features respectively; : activation function.
3. A video management method based on comprehensive operation of a science and technology park according to claim 2, characterized in that: The formula model of the video feature extraction and transformation process in step S3 is: ; Among them, the calculation results represents the video frame feature vector after the spatiotemporal attention mechanism and weighting, The temporal and spatial relationships of video frames are captured through the spatiotemporal attention mechanism calculation: ; : is the matrix transposition operator, which means the transposition operation of the matrix or vector. is a soft maximization function.
4. A video management method based on comprehensive operation of a science and technology park according to claim 2, characterized in that: The formula model of the sensor feature extraction and transformation process in step S3 is: ; Among them, the calculation results represents the sensor data features after temporal attention mechanism and weighting, The correlation between multiple sensors is captured through the temporal attention mechanism calculation: 。 5. A video management method based on comprehensive operation of a science and technology park according to claim 2, characterized in that: The formula model of the multimodal joint representation generation process in step S3 is: ; The dynamic weight factor is adjusted dynamically according to the current environment changes to meet the following requirements: ; in, is a normalization operation used to calculate the importance of the mode.
6. A video management method based on comprehensive operation of a science and technology park according to claim 2, characterized in that: In the scene state prediction process in step S3: Using a joint representation Input to the classification or regression module to achieve scene state prediction: ; in It is the output function of scene state prediction, and its specific form depends on the task.
7. The video management method based on comprehensive operation of a science and technology park according to claim 1 is characterized in that: In step S6, the digital twin model of the park consists of the following parts: Physical entity modeling: Use the Unity3D engine to build a virtual model of the park, including buildings, equipment, roads, and sensors; Data acquisition and synchronization: real-time synchronization of sensor and video data with corresponding entities in the virtual model, updating the state of the virtual environment; Spatiotemporal dynamic simulation: Use real-time data to adjust the state of objects in the virtual scene, simulate future scenes through prediction models, and conduct intelligent management; Prediction and decision optimization: Provide park equipment scheduling and resource allocation based on dynamic simulation results; The following symbols are defined: : The state vector of the digital twin virtual scene contains the state information of each entity in the park, including buildings, equipment, and personnel, and its size is ,in is the number of entities, It is the state information dimension of each entity; : The output of the scenario prediction model, indicating the park’s The state prediction of ; : Scenario optimization decision results to guide park resource scheduling; : The reward value for scene optimization, used as the value function in Q-learning.
8. A video management method based on comprehensive operation of a science and technology park according to claim 7, characterized in that: In step S6, when the digital twin model of the park is used: Input data and feature extraction: In the park, the current state of the park is obtained through real-time sensors and video data, that is, data is obtained through steps S1 and S2 and ; Dynamic scene update: Using spatiotemporal convolutional neural network ST-CNN to update the virtual state of the park Perform real-time updates to simulate dynamic changes in real-world entities such as equipment and buildings; Future state prediction: LSTM long short-term memory network model predicts the future state of the park based on historical data , and provide a basis for the next step of resource scheduling and optimization decision-making; Intelligent decision-making and optimization: Reinforcement learning algorithms adjust the use of campus resources at each time step to maximize the efficiency of campus operations and optimize decisions based on feedback rewards.
9. A video management method based on comprehensive operation of a science and technology park according to claim 8, characterized in that: In step S6, the park digital twin model includes: Multimodal data fusion: The video and sensor data are weightedly fused to generate the dynamic virtual state of the park. The fusion formula is: ; in, and It is a dynamic weight coefficient that is automatically adjusted according to the environmental conditions of the current time step, satisfying ; Spatiotemporal dynamic simulation and state update: The spatiotemporal convolutional neural network ST-CNN is used to perform spatiotemporal dynamic simulation of the park virtual scene. The state change of each entity is affected by the state of its surrounding environment, and the state is updated through convolution operations: ; Among them, the spatiotemporal convolutional neural network ST-CNN model captures the dependencies between spatial and temporal dimensions through convolution operations, simulating how each park entity changes under the current environmental conditions; Scenario prediction and dynamic adjustment: Based on the LSTM long short-term memory network, the future state of the park is predicted. The LSTM long short-term memory network can process long-term series data and generate state predictions at future moments: ; in, is the prediction result, which indicates the state of the park in the future time step; Scenario optimization decision and resource scheduling: Through the reinforcement learning algorithm, based on the virtual scenario prediction results, the park resource scheduling is automatically optimized. The specific Q-learning update formula is: ; in, Yes Status Take action The expected reward value, is the learning rate, is the discount factor, is the current reward, It is a possible subsequent action.
10. The video management method based on comprehensive operation of a science and technology park according to claim 1, characterized in that: In step S7, the process of adjusting the energy, space, and equipment allocation in the park includes: Energy dispatch; Space resource management; Equipment scheduling.
Citation Information
Cited By
Park data fusion and intelligent decision-making method and system based on intelligent operation center
CN120493184A
Park data fusion and intelligent decision-making method and system based on intelligent operation center
CN120493184B
Data center digital twinborn simulation and decision-making system oriented to intelligent management
CN120724914A