Hydraulic engineering construction project full-scene cooperative command visualization system and method

By acquiring multi-modal data across all scenarios and generating safety levels using multi-modal large language models, the problems of low supervision efficiency and slow emergency response in water conservancy construction projects have been solved. This has enabled collaborative command across all scenarios, improving the quality and efficiency of safety production management and emergency response capabilities.

CN122045284APending Publication Date: 2026-05-15YELLOW RIVER ENG CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YELLOW RIVER ENG CONSULTING CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional water conservancy construction projects suffer from problems such as low supervision efficiency, slow emergency response, high on-site inspection risks, and difficulty in emergency consultation. The lack of a full-scenario collaborative command and visualization system leads to delays in accident handling.

Method used

The system employs a full-scenario multimodal data acquisition module, a security level determination module, an event information generation module, and an event labeling module. It combines drones, individual soldier systems, and fixed-position cameras to collect multi-source data, uses a multimodal large language model to generate security levels, and marks event information on an electronic map. The backend command center then conducts emergency command.

Benefits of technology

It has achieved comprehensive supervision of the construction site, improved the quality and efficiency of safety production management and emergency response capabilities, reduced the occurrence of accidents, met the management needs of personnel at different levels, and has significant economic and social benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045284A_ABST
    Figure CN122045284A_ABST
Patent Text Reader

Abstract

The invention discloses a hydraulic engineering construction project full-scene cooperative command visualization system and method, and the system carries out the full-scene data collection at a front end from a hydraulic engineering construction site, carries out the analysis of multi-modal data, generates safety event information, and assists a rear end to make an emergency decision through the visual display of a safety event. The system can complete the functions of remote safety supervision of hydraulic engineering construction projects, safety analysis of on-site multi-modal data, visual emergency command and dispatch, multi-dimensional no-dead-corner hidden danger search, efficient hidden danger rectification and verification and the like, and improves the quality and efficiency of hydraulic engineering construction safety production management and the emergency processing capability. The problems of low project group supervision efficiency, high field inspection risk, lagging emergency response and the like in safety production management of traditional hydraulic engineering construction projects are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety production management technology for water conservancy projects, and is particularly applicable to a full-scene collaborative command visualization system and method for water conservancy construction projects. Background Technology

[0002] In the safety management of traditional water conservancy construction projects, there are many pain points and difficulties. From the enterprise's perspective, the safety management of project clusters lacks a clear logical relationship. Individual safety management personnel often need to supervise dozens of construction projects across the country, requiring frequent travel to various construction sites, resulting in low supervision efficiency. Moreover, when emergencies occur on-site, the distance between headquarters and projects makes it impossible to hold emergency consultations and unified command in a timely manner, making it difficult to concentrate emergency intelligence and resources to minimize casualties and property losses. In terms of on-site safety management, traditional manual inspections are time-consuming and labor-intensive. For high-risk areas such as high pier cap beams and landslide bodies, manual inspections are difficult, dangerous, and infrequent, and verifying the effectiveness of hazard rectification in dangerous areas is also difficult. Although drones, individual soldier systems, and supporting network technologies are mature and widely used in many fields with good results, a complete full-scenario collaborative command and visualization system and method based on these technologies has not yet been formed in the field of safety management. Summary of the Invention

[0003] The purpose of this invention is to provide a full-scene collaborative command visualization system and method for water conservancy construction projects, which is used to solve the problems in traditional safety production management, such as the difficulty in effectively organizing emergency consultations and achieving unified command, as well as the time-consuming and labor-intensive on-site safety management and the difficulty in verifying the rectification effect. This invention aims to improve the efficiency and effectiveness of water conservancy project group and construction site management, as well as the convenience of emergency consultations and unified command.

[0004] To achieve the above objectives, the water conservancy engineering construction project full-scene collaborative command visualization system of the present invention includes a full-scene multimodal data acquisition module, a safety level determination module, an event information generation module, an event annotation module, and a back-end command center; The full-scene multimodal data includes a video acquisition module, a voice acquisition module, and an image acquisition module, which are used to monitor and acquire full-scene multimodal data of the construction site in real time. The safety level determination module generates the safety level of the construction site based on multimodal data from the entire scenario and using a multimodal large language model. The event information generation module matches the full-scene multimodal data with the preset event template to obtain the target event template, and fills the target event template with security level information and keywords from the full-scene multimodal data to generate event information; The event labeling module is used to mark event information on the electronic map; The back-end command center determines the safety status of the construction site based on the markings on the electronic map, and issues emergency commands based on the safety status.

[0005] Furthermore, the video acquisition module collects multi-source video data based on drones, individual soldier systems, and fixed-position cameras; drones are used to film operations on high slopes and at heights; individual soldier systems, worn by workers, are used for real-time video transmission of long-distance tunnel scenes; fixed-position cameras monitor key areas of the construction site; the voice acquisition module collects voice information from communication devices carried by workers at the construction site; the image acquisition module collects image information of personnel at the construction site; feature vectors are extracted from the preprocessed video data based on a neural network model to obtain target video feature vectors; the voice information is recognized and converted into first text information, and feature vectors of the first text information are extracted based on a bag-of-words model to obtain target text feature vectors; feature vectors are extracted from the preprocessed image information based on a neural network model to obtain target image feature vectors.

[0006] Furthermore, the multimodal large language model is obtained by fine-tuning the initial large language model based on the multimodal corpus corresponding to the knowledge graph of the preset engineering domain.

[0007] Furthermore, matching the full-scene multimodal data with the preset event template to obtain the target event template specifically includes: inputting the target video feature vector and the target image feature vector into the multimodal large language model to generate second text information; merging the first text information and the second text information to generate target text information; extracting keywords from the target text information and matching them with field information in the preset event template to obtain the target event template; the target event template includes a security level field.

[0008] Furthermore, it also includes a dynamic simulation module, which is used to provide animated demonstrations of the causes, development process, judgment and prediction, and command operations of the incident after the incident has been handled.

[0009] The present invention provides a method for full-scene collaborative command visualization of water conservancy engineering construction projects, comprising the following steps: S1, Real-time monitoring of multi-modal data across the entire construction site, determination of the safety level of the construction site based on the multi-modal data, and generation of event information by combining the multi-modal data and the safety level; S2, mark the generated event information on the electronic map; S3 determines the safety status of the construction site based on the markings on the electronic map and issues emergency commands.

[0010] Furthermore, the full-scene multimodal data includes multi-source video data, voice data, and image data; the multi-source video data, voice information, and image information are used to generate the safety level of the construction site using a multimodal large language model.

[0011] Furthermore, the safety status includes the occurrence of a safety incident.

[0012] Furthermore, the emergency command includes notifying relevant departments and personnel to handle the incident.

[0013] The advantages of this invention lie in acquiring multimodal data of the entire construction site at the front end, generating safety levels based on a multimodal large language model, and matching the multimodal data with preset event templates to obtain target event templates. Safety level information and keywords from the multimodal data are then filled into the target templates to generate event information, which is then visualized on a map. The backend command center determines the safety status of the construction site based on the map's markings and issues emergency commands accordingly. This solves the problems of low efficiency in project group supervision, high on-site inspection risks, and delayed emergency response in traditional water conservancy construction project safety management, enabling comprehensive, blind-spot-free supervision of the construction site. Through multi-role service support, it meets the management needs of personnel at different levels, improving the quality and efficiency of safety management and emergency response capabilities, resulting in significant economic and social benefits. Attached Figure Description

[0014] Figure 1 This is a block diagram of the full-scene collaborative command and visualization system for water conservancy engineering construction projects according to the present invention. Detailed Implementation

[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0016] Taking a water conservancy construction company as an example, this company has water conservancy construction projects under construction in multiple regions across the country. Traditional safety management methods suffer from low supervision efficiency and slow emergency response. To solve these problems, the company applied the full-scene collaborative command and visualization system for water conservancy construction projects of this invention. For example... Figure 1 As shown, it includes a full-scene multimodal data acquisition module, a security level determination module, an event information generation module, an event annotation module, and a backend command center.

[0017] The full-scenario multimodal data acquisition module includes a video acquisition module, a voice acquisition module, and an image acquisition module. The video acquisition module uses drones, individual soldier systems, and fixed-position cameras to collect multi-source video data. Drones are used to film operations on high slopes and at heights. Individual soldier systems, worn by workers, are used to transmit real-time video footage from long-distance tunnels. Fixed-position cameras monitor key areas of the construction site. For example, based on the construction site conditions of each project, drones are deployed in the high slope and high-altitude operation areas of each project, with 2-3 drones per project, ready to take off and film as needed. Individual soldier systems are provided to personnel working in long-distance tunnels, with each work team equipped with 1-2 sets. Fixed-position cameras are installed at key locations on the construction site, such as near electrical boxes and lifting machinery, with 10-15 fixed-position cameras installed per project.

[0018] The voice acquisition module collects voice information from the communication devices carried by workers at the construction site.

[0019] The image acquisition module is used to collect image information of personnel at the construction site.

[0020] For the acquired video data, feature vectors are extracted from the preprocessed video data based on a neural network model to obtain the target video feature vectors. The full-scene multimodal data acquisition module is equipped with a high-performance video server, capable of real-time fusion processing and storage of multi-source video data from multiple projects.

[0021] For the acquired speech information, it is recognized and converted into first text information. The feature vector of the first text information is extracted based on a bag-of-words model to obtain the target text feature vector. For the acquired image information, after preprocessing, feature vector extraction is performed on the preprocessed image information based on a neural network model to obtain the target image feature vector.

[0022] The safety level determination module generates the safety level of the construction site based on multimodal data from the entire scene and utilizes a multimodal large language model. Specifically, the target video feature vector, target text feature vector, and target image feature vector are input into a preset multimodal large language model to generate the safety level of the construction site. The preset multimodal large language model is generated by fine-tuning an initial large language model based on a multimodal corpus corresponding to a knowledge graph in a preset engineering domain.

[0023] The event information generation module matches multimodal data from the entire scene with a preset event template to obtain a target event template. It then fills the target event template with security level information and keywords from the multimodal data to generate event information. Specifically, it inputs the target video feature vector and the target image feature vector into a preset multimodal large language model to generate second text information. The first text information and the second text information are then merged to generate target text information. Keywords from the target text information are extracted and matched with field information in the preset event template to obtain the target event template, which includes a security level field. Finally, security level information and keywords from the multimodal data are filled into the target event template to generate event information.

[0024] The event labeling module is used to mark event information on the electronic map. The backend command center determines the safety status of the construction site based on the map's markings and issues emergency commands accordingly. Safety status includes the occurrence of a safety accident; issuing emergency commands based on the safety status includes notifying relevant departments to handle the accident based on its type.

[0025] This invention also provides a visualization method for full-scene collaborative command of water conservancy engineering construction projects, specifically including the following steps: S1. Monitor the multi-modal data of the entire construction site in real time, determine the safety level of the construction site based on the multi-modal data of the entire construction site, and generate event information by combining the multi-modal data of the entire construction site and the safety level.

[0026] The full-scene multimodal data includes multi-source video data, voice data, and image data. Among them, multi-source video data is collected based on drones, individual soldier systems, and fixed-position cameras. Drones are used to film the operation on high slopes and at high altitudes. Individual soldier systems are worn by workers to transmit real-time video back to the scene of long-distance tunnels. Fixed-position cameras monitor key parts of the construction site.

[0027] The system collects voice information from the communication devices carried by workers at the construction site, as well as image information of the personnel on site.

[0028] The system performs preprocessing on multi-source video data, then fuses the preprocessed multi-source video data to generate target video data. Finally, it extracts feature vectors from the target video data based on a pre-defined neural network model to obtain the target video feature vectors. The speech information is used to perform speech recognition to obtain first text information, and the feature vector of the first text information is extracted based on a preset bag-of-words model to obtain the target text feature vector. Image information is extracted using a pre-defined neural network model to obtain the feature vector of the target image. The target video feature vector, target text feature vector, and target image feature vector are input into a preset multimodal large language model to generate the safety level of the construction site.

[0029] The preset multimodal large language model is generated by fine-tuning the initial large language model using multimodal corpus corresponding to the knowledge graph of the preset engineering domain.

[0030] The target video feature vector and the target image feature vector are input into a preset multimodal large language model to generate second text information. The first text information and the second text information are then merged to generate the target text information. Keywords from the target text information are extracted and matched with field information in a preset event template to obtain a target event template. The target event template includes a security level field. Security level information and keywords from the full-scene data are then populated into the target event template to generate event information.

[0031] S2. Mark the generated event information on the electronic map.

[0032] S3. Determine the safety status of the construction site based on the markings on the electronic map and issue emergency commands.

[0033] The safety status includes the occurrence of a safety accident, and issuing emergency commands based on the safety status includes notifying relevant departments and personnel to carry out accident handling according to the type of safety accident.

[0034] In some embodiments, the full-scene collaborative command visualization system for water conservancy construction projects of the present invention further includes a dynamic simulation module, which is used to perform animated demonstrations of the cause, development process, analysis and prediction, and command operations of the event information after the safety accident has been handled.

[0035] The types of safety accidents described in this invention include operator violations, fires, and rockfalls from high slopes. For example, staff at the back-end command center can use remote safety monitoring to view real-time video footage of key areas of each project. If they discover violations in the operation of cranes in a certain project, they can promptly send a reminder message to the on-site management personnel through the system. Upon receiving the message, the on-site management personnel will immediately require the operators to stop the violations and rectify the situation. When a fire occurs in a project, on-site personnel can urgently call the company's headquarters data center through the system. The system will immediately transmit the on-site video back, and the company's emergency leadership team can quickly formulate a rescue plan based on the video footage, allocate rescue resources, and promptly control the fire, reducing the accident losses. On-site safety management personnel can use the multi-dimensional, no-blind-spot hazard detection function and drones to discover rockfall hazards on high slopes of a project, and promptly arrange personnel to clear and reinforce them; they can use individual soldier systems to discover water seepage in long-distance tunnels and promptly take drainage and support measures; and they can use fixed-position cameras to discover damaged protective facilities in key areas and promptly arrange for repairs.

[0036] By applying the system and method of this invention, the water conservancy construction company has significantly improved the quality and efficiency of safety production management, reduced the occurrence of production safety accidents, and enhanced its emergency response capabilities, achieving good results.

[0037] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope defined in the claims of the present invention.

Claims

1. A full-scene collaborative command and visualization system for water conservancy engineering construction projects, characterized in that: It includes a full-scenario multimodal data acquisition module, a security level determination module, an event information generation module, an event annotation module, and a backend command center; The full-scene multimodal data includes a video acquisition module, a voice acquisition module, and an image acquisition module, which are used to monitor and acquire full-scene multimodal data of the construction site in real time. The safety level determination module generates the safety level of the construction site based on multimodal data from the entire scenario and using a multimodal large language model. The event information generation module matches the full-scene multimodal data with the preset event template to obtain the target event template, and fills the target event template with security level information and keywords from the full-scene multimodal data to generate event information; The event labeling module is used to mark event information on the electronic map; The back-end command center determines the safety status of the construction site based on the markings on the electronic map, and issues emergency commands based on the safety status.

2. The full-scene collaborative command and visualization system for water conservancy engineering construction projects as described in claim 1, characterized in that: The video acquisition module collects multi-source video data based on drones, individual soldier systems, and fixed-position cameras. Drones are used to film operations on high slopes and at heights; individual soldier systems, worn by workers, transmit real-time video of long-distance tunnels; fixed-position cameras monitor key areas of the construction site; the voice acquisition module collects voice information from communication devices carried by workers at the construction site; the image acquisition module collects image information of personnel at the construction site; feature vectors are extracted from preprocessed video data using a neural network model to obtain target video feature vectors; the voice information is converted into first text information, and feature vectors are extracted from the first text information using a bag-of-words model to obtain target text feature vectors; feature vectors are extracted from preprocessed image information using a neural network model to obtain target image feature vectors.

3. The full-scene collaborative command and visualization system for water conservancy engineering construction projects as described in claim 1, characterized in that: The multimodal large language model is obtained by fine-tuning the initial large language model based on the multimodal corpus corresponding to the knowledge graph of the preset engineering domain.

4. The full-scene collaborative command and visualization system for water conservancy engineering construction projects as described in claim 2, characterized in that: The process of matching multimodal data across the entire scene with a preset event template to obtain a target event template specifically includes: inputting the target video feature vector and the target image feature vector into the multimodal large language model to generate second text information; merging the first text information and the second text information to generate target text information; and extracting keywords from the target text information and matching them with field information in the preset event template to obtain the target event template. The target event template includes a security level field.

5. The full-scene collaborative command and visualization system for water conservancy engineering construction projects as described in claim 1, characterized in that: It also includes a dynamic simulation module, which is used to provide animated demonstrations of the causes, development process, judgment and prediction, and command operations of the incident after the incident has been handled.

6. A visualization method for full-scene collaborative command of water conservancy engineering construction projects, characterized in that, Includes the following steps: S1, Real-time monitoring of multi-modal data across the entire construction site, determination of the safety level of the construction site based on the multi-modal data, and generation of event information by combining the multi-modal data and the safety level; S2, mark the generated event information on the electronic map; S3 determines the safety status of the construction site based on the markings on the electronic map and issues emergency commands.

7. The full-scene collaborative command visualization method for water conservancy engineering construction projects as described in claim 6, characterized in that: The full-scene multimodal data includes multi-source video data, voice data, and image data; the multi-source video data, voice information, and image information are used to generate the safety level of the construction site using a multimodal large language model.

8. The full-scene collaborative command visualization method for water conservancy engineering construction projects as described in claim 6, characterized in that: The safety status includes the occurrence of a safety incident.

9. The full-scene collaborative command visualization method for water conservancy engineering construction projects as described in claim 6, characterized in that: The emergency command includes notifying relevant departments and personnel to handle the accident.