Building equipment remote multi-modal intelligent feed control system and method
By constructing a remote multimodal intelligent control system for construction equipment, multimodal data is collected and analyzed in real time to generate precise control commands, solving the problems of low efficiency and high safety risks in multi-equipment collaborative operations, and realizing intelligent construction management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI CONSTRUCTION GROUP CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-06-09
AI Technical Summary
The existing construction equipment management and control system operates independently, resulting in low efficiency and high safety risks in multi-equipment collaborative operations, and making it impossible to achieve refined management.
A remote multimodal intelligent feed control system for construction equipment is constructed. Data is collected in real time through multimodal sensing devices, and intelligent analysis and decision-making are performed using deep learning algorithms to generate precise control commands. These commands are then sent to the equipment controller for execution via a communication network, which monitors the equipment status in real time and triggers anomaly detection and safety strategies.
It enables real-time and intelligent control of multi-equipment collaborative operations, improves construction efficiency and safety, reduces safety risks, and supports reliable environmental perception and dynamic optimization decision-making around the clock.
Smart Images

Figure CN122172655A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a remote multimodal intelligent control system and method for construction equipment, belonging to the field of construction equipment technology. Background Technology
[0002] In the current field of construction equipment management and control, existing technologies typically employ independent systems to separately monitor the construction site, operate equipment, and schedule tasks. For example, site monitoring often relies on isolated video surveillance systems, whose collected data is mainly used for post-event tracing or manual monitoring by safety officers, lacking real-time, intelligent analysis capabilities for the site environment and equipment status. Meanwhile, equipment control primarily relies on manual operation by operators based on personal experience or simple point-to-point remote control. This approach cannot proactively perceive potential risks within the work area and makes precise coordination between multiple pieces of equipment difficult. Therefore, monitoring systems, equipment control systems, and project scheduling systems often operate as independent information silos, with a severe disconnect between perception and control. This separation leads to a heavy reliance on traditional communication methods such as walkie-talkies for on-site personnel during multi-equipment collaborative operations. This is not only inefficient but also highly susceptible to serious safety accidents such as equipment collisions and personnel accidentally entering dangerous areas due to information transmission delays, misunderstandings, or human error. In addition, traditional construction scheduling decisions are mostly based on static construction plans, which cannot be dynamically optimized and adjusted in a closed-loop feedback manner according to unexpected situations on site, real-time load of equipment and operating efficiency, resulting in low equipment utilization, excessive waiting time and limited overall construction efficiency.
[0003] Therefore, there is an urgent need for an integrated system and method that can deeply integrate multimodal perception, intelligent decision-making and collaborative control to systematically improve the safety, efficiency and intelligence of multi-equipment operations in large-scale construction projects. Summary of the Invention
[0004] In view of the problems in the existing technology where the monitoring, decision-making and control systems of construction equipment are independent and do not share information, resulting in low efficiency, high safety risks and inability to achieve refined management in multi-equipment collaborative operations, this invention provides a remote multimodal intelligent feed control system and method for construction equipment, which constructs an integrated intelligent closed loop of "perception-decision-control" to improve the efficiency and safety of multi-equipment collaborative operations.
[0005] To solve the above technical problems, the present invention includes the following technical solutions:
[0006] A remote multimodal intelligent feed control method for construction equipment includes the following steps:
[0007] S1. Real-time collection of multi-source heterogeneous data streams containing visual information, operator interaction intent information, and equipment operating status information through a network of field sensing devices;
[0008] S2. Based on deep learning algorithms, the collected visual information is analyzed in real time to realize the parsing from the original image to the high-level operation intention and form executable discretized operation instructions;
[0009] S3. Convert discretized operation commands into specific control target parameters that can be executed by the equipment;
[0010] S4. Encode the control target parameters into low-level command messages that conform to the industrial communication protocol and send them to the equipment controller for execution via the communication network;
[0011] S5. Continuously monitor the deviation between the actual state and the expected state of the equipment, perform anomaly detection through residual calculation, and trigger alarms or safety policies when the threshold is exceeded.
[0012] Furthermore, in step S1, an edge computing architecture is used for data preprocessing, and edge servers are deployed at the construction site to achieve data dimensionality reduction and feature extraction; for low-light scenes, an adaptive enhancement algorithm is used to improve image quality; and a weighted average method is used for multispectral image fusion. ,in Visible light image, Infrared image, For the first The weight corresponding to the sub-risk; weight By calculating image information entropy Adaptive adjustment ensures reliable sensing capabilities in all-weather operating environments.
[0013] Furthermore, in step S2, the collected visual information is analyzed in real time based on a deep learning algorithm to realize the parsing from the original image to the high-level operational intent, specifically as follows:
[0014] Using object detection networks to analyze images The system locates key targets and extracts a set of structured key feature points. , This represents the total number of key feature points.
[0015] Based on the key feature point set Calculate a feature vector that can characterize the operational intention D. ; ,in It is a mapping function used to calculate the geometric relationship between the relative positions, distances, or angles between feature points;
[0016] The calculated feature vector Input a pre-trained classifier It is parsed into a predefined discretized operation instruction pattern. ;
[0017] ,
[0018] in, It is a finite set of instructions.
[0019] Furthermore, in step S3, the parsed high-level operational intent is converted into specific control target parameters that the equipment can execute, specifically:
[0020] For multi-equipment scenarios, the optimization is extended to multi-target optimization; for single-equipment scenarios, the intent is directly mapped to the control target.
[0021] Discretization operation instruction mode Transformed into specific control objectives that can be executed by the equipment. , In the formula, In order to convey the operational intent Mapping to control target The decision function.
[0022] Furthermore, the interaction information includes gesture information, in step S2,
[0023] In gesture recognition scenarios, a hand keypoint detection model is used to extract the coordinates of the thumb tip. and the coordinates of the tip of the index finger Calculate the direction difference The gesture type is determined based on the threshold.
[0024] Furthermore, in step S5, the actual status of the equipment is continuously monitored. With the expected state To achieve precise control and anomaly monitoring of deviations, specifically including:
[0025] Construct vectors to describe the operational status of equipment , which includes location ,attitude Other key operating parameters;
[0026] The health status of system operation is assessed by calculating the norm of the state deviation, and the residual is defined. , when the residual Exceeding the preset dynamic or static threshold At that time, that is The system will determine this as an anomaly and trigger an alarm or corresponding security policy.
[0027] Furthermore, in step S5, a cognitive feedback and autonomous deductive control method based on a large visual language behavior model is used as the actual state. The monitoring and feedback methods include the following sub-steps:
[0028] S51. When the regular anomaly detection module detects the state deviation norm Exceeding the preset threshold At that time, visual information at the current moment Equipment state vector and deviation information The input is fed into a pre-trained visual language behavior model; the visual language behavior model parses the abnormal scene and outputs two items: a) a natural language text describing the nature of the abnormality. b) A measure of cognitive uncertainty This value quantifies the model's confidence in its understanding of the current scene;
[0029] S52. Measurement of cognitive uncertainty With a preset threshold of cognitive uncertainty Compare: a) If This indicates that the model has a high degree of confidence in the current situation, and the system will autonomously enter sub-step S53; b) If This indicates that the model has a vague understanding of the current situation or has encountered a completely new context. The system will automatically trigger the human-machine collaboration mode and translate the natural language description. The system presents relevant visual information to remote operators and provides a dialogue interface, allowing operators to lead or assist the model in making decisions.
[0030] S53. Enter autonomous decision-making mode; the system will send a high-level text instruction. The input is fed into the VLA model; the VLA model dynamically generates a sequence of sub-actions based on its understanding of the scene. The solution sequence is constructed, and the behavioral entropy of the sequence is calculated simultaneously. :
[0031] ,
[0032] in, The model generates the first Size of movement The probability of;
[0033] S54. Based on the behavioral entropy The value is used to apply a hierarchical verification strategy to the generated action sequence:
[0034] a) If If the risk level is low, the physical equipment will be deployed directly after a rapid and secure simulation is performed in the digital twin environment.
[0035] b) If If the simulation is in a high-risk zone, explicit authorization from a remote operator is required before execution.
[0036] Furthermore, the aforementioned remote multimodal intelligent feed control method for construction equipment also includes:
[0037] S6. By transforming coordinates, the physical equipment status is mapped to the virtual 3D environment in real time, and the model rendering and performance index display are dynamically updated. Specifically:
[0038] Use the WebGL framework to build 3D visualization scenes;
[0039] Establish basic scene elements including ground grid, coordinate axes, and lighting;
[0040] Load the 3D model of the equipment and set different initial poses and colors according to the equipment type;
[0041] Equipment status data is pushed to the front end via a WebSocket server at a preset frequency;
[0042] After receiving the data, the front end updates the rotation matrix and position coordinates of the 3D model in real time: for rotating parts, the rotation angle is accumulated according to the gesture state. For moving parts, the displacement is accumulated based on the gesture state. The equipment's operational status is displayed intuitively through color gradients and numerical labels.
[0043] Furthermore, the aforementioned remote multimodal intelligent feed control method for construction equipment also includes:
[0044] S7. Establish a comprehensive risk assessment index that integrates multiple risk sources, and trigger tiered response strategies based on risk levels, up to and including emergency braking; specifically:
[0045] Establish comprehensive risk indicators By using a weighted summation method, risk sources from different sources can be integrated:
[0046] ,
[0047] in, Representing various sub-risks, Its corresponding weight;
[0048] According to risk indicators The value and the preset multi-level threshold Comparisons are made to trigger different levels of response strategies.
[0049] Accordingly, the present invention also provides a remote multimodal intelligent feed control system for construction equipment, comprising:
[0050] The construction site perception module is responsible for the real-time acquisition and fusion processing of multi-source heterogeneous information, including a multi-source visual acquisition unit, an environmental monitoring unit, and a positioning and tracking unit.
[0051] The equipment status recognition module uses computer vision and deep learning technologies to achieve intelligent recognition and status analysis of construction equipment; this module includes an equipment detection unit, an attitude estimation unit, and an operation recognition unit.
[0052] The collaborative control module enables seamless integration and collaborative control with the construction equipment control system; this module includes an industrial interface unit, a fieldbus unit, a wireless communication unit, and a safety interlock unit.
[0053] The scheduling decision module uses artificial intelligence algorithms to realize intelligent scheduling and optimization decisions for construction equipment, and supports cognitive decision-making based on a large visual language behavior model, including operation planning unit, path optimization unit, conflict resolution unit and load balancing unit.
[0054] The digital twin module constructs a virtual mapping of the construction site, enabling real-time synchronization between the physical and digital worlds. It includes a model integration unit, a real-time rendering unit, a physical simulation unit, a data mapping unit, and an interactive control unit.
[0055] The present invention, by adopting the above technical solutions, has the following advantages and positive effects compared with the prior art: The remote multimodal intelligent feed control method for construction equipment provided by the present invention collects multi-source heterogeneous information such as vision, environment and location in real time through multimodal perception modules deployed at the construction site, uses artificial intelligence technology to intelligently identify and estimate the state of the equipment, and performs collaborative scheduling and dynamic path optimization of the equipment based on this. Finally, the optimization decision is transformed into precise collaborative control commands and issued to each piece of equipment for execution, thereby constructing an integrated intelligent operation process from environmental perception, intelligent decision-making to closed-loop control, which can solve the problems of low efficiency and high safety risks of multi-equipment collaborative operation in the traditional mode. Attached Figure Description
[0056] Figure 1 This is a flowchart of remote multimodal intelligent feed control of construction equipment in one embodiment of the present invention. Detailed Implementation
[0057] The remote multimodal intelligent feed control system and method for construction equipment provided by the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of the present invention will become clearer from the following description. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.
[0058] Example 1
[0059] like Figure 1 As shown, this embodiment provides a remote multimodal intelligent feed control method for construction equipment, including the following steps:
[0060] Step S1. Multimodal Information Acquisition: Real-time acquisition of multi-source heterogeneous data streams containing visual information, operator interaction intent information, and equipment operating status information via a network of field sensing devices. Visual information is a sequence of dynamic images captured by a camera network. Interaction information is captured through non-contact methods such as visual analysis to obtain operator control intent information, such as gestures and postures. Status information includes the equipment's own pose, operating parameters, and environmental state vectors acquired by various sensors.
[0061] Step S2. Intelligent Recognition and Intent Parsing: Based on deep learning algorithms, the collected visual information is analyzed in real time to realize the parsing from the raw image to the advanced operation intent, forming executable discretized operation commands. This step is executed through the equipment status recognition module. Specifically, it includes:
[0062] (1) Key target detection and feature extraction: A target detection network is used to extract key targets from images. The system locates key targets (such as the operator's hand) and extracts a set of structured key feature points. . For key feature point set The first in One key point, =1,2,…,n, where n is the total number of key feature points.
[0063] (2) Construction of task intent features: based on the set of key feature points Calculate an eigenvector that can characterize the D-map of the operation. This calculation process can be represented as a function. ,in It is a mapping function used to calculate the geometric relationships such as relative position, distance or angle between feature points.
[0064] (3) Intent classification: The calculated feature vectors Input a pre-trained classifier It is parsed into a predefined discretized operation instruction pattern. .
[0065] ,
[0066] in, Belongs to a finite set of instructions, for example .
[0067] Step S3. Control Decision Generation: This step transforms discretized operation commands into specific control target parameters executable by the equipment. This step is executed by the scheduling decision module. For multi-equipment scenarios, this is expanded to multi-objective optimization; for single-equipment scenarios, the intention is directly mapped to the control target. This module acts as the control hub, converting the discretized operation command patterns parsed in the previous step... Transformed into specific control objectives that can be executed by the equipment. .
[0068] This process is a decision mapping function: .
[0069] in, This includes a set of control parameters, such as target velocity, angular velocity, or target pose. In complex scenarios such as multi-equipment collaboration, this step can be extended to solve a multi-objective optimization problem with time, safety, and other objectives as the goals.
[0070] Step S4. Control Command Generation and Issuance: The control target parameters are encoded into low-level command messages conforming to industrial communication protocols and issued to the equipment controller for execution via the communication network. This step is performed by the collaborative control module. The abstract control target generated in the previous step is then... This is encoded into low-level command messages that the target equipment controller can directly recognize and execute. This step includes:
[0071] (1) Instruction encoding: to control the target Convert into instruction packets that conform to a specific industrial communication protocol format. This process can be represented as In the formula, It is to control the target Encoded as instruction messages The function.
[0072] (2) Command issuance: The command packet is transmitted via a communication network such as fieldbus or industrial Ethernet. Reliably transmit data to the target equipment's controller (such as a PLC) and trigger execution.
[0073] Step S5. Real-time Status Monitoring and Feedback: Continuously monitor the deviation between the actual and desired equipment status, perform anomaly detection through residual calculation, and trigger alarms or safety strategies when thresholds are exceeded. This step is integral throughout the system, forming the core of closed-loop control. This is achieved by continuously monitoring the actual status of the equipment. With the expected state To achieve precise control and anomaly monitoring by addressing deviations between parameters. This includes:
[0074] (1) Definition of state vector: Construct a vector to describe the operating state of the equipment. , which includes location ,attitude And other key operating parameters.
[0075] (2) Anomaly Detection: The health status of the system is assessed by calculating the norm of the state deviation. The residual is defined. . When the residual Exceeding the preset dynamic or static threshold At that time, that is The system will determine this as an anomaly and trigger an alarm or corresponding security policy.
[0076] The remote multimodal intelligent feed control method for construction equipment provided in this embodiment collects multi-source heterogeneous information such as vision, environment, and location in real time through multimodal perception modules deployed at the construction site. It uses artificial intelligence technology to intelligently identify and estimate the state of the equipment, and performs collaborative scheduling and dynamic path optimization based on this. Finally, the optimization decision is transformed into precise collaborative control commands and issued to each piece of equipment for execution, thereby constructing an integrated intelligent operation process from environmental perception and intelligent decision-making to closed-loop control. This method can solve the problems of low efficiency and high safety risks in multi-equipment collaborative operation under the traditional mode.
[0077] In one specific embodiment, the remote multimodal intelligent feed control method for construction equipment further includes:
[0078] Step S6. Digital Twin Visualization Synchronization: The physical equipment status is mapped to the virtual 3D environment in real time through coordinate transformation, and the model rendering and performance indicator display are dynamically updated. This step is performed by the digital twin module. Specifically, it includes:
[0079] (1) State mapping: Through homogeneous coordinate transformation, the equipment state (such as pose) in the physical world is mapped to the digital twin virtual environment in real time through the transformation matrix.
[0080] (2) Real-time rendering: Based on the mapped pose information, the model matrix of the digital twin model in the three-dimensional scene is updated in real time. This matrix is composed of translation, rotation and scaling matrices. At the same time, key performance indicators (such as velocity and displacement) are overlaid on the model in the form of labels or charts to achieve information-enhanced visualization.
[0081] In one specific embodiment, the remote multimodal intelligent feed control method for construction equipment further includes:
[0082] S7. Safety Early Warning and Emergency Response: Establish a comprehensive risk assessment index that integrates multiple risk sources, triggering tiered response strategies based on risk levels up to emergency braking. This step is executed collaboratively by the collaborative control module and the scheduling decision module, providing proactive safety assurance for the entire operation process. Specifically:
[0083] (1) Multi-level risk assessment: Establish a comprehensive risk index By using a weighted summation method, risk sources from different sources can be integrated:
[0084] ,
[0085] in, These represent various sub-risks, such as collision risk, over-limit risk, and personnel intrusion risk. Set its corresponding weight.
[0086] (2) Tiered response: based on risk indicators The value and the preset multi-level threshold Comparisons are made to trigger different levels of response strategies. For example, when When, trigger a visual warning (such as changing the model's color in a digital twin); when When the equipment speed is reached, the speed is automatically limited; when the critical threshold is reached... If necessary, immediately execute the emergency braking or safety shutdown procedure.
[0087] Repeat steps S1 to S7 to form a continuous control loop of "perception-recognition-decision-execution-feedback" until all predetermined tasks are completed, thereby realizing remote, real-time, and closed-loop intelligent control of the construction equipment.
[0088] In one specific embodiment, in step S1, an edge computing architecture is used for data preprocessing. Edge servers are deployed at the construction site to achieve data dimensionality reduction and feature extraction, effectively reducing transmission bandwidth requirements. For low-light scenes, an adaptive enhancement algorithm is used to improve image quality. Multispectral image fusion uses a weighted average method. ,in Visible light image, For infrared images, weights By calculating image information entropy Adaptive adjustment ensures reliable sensing capabilities in all-weather operating environments.
[0089] In one specific embodiment, the operator's interaction intent information includes gesture information. In step S2, for a gesture recognition scenario, a hand key point detection model is used to extract the coordinates of the thumb tip. and the coordinates of the tip of the index finger Calculate the direction difference The gesture type is determined based on a threshold: when... When the pixel is determined to be a left-finger gesture, If the pixel value is within the specified range, it is determined to be a right-pointing gesture; otherwise, it is determined to be a no-gesture state. Simultaneously, extract... Parameters such as these are used for more complex gesture recognition.
[0090] In one specific embodiment, in step S3, for a single-equipment gesture control scenario, a control decision is directly generated based on the recognized gesture type: a left-finger gesture corresponds to a counter-clockwise rotation or leftward movement command, a right-finger gesture corresponds to a clockwise rotation or rightward movement command, and no gesture generates a stop action command. Control parameters, such as rotation speed set to 0.05 rad / s and movement speed set to 0.001 m / step, can be adjusted according to actual needs.
[0091] In one specific embodiment, step S5 employs a cognitive feedback and autonomous deductive control method based on a vision-language-behavior (VLA) big model as the real-time state monitoring and feedback method, including the following sub-steps:
[0092] S51. Cognitive Understanding and Quantitative Assessment of Abnormalities.
[0093] When the conventional anomaly detection module detects the state deviation norm Exceeding the preset threshold At that time, visual information at the current moment Equipment state vector and deviation information The input is fed into a pre-trained Visual-Language-Behavior (VLA) large-scale model; the VLA model parses the abnormal scene and outputs two items: a) a natural language text describing the nature of the abnormality. b) A measure of cognitive uncertainty This value quantifies the model's confidence in its understanding of the current scene.
[0094] S52. Dynamic switching of decision-making power in human-machine collaboration.
[0095] The cognitive uncertainty measure With a preset threshold of cognitive uncertainty Compare: a) If This indicates that the model has a high degree of confidence in the current situation, and the system will autonomously enter sub-step (3); b) if This indicates that the model has a vague understanding of the current situation or has encountered a completely new context. The system will automatically trigger the human-machine collaboration mode and translate the natural language description. The system presents relevant visual information to remote operators and provides a dialogue interface, allowing operators to lead or assist the model in making decisions.
[0096] S53. Autonomous Generative Behavior Planning and Risk Assessment.
[0097] In autonomous decision-making mode, the system will send a high-level text instruction. (For example, automatically generated system-generated or pre-text instructions such as "safely bypass the obstacle") are input into the VLA model; the VLA model dynamically generates a sequence of sub-actions based on its understanding of the scene. The solution sequence is constructed, and the behavioral entropy of the sequence is calculated simultaneously. :
[0098] ,
[0099] in, The model generates the first Size of movement The probability of the behavior entropy. To assess the risk and innovativeness of the generation strategy: lower entropy indicates a conventional, high-confidence approach; higher entropy indicates an unconventional, exploratory approach that may require a higher level of security monitoring or operator verification.
[0100] S54. Hierarchical verification and closed-loop self-learning.
[0101] According to the behavioral entropy The value is used to perform a hierarchical verification strategy on the generated action sequence: a) If If the risk level is low, then a rapid and secure simulation is performed directly in the digital twin environment before the physical equipment is deployed for execution; b) If If the scenario is in a high-risk zone, explicit authorization from a remote operator is required after simulation before execution. Successful physical execution will include (abnormal scenarios, language interaction, solutions, execution results, and cognitive uncertainty). Behavioral entropy The complete data stream is packaged into an experience sample for online fine-tuning of the VLA large model, thereby enabling the system to learn and evolve from experience in handling anomalies.
[0102] In one specific embodiment, in step S6, a 3D visualization scene is constructed using the WebGL framework. Basic scene elements, including a ground mesh, coordinate axes, and lighting, are established. The 3D model of the equipment (e.g., in GLB format) is loaded, and different initial poses and colors are set according to the equipment type. Equipment status data, including gesture states, rotation angles, and displacements, is pushed to the front end via a WebSocket server at 10ms intervals. After receiving the data, the front end updates the rotation matrix and position coordinates of the 3D model in real time: for rotating components, the rotation angle is accumulated based on the gesture state. For moving parts, the displacement is accumulated based on the gesture state. The equipment's operating status is displayed intuitively through color gradients (such as linear interpolation from the initial color to the warning color) and numerical labels (such as displaying "0.05rad / s", "10.25mm", etc.).
[0103] In one specific embodiment, the entire control loop from steps S1 to S7 runs with a period of 10ms to 100ms to ensure the real-time performance of gesture recognition, control decision-making, PLC communication, and digital twin visualization. Multithreading technology is used to run the image processing thread, PLC communication thread, and WebSocket server thread in parallel, using a thread lock mechanism to protect shared data structures and avoid data contention. The main thread is responsible for image acquisition and gesture recognition, the communication thread is responsible for issuing PLC commands, and the server thread is responsible for pushing status data; the three work together to improve system response speed and stability.
[0104] Example 2
[0105] This embodiment provides a remote multimodal intelligent feed control system for construction equipment. It adopts a modular and layered architecture design to achieve intelligent control of the entire process of perception, identification, decision-making, control and visualization, and includes the following modules.
[0106] (1) Construction Site Perception Module. The construction site perception module is responsible for the real-time acquisition and fusion processing of multi-source heterogeneous information. This module includes a multi-source visual acquisition unit, an environmental monitoring unit, and a positioning and tracking unit. The multi-source visual acquisition unit deploys a high-definition camera array at key locations on the construction site to achieve multi-angle, full-coverage visual information acquisition, ensuring continuous tracking of dynamic targets, and supporting multi-spectral imaging such as visible light and infrared to adapt to different lighting conditions. The environmental monitoring unit integrates multiple types of environmental sensors such as wind speed and direction sensors, tilt sensors, load sensors, and temperature and humidity sensors to monitor environmental parameters at the construction site in real time, providing data support for safe operations. The positioning and tracking unit adopts multi-source positioning fusion technology, combined with ultra-wideband positioning (UWB), satellite positioning (GNSS), and other methods to achieve high-precision positioning and tracking of equipment and personnel.
[0107] (2) Equipment Status Recognition Module. The equipment status recognition module uses computer vision and deep learning technologies to achieve intelligent recognition and status analysis of construction equipment. This module consists of an equipment detection unit, an attitude estimation unit, and a job recognition unit. The equipment detection unit is based on a deep learning target detection algorithm to achieve real-time detection and classification of construction equipment, supports multi-scale detection, and can identify equipment targets at different distances and angles. The attitude estimation unit calculates the multi-degree-of-freedom pose parameters (position and attitude) of key components of the equipment through key point detection and spatial geometric calculation, providing basic data for precise control. The job recognition unit analyzes the motion mode and operation characteristics of the equipment, identifies different job types (such as excavation, hoisting, walking, etc.), and judges the job progress and completion status by combining time series analysis technology.
[0108] (3) Collaborative Control Module. The collaborative control module enables seamless integration and collaborative control with the construction equipment control system. This module includes an industrial interface unit, a fieldbus unit, a wireless communication unit, and a safety interlock unit. The industrial interface unit supports multiple industrial control protocols (such as Modbus and Profinet) and interface standards, achieving compatibility with equipment from different manufacturers (such as PLCs) through protocol conversion. The fieldbus unit uses standard fieldbus technologies (such as CAN and EtherCAT) for the transmission of equipment-level control signals, supporting real-time data exchange and remote parameter configuration to ensure reliable transmission and rapid response of control commands. The wireless communication unit uses technologies such as 5G or Wi-Fi 6 to support high-bandwidth, low-latency data transmission, enabling remote control and status monitoring of mobile equipment. The safety interlock unit establishes a multi-level safety protection mechanism, including software interlocks, hardware interlocks, and emergency stop functions, to prevent conflicts in multiple equipment operations and ensure rapid response in abnormal situations.
[0109] (4) Scheduling Decision Module. The scheduling decision module uses artificial intelligence algorithms to realize intelligent scheduling and optimization decisions for construction equipment, supporting cognitive decision-making based on the Visual-Language-Behavior (VLA) model. This module includes a job planning unit, a path optimization unit, a conflict resolution unit, and a load balancing unit. The job planning unit automatically generates equipment operation sequences based on Building Information Modeling (BIM) and construction schedule, considering process dependencies and resource constraints to optimize the construction process arrangement. The path optimization unit uses intelligent path planning algorithms to generate optimal movement trajectories for equipment, comprehensively considering factors such as path length, energy consumption, and obstacle avoidance, and supports dynamic path adjustment. The conflict resolution unit detects potential conflicts in multi-equipment operations in real time, including spatial conflicts, temporal conflicts, and resource conflicts, and automatically resolves conflicts and redistributes tasks through priority scheduling and negotiation mechanisms. The load balancing unit monitors the operating load status of each piece of equipment, and achieves equipment load balancing through task redistribution and scheduling optimization, thereby improving equipment utilization and extending equipment lifespan.
[0110] (5) Digital Twin Module. The digital twin module constructs a virtual mapping of the construction site, realizing real-time synchronization between the physical and digital worlds. This module includes a model integration unit, a real-time rendering unit, a physical simulation unit, a data mapping unit, and an interactive control unit. The model integration unit imports and manages BIM data, including building structures, equipment models, and site information, supports multiple data formats, and enables dynamic updates and version management of the model. The real-time rendering unit uses high-performance graphics rendering technology to create a three-dimensional visualization scene, supporting visual effects such as lighting, shadows, and materials, providing an immersive user experience. The physical simulation unit performs equipment dynamics simulation and collision detection based on a physics engine, simulating the motion behavior, load distribution, and structural stress of the equipment, and predicting potential risks. The data mapping unit establishes a two-way data channel between the physical entity and the digital model, realizing real-time synchronization of status information, supporting historical data playback and future trend prediction. The interactive control unit provides multiple human-computer interaction methods, including touch operation, gesture control, and voice commands, supports multi-user collaborative operation, and realizes remote monitoring and control functions.
[0111] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0112] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A remote multimodal intelligent feed control method for construction equipment, characterized in that, Includes the following steps: S1. Real-time collection of multi-source heterogeneous data streams containing visual information, operator interaction intent information, and equipment operating status information through a network of field sensing devices; S2. Based on deep learning algorithms, the collected visual information is analyzed in real time to realize the parsing from the original image to the high-level operation intention and form executable discretized operation instructions; S3. Convert discretized operation commands into specific control target parameters that can be executed by the equipment; S4. Encode the control target parameters into low-level command messages that conform to the industrial communication protocol and send them to the equipment controller for execution via the communication network; S5. Continuously monitor the deviation between the actual state and the expected state of the equipment, perform anomaly detection through residual calculation, and trigger alarms or safety policies when the threshold is exceeded.
2. The remote multimodal intelligent feed control method for construction equipment as described in claim 1, characterized in that, In step S1, an edge computing architecture is used for data preprocessing, and edge servers are deployed at the construction site to achieve data dimensionality reduction and feature extraction. For low-light scenes, an adaptive enhancement algorithm is used to improve image quality. Multispectral image fusion uses a weighted average method. ,in Visible light image, Infrared image, For the first The weight corresponding to the sub-risk; Weight By calculating image information entropy Adaptive adjustment ensures reliable sensing capabilities in all-weather operating environments.
3. The remote multimodal intelligent feed control method for construction equipment as described in claim 1, characterized in that, In step S2, the collected visual information is analyzed in real time based on deep learning algorithms to realize the parsing from the original image to the high-level operation intent, specifically as follows: Using object detection networks to analyze images The system locates key targets and extracts a set of structured key feature points. , This represents the total number of key feature points. Based on the key feature point set Calculate a feature vector that can characterize the operational intention D. ; ,in It is a mapping function used to calculate the geometric relationship between the relative positions, distances, or angles between feature points; The calculated feature vector Input a pre-trained classifier It is parsed into a predefined discretized operation instruction pattern. ; , in, It is a finite set of instructions.
4. The remote multimodal intelligent feed control method for construction equipment as described in claim 3, characterized in that, In step S3, the parsed high-level operational intent is converted into specific control target parameters that the equipment can execute, specifically: For multi-equipment scenarios, the optimization is extended to multi-target optimization; for single-equipment scenarios, the intent is directly mapped to the control target. Discretization operation instruction mode Transformed into specific control objectives that can be executed by the equipment. , In the formula, In order to convey the operational intent Mapping to control target The decision function.
5. The remote multimodal intelligent feed control method for construction equipment as described in claim 1, characterized in that, The interactive information includes gesture information, in step S2. In gesture recognition scenarios, a hand keypoint detection model is used to extract the coordinates of the thumb tip. and the coordinates of the tip of the index finger Calculate the direction difference The gesture type is determined based on the threshold.
6. The remote multimodal intelligent feed control method for construction equipment as described in claim 1, characterized in that, In step S5, the actual status of the equipment is continuously monitored. With the expected state To achieve precise control and anomaly monitoring of deviations, specifically including: Construct vectors to describe the operational status of equipment , which includes location ,attitude Other key operating parameters; The health status of system operation is assessed by calculating the norm of the state deviation, and the residual is defined. , when the residual Exceeding the preset dynamic or static threshold At that time, that is The system will determine this as an anomaly and trigger an alarm or corresponding security policy.
7. The remote multimodal intelligent feed control method for construction equipment as described in claim 6, characterized in that, In step S5, a cognitive feedback and autonomous deductive control method based on a large visual language behavior model is used as the actual state. The monitoring and feedback methods include the following sub-steps: S51. When the regular anomaly detection module detects the state deviation norm Exceeding the preset threshold At that time, visual information at the current moment Equipment state vector and deviation information The input is fed into a pre-trained visual language behavior model; the visual language behavior model parses the abnormal scene and outputs two items: a) a natural language text describing the nature of the abnormality. b) A measure of cognitive uncertainty This value quantifies the model's confidence in its understanding of the current scene; S52. Measurement of cognitive uncertainty With a preset threshold of cognitive uncertainty Compare: a) If This indicates that the model has a high degree of confidence in the current situation, and the system will autonomously enter sub-step S53; b) If This indicates that the model has a vague understanding of the current situation or has encountered a completely new context. The system will automatically trigger the human-machine collaboration mode and translate the natural language description. The system presents relevant visual information to remote operators and provides a dialogue interface, allowing operators to lead or assist the model in making decisions. S53. Enter autonomous decision-making mode; the system will send a high-level text instruction. The input is fed into the VLA model; the VLA model dynamically generates a sequence of sub-actions based on its understanding of the scene. The solution sequence is constructed, and the behavioral entropy of the sequence is calculated simultaneously. : , in, The model generates the first Size of movement The probability of; S54. Based on the behavioral entropy The value is used to apply a hierarchical verification strategy to the generated action sequence: a) If If the risk level is low, the physical equipment will be deployed directly after a rapid and secure simulation is performed in the digital twin environment. b) If If the simulation is in a high-risk zone, explicit authorization from a remote operator is required before execution.
8. The remote multimodal intelligent feed control method for construction equipment as described in claim 1, characterized in that, Also includes: S6. By transforming coordinates, the physical equipment status is mapped to the virtual 3D environment in real time, and the model rendering and performance index display are dynamically updated. Specifically: Use the WebGL framework to build 3D visualization scenes; Establish basic scene elements including ground grid, coordinate axes, and lighting; Load the 3D model of the equipment and set different initial poses and colors according to the equipment type; Equipment status data is pushed to the front end via a WebSocket server at a preset frequency; After receiving the data, the front end updates the rotation matrix and position coordinates of the 3D model in real time: for rotating parts, the rotation angle is accumulated according to the gesture state. ; For moving parts, the displacement is accumulated based on the gesture state. The equipment's operational status is displayed intuitively through color gradients and numerical labels.
9. The remote multimodal intelligent feed control method for construction equipment as described in claim 8, characterized in that, Also includes: S7. Establish a comprehensive risk assessment index that integrates multiple risk sources, and trigger tiered response strategies based on risk levels, up to and including emergency braking; specifically: Establish comprehensive risk indicators By using a weighted summation method, risk sources from different sources can be integrated: , in, Representing various sub-risks, Its corresponding weight; According to risk indicators The value and the preset multi-level threshold Comparisons are made to trigger different levels of response strategies.
10. A remote multimodal intelligent feed control system for construction equipment, characterized in that, include: The construction site perception module is responsible for the real-time acquisition and fusion processing of multi-source heterogeneous information, including a multi-source visual acquisition unit, an environmental monitoring unit, and a positioning and tracking unit. The equipment status recognition module uses computer vision and deep learning technologies to achieve intelligent recognition and status analysis of construction equipment; this module includes an equipment detection unit, an attitude estimation unit, and an operation recognition unit. The collaborative control module enables seamless integration and collaborative control with the construction equipment control system; this module includes an industrial interface unit, a fieldbus unit, a wireless communication unit, and a safety interlock unit. The scheduling decision module uses artificial intelligence algorithms to realize intelligent scheduling and optimization decisions for construction equipment, and supports cognitive decision-making based on a large visual language behavior model, including operation planning unit, path optimization unit, conflict resolution unit and load balancing unit. The digital twin module constructs a virtual mapping of the construction site, enabling real-time synchronization between the physical and digital worlds. It includes a model integration unit, a real-time rendering unit, a physical simulation unit, a data mapping unit, and an interactive control unit.