A Method and System for Understanding Lift Operation Scenarios by Integrating Multimodal Perception Models

By collecting and integrating multimodal perception information from the lifting operation scenario, the limitations of single-modal perception methods have been overcome, enabling a comprehensive understanding and dynamic adaptation to complex operation scenarios, thereby improving the safety and efficiency of operations.

CN121686121BActive Publication Date: 2026-05-26NANTONG INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANTONG INST OF TECH
Filing Date
2026-02-12
Publication Date
2026-05-26

Smart Images

  • Figure CN121686121B_ABST
    Figure CN121686121B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for understanding lift operation scenarios by integrating a multimodal perception model, relating to the field of artificial intelligence technology. First, it collects scene-triggered perception information such as visual, mechanical, and positional data of the lift operation scenario to form a trigger perception information set. Then, it performs cross-modal interactive mapping processing on the trigger perception information set to obtain interactive mapping perception information. Based on this, it performs scene adaptation calibration processing to obtain a scene-adapted perception model. The interactive mapping perception information is input into the scene-adapted perception model to perform multi-dimensional situational deduction and generate a panoramic situational description. Finally, based on the panoramic situational description, it generates dynamically adapted lift operation control commands and transmits them to the execution control unit, achieving accurate understanding and intelligent control of the lift operation scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method and system for understanding lift operation scenarios by integrating a multimodal perception model. Background Technology

[0002] In industrial production, lifting platforms serve as crucial material handling and operational equipment, operating in complex and ever-changing scenarios involving the dynamic interaction of various operational elements. Accurately understanding these scenarios is essential for ensuring operational safety, improving efficiency, and optimizing workflows. However, existing methods for understanding lifting platform operational scenarios have several limitations.

[0003] Traditional scene understanding methods often rely on single-modal perception information, such as acquiring image information of the work scene solely through visual sensors, or acquiring force information of the lift solely through mechanical sensors. These single-modal perception methods cannot comprehensively and accurately reflect the complex state of the work scene, easily overlooking key information and leading to biased understanding of the scene. For example, relying solely on visual information may not accurately determine the load weight borne by the lift, while relying solely on mechanical information makes it difficult to determine the specific location and shape of the load.

[0004] Furthermore, existing methods lack effective cross-modal interaction and integration mechanisms when processing perceptual information from different modalities. Perceptual information from different modalities is typically processed independently, without fully considering their inherent connections and mutual influences, thus failing to form a comprehensive and integrated understanding of the operational scenario. Moreover, existing perceptual models are often static, making it difficult to adapt to dynamic changes in operational scenarios and unable to adjust the model's response logic in a timely manner. This results in poor accuracy and adaptability when facing complex and ever-changing operational scenarios. Summary of the Invention

[0005] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for understanding lift operation scenarios by fusing a multimodal perception model, the method comprising:

[0006] Scene-triggered perception information of the lifting operation scenario is collected to form a set of trigger perception information. The scene-triggered perception information includes visual modal perception information, mechanical modal perception information, and position modal perception information. All perception information is directly triggered and generated by dynamic changes in the operation scenario.

[0007] Cross-modal interactive mapping processing is performed on the set of triggered sensing information. Different modal sensing information is compared, associated and integrated based on the real-time association of work scene elements to obtain interactive mapping sensing information.

[0008] Scene adaptation calibration is performed based on interactive mapping perception information to keep the response logic of the perception model synchronized with the dynamic changes of the work scene, thus obtaining a scene-adapted perception model.

[0009] The interactive mapping perception information is input into the scene adaptation perception model, which performs multi-dimensional situational inference processing to generate a panoramic situational description that includes the associated operation trajectory and potential change trend of each element in the operation scene.

[0010] Based on the panoramic situation description, dynamically adapted lift operation control commands are generated and transmitted to the lift's execution control unit.

[0011] In another aspect, embodiments of the present invention also provide a lifting operation scene understanding system that integrates a multimodal perception model, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or code. The processor is used to run the programs, instructions or code in the machine-readable storage medium to implement the above-described method.

[0012] Based on the above, this embodiment of the invention collects multimodal scene-triggered perception information, including visual, mechanical, and positional information, of the lifting operation scenario to form a trigger perception information set. Cross-modal interactive mapping processing is performed on this set, comparing, associating, and integrating different modal perception information based on the real-time correlation of operation scenario elements. This allows for a more comprehensive and accurate reflection of the complex state of the operation scenario. Scene adaptation calibration processing based on the interactive mapping perception information ensures that the response logic of the perception model remains synchronized with the dynamic changes of the operation scenario, enhancing the model's adaptability and real-time performance. This enables timely capture of changes in the operation scenario and accurate responses. The interactive mapping perception information is input into the scene adaptation perception model for multi-dimensional situational analysis, generating a panoramic situational description containing the associated operational trajectories and potential change trends of various elements in the operation scenario. Based on this panoramic situational description, dynamically adapted lifting operation control commands are generated, enabling timely adjustment of the lifting operation parameters according to real-time changes in the operation scenario, ensuring the safety and efficiency of lifting operations. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the execution flow of the lift operation scenario understanding method that integrates multimodal perception models provided in the embodiments of the present invention.

[0014] Figure 2 This is a schematic diagram of the hardware architecture of the lift operation scenario understanding system that integrates a multimodal perception model, provided in an embodiment of the present invention. Detailed Implementation

[0015] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a method for understanding lift operation scenarios by fusing a multimodal perception model, provided in one embodiment of the present invention. The following is a detailed description of this method for understanding lift operation scenarios by fusing a multimodal perception model.

[0016] Step S110: Collect scene-triggered perception information of the lifting operation scene to form a set of trigger perception information. The scene-triggered perception information includes visual modal perception information, mechanical modal perception information, and position modal perception information. All perception information is directly triggered and generated by dynamic changes in the operation scene.

[0017] In this embodiment, a scenario of a lift in an auto repair shop lifting a small car is used as an example. First, high-definition industrial cameras are installed on the four support arms of the lift. The camera's shooting angle covers the movement range of the lift's actuators, the overall shape of the object being lifted (i.e., the small car), and the environmental structure within a 2-meter radius of the lift. When the lift starts and the support arms contact the car's chassis and gradually lift, the cameras are triggered to begin acquiring image information, capturing one frame every 0.02 seconds to form visual modal perception information. Simultaneously, a three-dimensional force sensor is installed at the top of each support arm. When pressure is generated when the support arm contacts the car's chassis, the force sensor is triggered, collecting in real time the vertical pressure value, the horizontal force component value, and the coordinates of the force's point of application. The sampling frequency is 100 Hz, forming mechanical modal perception information. In addition, laser positioning sensors are installed at the base of the lift, the joints of the support arms, and near the four wheels of the car. When the lift's support arms begin to move or the car's position changes, the positioning sensors are triggered to collect the three-dimensional coordinates of various components of the lift and the three-dimensional coordinates of key parts of the car (such as chassis support points and wheel centers) in real time. The sampling frequency is 50 Hz, forming position modal perception information. The perception information of the above three modalities is timestamped during the acquisition process to ensure time synchronization, and is finally integrated to form a set of trigger perception information.

[0018] Step S120: Perform cross-modal interactive mapping processing on the trigger perception information set, compare, associate and integrate the different modal perception information based on the real-time association of the work scene elements to obtain interactive mapping perception information.

[0019] Step S121: Based on the lifting operation scenario, define the core element types, which include lifting execution components, operation load objects, and surrounding environmental structures, and extract operating characteristics and interaction boundary parameters for each type of core element.

[0020] In this embodiment, the lifting machine's execution components include four support arms, a hydraulic lifting column, and a control panel; the work-bearing object is a small car to be repaired, specifically including the car's chassis, wheels, and body; the surrounding environment includes a repair tool cabinet around the lifting machine, safety warning lines on the ground, and another lifting machine nearby. For the lifting machine's execution components, its operational characteristics are extracted, including the extension length, lifting height, and rotation angle of the support arms; the lifting speed and pressure value of the hydraulic lifting column; and the button status of the control panel. Interactive boundary parameters include the maximum extension range and minimum retracted length of the support arms; and the maximum lifting height and minimum ground clearance of the hydraulic lifting column. For the work-bearing object, operational characteristics include the overall weight of the car, chassis height, and wheel spacing; interactive boundary parameters include the range of support point locations on the car chassis, and the maximum width and length of the car body. For the surrounding environment, operational characteristics include the location coordinates of the repair tool cabinet and the range dimensions of the safety warning lines; interactive boundary parameters include the minimum safe distance between the repair tool cabinet and the lifting machine, and the boundary of the area enclosed by the safety warning lines.

[0021] Step S122: Extract scene representation details of visual modal perception information from the trigger perception information set to form a visual modal representation set. The scene representation details include the morphological change features, color distribution features, and motion trajectory features of the core elements.

[0022] Step S1221: Perform frame segmentation on the visual modal perception information in the trigger perception information set to obtain a continuous visual frame sequence, where each visual frame contains a complete picture of the work scene.

[0023] In this embodiment, the acquired visual modal perception information is arranged in timestamp order. Since the camera's sampling interval is 0.02 seconds, 50 frames can be obtained per second. Frame-by-frame extraction is performed on these images, removing blurred frames caused by lighting changes or sensor jitter, retaining only clear visual frames, ultimately forming a continuous sequence of visual frames. Each visual frame has a resolution of 1920×1080 pixels and contains a complete view of the lift's actuators, the object being carried (a small car), and the surrounding environmental structure.

[0024] Step S1222: For each visual frame, extract the contour morphology parameters of the lifting machine's actuators, calculate the contour morphology difference value of the actuators in adjacent visual frames, and determine the extension, contraction, and rotation morphological changes of the actuators based on the difference value to form morphological change features.

[0025] In this embodiment, taking the support arm of a lift as an example, for each visual frame, the contour of the support arm is extracted through an edge detection algorithm. The contour shape parameters of the support arm include the coordinates of each feature point on the contour, the perimeter of the contour, the area, etc. Calculate the coordinate differences of the corresponding feature points of the support arm contour in two adjacent frames to obtain the difference values. If the difference value is positive and gradually increasing in the horizontal direction, it is determined that the support arm is extending; if the difference value is negative and the absolute value is gradually increasing in the horizontal direction, it is determined that the support arm is contracting; if the angle of the feature point around the support arm joint point changes, it is determined that the support arm is rotating. For example, in the visual frames from the 100th frame to the 150th frame, the horizontal coordinate of a certain feature point of the support arm changes from x1 to x2 (x2 > x1), and the vertical coordinate remains basically unchanged. It is calculated that the difference value in the horizontal direction gradually increases, thereby determining that the support arm is in the extended state, forming an extended shape change feature; in the 200th frame to the 250th frame, the horizontal coordinate of this feature point changes from x3 to x4 (x4 < x3), determining that the support arm is in the contracted state, forming a contracted shape change feature; in the 300th frame to the 350th frame, the angle of this feature point around the joint point changes from θ1 to θ2, determining that the support arm is rotating, forming a rotational shape change feature.

[0026] Step S1223: Analyze the color distribution of the core elements in each visual frame, extract the main color tone information and color uniformity information of each core element, compare the color changes of the same core element in adjacent visual frames, record the changes in color depth and brightness, and form color distribution features.

[0027] In this embodiment, for a small sedan as the operation bearing object, in each visual frame, the image is converted from the RGB color space to the HSV color space through color space conversion, and then color clustering is performed on the sedan area, and the H (hue), S (saturation), and V (brightness) values corresponding to the clustering center are extracted as the main color tone information. The color uniformity information is represented by calculating the variance of the HSV values of each pixel point in the sedan area and the main color tone HSV values. The smaller the variance, the higher the color uniformity. Compare the main color tone V values of the sedan in adjacent visual frames. If the V value increases, it means that the color of the sedan becomes brighter; if the V value decreases, it means that the color becomes darker. For example, in the 50th frame to the 100th frame, due to the change in the angle of the workshop lights, the main color tone V value of the sedan body increases from v1 to v2 (v2 > v1), which is recorded as the color becoming brighter; in the 150th frame to the 200th frame, the V value decreases from v3 to v4 (v4 < v3), which is recorded as the color becoming darker. At the same time, calculate the variance value of the color uniformity of the sedan in each frame. If the variance value changes little in adjacent frames, it means that the color distribution is relatively stable, otherwise it means that the color distribution has changed.

[0028] Step S1224: In a continuous visual frame sequence, mark the key feature points of each core element, track the position changes of the key feature points in different visual frames, connect the key feature points in each frame to form the motion trajectory line segments of the core element, and extract information on the direction, curvature, and extension speed of the trajectory line segments to form motion trajectory features.

[0029] In this embodiment, the center of the top of the hydraulic lifting column of the hoist is marked as a key feature point. In a continuous sequence of visual frames, the coordinates (x, y) of this key feature point in each frame are determined using a feature point matching algorithm. The coordinates of the key feature points in adjacent frames are connected sequentially to form a motion trajectory segment. The direction of the trajectory segment is determined by calculating the angle between the segment and the horizontal direction; an angle of 0 degrees indicates horizontal to the right, 90 degrees indicates vertical upward, and so on. The degree of curvature is represented by calculating the curvature of the trajectory segment; the greater the curvature, the more curved the trajectory. The extension speed is obtained by dividing the distance between the key feature points in two adjacent frames by the time interval (0.02 seconds). For example, in frames 100 to 200, the coordinates of the key feature point at the top of the hydraulic lifting column change from (x100, y100) to (x200, y200). The trajectory line segment formed by connecting these coordinates has an angle of 85 degrees with the horizontal direction, indicating that the direction is close to vertically upward. The curvature of this trajectory line segment is calculated to be 0.01 / pixel, indicating a small degree of bending. The extension speed is calculated to be between 10-15 cm / s through the coordinates of each adjacent frame, forming a motion trajectory feature.

[0030] Step S1225: Perform correlation annotation on the morphological change features, color distribution features, and motion trajectory features corresponding to each visual frame, and define the correspondence between the morphological change features, color distribution features, and motion trajectory features of the same core element.

[0031] In this embodiment, for the core element of the lift support arm, in the 100th visual frame, its morphological change characteristic is extension, its color distribution characteristics are main hues H1, S1, and V1, its color uniformity variance is d1, and its motion trajectory characteristics are direction θ1, bending degree c1, and extension speed v1. By adding the same core element identifier (e.g., "support arm A") and timestamp (e.g., "frame 100") to these features, associative annotation is achieved, indicating that these features all belong to the visual features of support arm A in frame 100. Similarly, similar annotations are applied to the features of other core elements in each visual frame to ensure that different visual features of the same core element at the same time can correspond to each other.

[0032] Step S1226: Integrate the scene representation details of all visual frames in chronological order to form a visual feature sequence of each core element at different time points. Classify the visual feature sequences of all core elements according to the core element type to construct a visual feature classification set.

[0033] In this embodiment, the morphological change features, color distribution features, and motion trajectory features of support arm A in all visual frames are arranged in time-stamp order to form a visual feature sequence of support arm A. Similarly, visual feature sequences of core elements such as support arms B, C, and D, hydraulic lifting column, car, and repair tool cabinet are formed respectively. Then, the visual feature sequences of all support arms are classified into the "lifting machine actuator" category, the visual feature sequences of the car are classified into the "work load object" category, and the visual feature sequences of the repair tool cabinet, etc., are classified into the "surrounding environment structure" category, thus constructing a visual feature classification set.

[0034] Step S1227: Perform redundancy removal on the feature data in the visual feature classification set, add visual modality labels and core element labels to the redundancy-removed feature data, and integrate all labeled feature data to form a visual modality representation set.

[0035] In this embodiment, for the visual feature sequence of support arm A under the "lifting machine execution component" category in the visual feature classification set, it is checked whether there are consecutive frames of feature data that are completely identical. If so, the first frame of data is retained, and subsequent redundant data is deleted. For example, if the shape change features, color distribution features, and motion trajectory features of support arm A are the same in frames 200-210, then only the feature data of frame 200 is retained. After redundancy removal, a visual modality identifier (such as "visual-001") and a core element identifier (such as "support arm A") are added to each feature data. Then, all the identified feature data are integrated together to form a visual modality representation set.

[0036] Step S123: Extract the force transmission characteristics of the mechanical modal perception information from the trigger perception information set to form a mechanical modal characterization set. The force transmission characteristics include the force distribution characteristics, reaction force feedback characteristics, and force transmission attenuation characteristics between core elements.

[0037] Step S1231: Perform signal analysis on the mechanical modal sensing information in the trigger sensing information set, and extract the force value change data, force direction data, and force application point data from the original mechanical signal.

[0038] In this embodiment, the mechanical modal sensing information comes from a three-dimensional force sensor installed at the top of the lifting arm support. The original mechanical signal is an analog signal, which is converted into a digital signal by an A / D converter. The digital signal is filtered to remove noise interference, and then the force value change data is extracted, that is, the sequence of force values ​​of the support arm in the x, y, and z directions over time. The force direction data is obtained by calculating the resultant force direction of the force values ​​in the three directions, represented by a direction vector (dx, dy, dz). The force application point data is obtained through the positioning module built into the force sensor, which is a three-dimensional coordinate (px, py, pz). For example, at a certain moment, the force value change data of the support arm A is f_z(t) in the z direction (vertical direction), f_x(t) in the x direction (horizontal transverse direction), and f_y(t) in the y direction (horizontal longitudinal direction); the force direction vector is (dx1, dy1, dz1); and the force application point coordinates are (px1, py1, pz1).

[0039] Step S1232: Based on the coordinates of the point of application of the force and the spatial coordinates of each core element, determine the core element and the position of application corresponding to each force value, establish the force transmission path between different core elements, and mark the direction of force transmission.

[0040] In this embodiment, the spatial coordinate ranges of the lift support arm and the car chassis are known. The coordinates of the force's point of application (px, py, pz) are compared with the spatial coordinates of each core element. If the coordinates of the point of application fall within the spatial coordinate range of the car chassis, then the core element corresponding to that force value is determined to be the car chassis, and the point of application is a specific area of ​​the chassis (e.g., the left front support point). In this case, the force transmission path is lift support arm → car chassis. Based on the force's direction vector, the direction of force transmission is determined to be from the support arm to the car chassis. For example, if the coordinates of the force's point of application (px1, py1, pz1) of support arm A fall within the spatial coordinate range of the left front support point of the car chassis, then the core element corresponding to that force value is determined to be the car chassis, the point of application is the left front support point, the force transmission path is support arm A → left front support point of the car chassis, and the transmission direction is from support arm A to that support point.

[0041] Step S1233: Analyze the force value change data on each transmission path, extract the force value distribution on different core elements, including the magnitude distribution of the force value, the shape characteristics of the distribution area, and the uniformity of the distribution, to form the force distribution characteristics.

[0042] In this embodiment, taking the transmission path from support arm A to the left front support point of the car chassis as an example, the force value change data f_z(t) (vertical force value) is analyzed. The distribution of force values ​​is represented by the frequency of occurrence of different force value intervals over a period of time. For example, the frequency of force values ​​occurring in the 500-600 N interval is 30%, and the frequency of occurrence in the 600-700 N interval is 50%. The shape characteristics of the distribution area are determined by the distribution of force application points on the car chassis. If multiple force application points form a circular area, the shape characteristic is circular; if they form a rectangular area, the shape characteristic is rectangular. The uniformity of distribution is measured by calculating the variance of the force application point coordinates. The smaller the variance, the more uniform the distribution. For example, the small variance of the force application point coordinates of support arm A within 10 seconds indicates that the force is relatively evenly distributed in the left front support point area of ​​the car chassis, forming a force distribution characteristic.

[0043] Steps S1234: Extract the force value data, direction data, and point of application data of the reaction force for each action force, analyze the correspondence between the reaction force and the action force, including the proportional relationship of the force values, the degree of oppositeity of the action directions, and the corresponding position of the point of application, and form the reaction force feedback characteristics.

[0044] In this embodiment, when support arm A applies a force F to the left front support point of the car chassis, the car chassis will generate a reaction force F' on support arm A. Based on the data collected by the force sensor, the force values ​​f'_x(t), f'_y(t), and f'_z(t), the direction vectors (dx', dy', dz'), and the coordinates of the point of application (px', py', pz') of the reaction force F' are extracted. The proportional relationship between the force values ​​is analyzed, and the ratio of F' to F is calculated. Ideally, this ratio should be 1, but in practice, it may deviate slightly due to measurement errors. The degree of opposition in the directions of action is determined by calculating the angle between the direction vectors of the action and reaction forces; when the angle is 180 degrees, the directions are completely opposite. The corresponding positions of the points of application are determined by comparing the coordinate differences between the points of application of the action and reaction forces; they should be basically coincident. For example, the force F exerted by the support arm A on the car chassis is 650 N, and the force F' of the reaction force is 648 N, with a ratio close to 1; the direction vector of the force is (0, 0, 1) (vertically upward), and the direction vector of the reaction force is (0, 0, -1) (vertically downward), with an angle of 180 degrees; the coordinates of the points of application are basically the same, thus forming the reaction force feedback characteristic.

[0045] Step S1235: Track the transmission process of force along the transmission path, compare the force value changes on different core elements, calculate the attenuation of force during transmission, analyze the relationship between attenuation and transmission distance and material properties of core elements, extract the transmission attenuation law of force, and form the transmission attenuation characteristics of force.

[0046] In this embodiment, the transmission path is taken as follows: hydraulic system of the lift → hydraulic lifting column → support arm → car chassis. Force value F1 is measured at the output end of the hydraulic system, force value F2 is measured at the top of the hydraulic lifting column, force value F3 is measured at the end of the support arm, and force value F4 is measured at the car chassis. The force attenuation ΔF1 = F1 - F2, ΔF2 = F2 - F3, ΔF3 = F3 - F4. The transmission distances are L1 from the hydraulic system to the top of the lifting column, L2 from the top of the lifting column to the end of the support arm, and L3 from the end of the support arm to the car chassis. The core material properties include the elastic modulus E1 of the steel in the hydraulic lifting column and the elastic modulus E2 of the aluminum alloy in the support arm. Analysis shows that the attenuation ΔF is approximately proportional to the transmission distance L and approximately inversely proportional to the elastic modulus E, i.e., ΔF = k × L / E (k is a proportionality coefficient). From this, the force transmission attenuation law is extracted, forming the force transmission attenuation characteristic.

[0047] Step S1236: Identify the transmission path for each force distribution feature, reaction force feedback feature, and force transmission attenuation feature, and define the force transmission path to which the feature belongs.

[0048] In this embodiment, the force distribution characteristics, reaction force feedback characteristics, and force transmission attenuation characteristics along the transmission path from support arm A to the left front support point of the car chassis are marked with the identifier "path-A-left front", indicating that these characteristics belong to the force transmission path from support arm A to the left front support point of the car chassis. Similarly, corresponding transmission path identifiers are added to the characteristics of other transmission paths.

[0049] Step S1237: Classify the mechanical features according to the interaction relationship between the core elements, and group the force distribution features, reaction force feedback features, and force transmission attenuation features that belong to the same interaction relationship into one group.

[0050] In this embodiment, the interaction relationships between core elements include the support interaction between the support arm and the car chassis, and the connection interaction between the hydraulic lifting column and the support arm. The force distribution characteristics, reaction force feedback characteristics, and force transmission attenuation characteristics belonging to the support interaction between support arm A and the left front support point of the car chassis are grouped together and named "Support Arm A - Left Front Support Interaction Group"; the relevant characteristics belonging to the connection interaction between the hydraulic lifting column and support arm A are grouped into "Hydraulic Lifting Column - Support Arm A Connection Interaction Group", etc.

[0051] Step S1238: Sort each set of mechanical features in time sequence, organize the feature data according to the force transmission time order, form a mechanical feature sequence for each interaction relationship, add modal identifier and interaction relationship identifier to each mechanical feature, and define the perception modality to which the feature belongs and the corresponding core element interaction relationship.

[0052] In this embodiment, for the "Support Arm A - Front Left Support Interaction Group of the Car", the force distribution characteristics, reaction force feedback characteristics, and force transmission attenuation characteristics within the group are arranged in time stamp order to form a mechanical feature sequence. A modal identifier "Mechanical-001" and an interaction relationship identifier "Support Arm A - Front Left Left Support of the Car" are added to each mechanical feature. Then, all identified mechanical feature sequences are integrated to form a mechanical modal characterization set.

[0053] Step S1239: Integrate all labeled mechanical feature sequences to form a set of mechanical modal representations.

[0054] Step S124: Extract the spatial occupancy features of the position modality perception information in the trigger perception information set to form a position modality representation set. The spatial occupancy features include the coordinate distribution features, relative distance features, and spatial posture features of the core elements.

[0055] In this embodiment, the position modality perception information comes from a laser positioning sensor. For the support arm in the lift's actuator, the laser positioning sensor collects the coordinates of multiple feature points in three-dimensional space, such as the coordinates of the two ends and joints of the support arm. The set of these coordinates forms a coordinate distribution feature. The shortest distance between the support arm and the car chassis, and the shortest distance between the support arm and the surrounding repair tool cabinet are calculated to form a relative distance feature. The spatial attitude angles (pitch angle, yaw angle, and roll angle) of the support arm are calculated using the coordinates of three non-collinear feature points on the support arm to form a spatial attitude feature. Similarly, the corresponding coordinate distribution features, relative distance features, and spatial attitude features are extracted from the car and the surrounding environmental structure, and integrated to form a position modality representation set.

[0056] Step S125: Establish cross-modal interaction association rules. The content of the cross-modal interaction association rules is that features describing the same core element in different modal representation sets must corroborate each other, and features describing the interaction of different core elements must be related to each other.

[0057] In this embodiment, the cross-modal interaction association rules specifically stipulate that: for the same core element, such as a car chassis, the morphological change features (such as chassis height changes) in its visual modal representation set should corroborate the coordinate distribution features (such as z-axis coordinate changes) in its position modal representation set. If the chassis height is visually observed to increase, the z-axis coordinate in the position modality should increase. To describe the interaction of different core elements, such as the interaction between the support arm and the car chassis, the contact state features between the support arm and the chassis in the visual modality should be correlated with the force distribution features of the support arm on the chassis in the mechanical modality. If the support arm is visually in contact with the chassis, the corresponding force should be detected in the mechanical modality.

[0058] Step S126: Based on the cross-modal interaction association rules, compare and associate the features of the same core element in the visual modal representation set and the mechanical modal representation set, and use the visual morphological change features to explain the cause of the mechanical force distribution features.

[0059] In this embodiment, taking the car chassis as a core element, the morphological change features of the car chassis in the visual modal representation set show that during frames 100-200, the chassis underwent upward displacement, and the contact area between the chassis and the support arm gradually expanded. The force distribution features in the mechanical modal representation set corresponding to the same time period show that the force distribution area of ​​the support arm on the chassis gradually expanded, and the force value gradually increased. Through comparison and correlation, the morphological change feature of the expanded contact area between the chassis and the support arm is used to explain the reason for the expansion of the force distribution area and the increase in force value in mechanics. That is, due to the increased contact area, the force distribution range expands, and at the same time, in order to support the chassis's upward movement, the force value also increases accordingly.

[0060] Step S127: Connect and associate the features of the interaction between the mechanical modal representation set and the position modal representation set corresponding to different core elements, and use the force effect transmission feature to supplement the dynamic change basis of the spatial occupancy feature.

[0061] In this embodiment, the force transmission characteristics of the support arm to the car chassis in the mechanical modal characterization set show that the vertical force exerted by the support arm on the chassis gradually increases during frames 200-300. The spatial occupancy characteristics of the car chassis in the position modal characterization set show that the z-axis coordinate (height) of the chassis gradually increases during the same time period. By connecting and relating these two features, the increase in the vertical force in the force transmission characteristics is used to supplement the explanation of the dynamic change in chassis height increase in the spatial occupancy characteristics. That is, the increase in the vertical force exerted by the support arm on the chassis overcomes the weight of the car, causing the chassis height to rise.

[0062] Step S128: Complementarily associate the spatial pose features of the corresponding core elements in the position modality representation set and the visual modality representation set, and use the spatial occupancy features to improve the spatial dimension information of the visual morphological change features.

[0063] In this embodiment, the spatial occupancy features of the lift support arm in the position modality representation set include the spatial attitude angles of the support arm (pitch angle α, yaw angle β, roll angle γ). The morphological change features of the support arm in the visual modality representation set show that the support arm has rotated, but it is difficult to accurately determine the angle and direction of rotation from visual perspective alone. Through complementary correlation, the spatial attitude angle data in the position modality is used to improve the spatial dimension information of the visual morphological change features, clarifying that the rotation of the support arm is a change of α degrees in the pitch angle around the x-axis, a change of β degrees in the yaw angle around the y-axis, and a change of γ degrees in the roll angle around the z-axis.

[0064] Step S129: Integrate all associated cross-modal features to form an associated feature set containing mutual explanation logic and supplementary evidence. Based on the interaction strength of each modal feature in the associated feature set, organize them hierarchically according to the core element type and interaction relationship to construct a cross-modal interaction mapping framework. Each feature node in the cross-modal interaction mapping framework is accompanied by explanation information and supplementary data of the associated modality, and finally obtains the interaction mapping perception information.

[0065] In this embodiment, visual, mechanical, and positional modal features, after comparison, association, and complementary association, are integrated to form an associated feature set. Each feature in the associated feature set contains mutual interpretation logic and supplementary evidence with other modal features. For example, the feature of increased car chassis height is explained by morphological changes in the visual modality and force transmission in the mechanical modality, with supplementary evidence from coordinate data in the positional modality. Then, based on the interaction strength of each modal feature (such as the closeness of the association between visual and mechanical features), the associated feature set is hierarchically organized according to core element type (lifting machine actuator, work-bearing object, surrounding environmental structure) and interaction relationship (support interaction, connection interaction, etc.). When constructing the cross-modal interaction mapping framework, each feature node (such as the force distribution feature node of support arm A) is accompanied by explanatory information of the associated modality (such as the contact morphological changes between support arm A and the car chassis in the visual modality) and supplementary data (such as the coordinate change data of support arm A in the positional modality). Finally, this cross-modal interaction mapping framework and all the feature nodes and association information it contains together constitute the interactive mapping perception information.

[0066] Step S130: Perform scene adaptation calibration processing based on interactive mapping perception information to keep the response logic of the perception model synchronized with the dynamic changes of the work scene, and obtain a scene-adapted perception model.

[0067] Step S131: Analyze the initial response logic of the preset perception model, and deduce the processing priority, feature extraction path, and correlation analysis method of each modality perception information in the initial response logic.

[0068] In this embodiment, the preset perception model is a pre-trained deep learning model for understanding the lifting operation scenario. The initial response logic is analyzed by examining the model's network structure and parameter configuration file. The processing priority for each modality of perception information is as follows: visual modality is prioritized over mechanical modality, and mechanical modality is prioritized over positional modality. Regarding the feature extraction path, visual modality information first passes through convolutional layers to extract low-level features, then through pooling layers for dimensionality reduction, and finally inputs into fully connected layers to extract high-level features; mechanical modality information is directly input into fully connected layers for feature extraction; positional modality information undergoes coordinate transformation and is then input into a recurrent neural network to extract temporal features. The association analysis method uses simple feature concatenation followed by input into a classifier for decision-making.

[0069] Step S132: Extract scene dynamic change features from the interactive mapping perception information to form a dynamic change feature set. The scene dynamic change features include the operation state transformation features of core elements, the interaction relationship adjustment features, and the environmental impact change features.

[0070] Step S1321: Analyze the cross-modal association data in the interactive mapping perception information and extract the current running state information of each core element. The current running state information includes motion speed, force magnitude, spatial position, and morphological state.

[0071] In this embodiment, the cross-modal association data in the interactive mapping perception information includes the features of each core element in different modalities and their relationships. For the lift support arm A, its motion speed is extracted from the visual modal features (calculated by the position changes of feature points in adjacent frames), its force magnitude (the vertical force value on the support arm) is extracted from the mechanical modal features, its spatial position (three-dimensional coordinates) is extracted from the position modal features, and its morphological state (extended, contracted, or rotated) is extracted from the visual modal features. For example, the current operating state information of support arm A is: motion speed v=[vx, vy, vz], force magnitude f=[fx, fy, fz], spatial position p=[px, py, pz], and morphological state is extended. Similarly, the current operating state information of core elements such as the car, other support arms, and surrounding environmental structures is extracted.

[0072] Step S1322: Compare the current operating status information of the same core element at different time points, identify the type of operating status transformation, record the starting conditions, transformation duration, and post-transformation status characteristics of the transformation process, and form operating status transformation characteristics.

[0073] In this embodiment, taking the car's operating state as an example, the current operating state information at the 100th second and the 200th second is compared. At the 100th second, the car's speed is 0 (stationary), its spatial position is the initial position of the lift, and its shape is horizontal. At the 200th second, the car's speed is v=[0, 0, 0.1] (moving vertically upwards), its z-coordinate increases, and its shape remains horizontal. The car's operating state transition type is identified as a transition from a stationary state to a vertically upward movement state. The initial condition is that the force exerted by the lift support arm on the car reaches the car's weight. The transition time is from the 150th second to the 200th second (50 seconds). The characteristics of the transition state are a speed of 0.1 m / s, a continuously increasing z-coordinate, and a horizontal shape. This forms the operating state transition characteristics of the car.

[0074] Step S1323: Analyze the cross-modal correlation strength and correlation type between different core elements, track the changes in correlation, record the adjustment process of the establishment, enhancement, weakening and termination of correlation, extract the key triggering factors and manifestations in the adjustment process, and form interactive relationship adjustment characteristics.

[0075] In this embodiment, the cross-modal correlation strength (measured by the number and closeness of correlations between the features of the two in the correlation feature set) and correlation type (support interaction) of support arm A and car chassis are analyzed. Before the lifting operation begins, support arm A and car chassis are not correlated; when support arm A contacts the chassis, a correlation is established, and the correlation strength gradually increases; as the force exerted by the support arm on the chassis increases, the correlation strength strengthens; when the lifting is completed, the correlation strength remains stable; when the lift descends and the support arm separates from the chassis, the correlation terminates. Key triggering factors during the adjustment process are recorded: support arm contacts the chassis (establishment), force increases (strengthens), force stabilizes (stable), and support arm leaves the chassis (termination). This is represented by a curve showing the change in correlation strength values. This forms the interaction relationship adjustment feature between support arm A and car chassis.

[0076] Step S1324: Extract cross-modal feature data related to the working environment structure from the interactive mapping perception information, calculate the occlusion range and relative distance changes of the environmental structure on the lifting machine's execution components and the working load objects, and form environmental impact change features.

[0077] In this embodiment, the maintenance tool cabinet and the support arm of the lifting mechanism's actuator in the surrounding environment have a potential impact. The coordinate distribution features of the maintenance tool cabinet and the support arm are extracted from the interactive mapping perception information. The occlusion range of the maintenance tool cabinet on the support arm (the visual area of ​​the support arm obstructed by the tool cabinet when it moves to a certain position) and the change in the relative distance between the support arm and the tool cabinet (the change in distance between them as the support arm extends or retracts) are calculated. For example, when the support arm extends to its maximum length, the relative distance to the tool cabinet decreases to 1 meter. At this point, the tool cabinet occludes a portion of the visual area of ​​the support arm, specifically 10% of the area at the end of the support arm. This forms the environmental impact variation characteristics.

[0078] Step S1325: Mark the time sequence of the operating state transition features, interaction relationship adjustment features, and environmental impact change features, and define the time nodes and duration periods corresponding to each type of feature.

[0079] In this embodiment, time node markers are added to the operating state transition characteristics of the car, with a start time node of 150 seconds, an end time node of 200 seconds, and a duration of 50 seconds. Time node markers are added to the interaction relationship adjustment characteristics between support arm A and the car chassis, with an establishment time node of 140 seconds, an enhancement start time node of 150 seconds, a stabilization time node of 200 seconds, and an end time node of 500 seconds, with durations of 10 seconds, 50 seconds, 300 seconds, etc., for each stage. Time node markers are added to the environmental impact change characteristics of the maintenance tool cabinet on the support arm, with an occlusion start time node of 300 seconds, an end time node of 350 seconds, and a duration of 50 seconds.

[0080] Step S1326: Analyze the interaction relationships among the operational state transition features, interaction relationship adjustment features, and environmental impact change features; identify the triggering effect of operational state transition features on interaction relationship adjustment features and the impact of environmental impact change features on operational state transition features; and establish causal relationship links among features.

[0081] In this embodiment, analysis revealed that the transition of the car's operating state (from stationary to upward movement) triggered an enhancement phase in the interaction relationship adjustment characteristics between the support arm A and the car chassis; that is, the start of the car's movement was the cause of the enhanced interaction relationship. Simultaneously, the occlusion of the support arm by the tool cabinet in the environmental impact change characteristics affected the extraction of the support arm's visual modal features, thus impacting the accuracy of the support arm's operating state transition judgment. In other words, the environmental impact change characteristics negatively affected the operating state transition characteristics. Based on these interactions, a causal link was established: increased force on the support arm → car operating state transition (stationary → upward movement) → interaction relationship adjustment (enhancement); tool cabinet occlusion → abnormal visual feature extraction of the support arm → delay in operating state transition judgment.

[0082] Step S1327: Classify and organize all dynamic change features according to core element type and feature type, and construct a dynamic change feature classification framework.

[0083] In this embodiment, dynamic change characteristics are categorized into three types based on core element type: lifting machine execution components, work-bearing objects, and surrounding environment structures. Under the lifting machine execution components category, characteristics are further divided into operating state transition characteristics (e.g., support arm extension state transition), interaction relationship adjustment characteristics (e.g., interaction adjustment between the support arm and the hydraulic lifting column), and environmental impact change characteristics (e.g., the obstruction effect of the tool cabinet on the support arm). Similarly, the work-bearing objects and surrounding environment structures categories are further subdivided to construct a dynamic change characteristic classification framework.

[0084] Step S1328: Add details to the classified dynamic change features, improve the key information in the dynamic feature description, so that the features can accurately reflect the essence of scene changes.

[0085] In this embodiment, for the operating state transition feature of the car, key information such as acceleration changes and the smoothness of speed changes during the transition process is added, so that the feature not only includes the type and time of the state transition, but also reflects the dynamic characteristics of the transition process. For the interaction relationship adjustment feature between the support arm and the car chassis, information such as the specific numerical range of the correlation strength and the force value change threshold at different stages is added to improve the feature description.

[0086] Step S1329: Add an associated modality identifier to each dynamically changing feature to define which modal interaction mapping information the feature data originates from.

[0087] In this embodiment, the data for the car's operational state transition features are derived from the motion trajectory features of the visual modality and the coordinate distribution features of the position modality; therefore, an associated modality identifier "visual + position" is added. The data for the interaction relationship adjustment features between the support arm and the car chassis are derived from the contact morphology features of the visual modality and the force distribution features of the mechanical modality; an associated modality identifier "visual + mechanical" is added. The data for the environmental impact changes of the maintenance tool cabinet on the support arm are derived from the occlusion features of the visual modality and the relative distance features of the position modality; an associated modality identifier "visual + position" is added.

[0088] Step S13210: Integrate all the classified and organized dynamic change features, causal relationship links, and related modal identifiers to form a dynamic change feature set.

[0089] In this embodiment, the operational state transition features, interaction relationship adjustment features, environmental impact change features, and established causal relationship links, which have been categorized, supplemented with details, and marked with associated modal identifiers, are integrated together to form a dynamic change feature set. This dynamic change feature set comprehensively reflects the dynamic changes of core elements and their interrelationships in the lift operation scenario.

[0090] Step S133: Based on the dynamic change feature set, extract the actual response output of the initial response logic under the driving of the scene dynamic change features, compare the actual response output with the actual scene state reflected by the dynamic change feature set, and identify the response deviation of the initial response logic in terms of processing priority, feature extraction path, and correlation analysis method.

[0091] In this embodiment, a set of dynamic change features is input into a preset perception model to obtain the model's actual response output, such as the recognition result of the car's operating state and the judgment result of the interaction relationship between the support arm and the chassis. The above actual response output is compared with the actual scene state reflected by the set of dynamic change features. For example, the set of dynamic change features shows that the car has started to move upward at 150 seconds, but the preset perception model only recognizes this state at 170 seconds, which is a delay of 20 seconds. This indicates that the initial response logic may have a deviation in processing priority and the processing of position modal information is not timely enough. The model's judgment on the enhanced interaction relationship between the support arm and the chassis relies solely on mechanical features and ignores the supplementary role of visual features, indicating a deviation in the correlation analysis method. The model failed to effectively extract the environmental impact change feature of the maintenance tool cabinet obstruction, indicating a deviation in the feature extraction path and the absence of an extraction channel for environmental impact features.

[0092] Step S134: To address the response bias in processing priorities, based on the frequency of operational state transitions of each core element in the dynamically changing feature set, the processing order of different modal sensing information is reallocated, so that the sensing information corresponding to elements whose operational state transitions occur continuously receives higher priority processing resources.

[0093] In this embodiment, the dynamic change feature set shows that the object being lifted (the car) has the highest frequency of state transitions during the lifting process (e.g., from stationary to moving, speed changes, etc.), followed by the lifting machine's actuators, and then the surrounding environmental structure. The initial response logic's prioritization of visual modal over mechanical modal processing leads to a delay in responding to the car's motion state. Therefore, processing priorities are reallocated: position modal perception information (directly reflecting spatial position and motion state) is given the highest priority, followed by mechanical modal, and then visual modal. This way, the position modal information corresponding to the car, which experiences frequent state transitions, receives more priority processing resources, thereby accelerating the response speed to its state changes.

[0094] Step S135: To address the response deviation in the feature extraction path, adjust the feature extraction nodes of the preset perception model according to the manifestation of features in the dynamically changing feature set, add extraction channels that adapt to new dynamic features, and delete extraction links that are irrelevant to the current dynamic changes.

[0095] Step S1351: Analyze the original feature extraction path of the preset perception model, and define the feature type, extraction algorithm, and output format corresponding to each extraction node in the original feature extraction path.

[0096] In this embodiment, the original feature extraction path of the preset perception model includes visual modality feature extraction nodes, mechanical modality feature extraction nodes, and positional modality feature extraction nodes. The feature types corresponding to the visual modality extraction nodes are morphological change features and color distribution features, the extraction algorithm is a convolutional neural network, and the output format is a feature vector; the feature type corresponding to the mechanical modality extraction nodes is force magnitude features, the extraction algorithm is a fully connected network, and the output format is a feature vector; the feature type corresponding to the positional modality extraction nodes is coordinate features, the extraction algorithm is a recurrent neural network, and the output format is a feature vector.

[0097] Step S1352: Compare the feature types output by each extraction node in the original feature extraction path with the feature types contained in the dynamically changing feature set, and mark the extraction nodes that cannot output any feature type in the dynamically changing feature set as invalid nodes.

[0098] In this embodiment, the dynamic change feature set includes environmental impact change features (such as occlusion range), but there is no corresponding extraction node in the original feature extraction path that can output this feature type. Therefore, the visual modality extraction sub-node that is unrelated to environmental impact change features and only extracts color distribution features is marked as an invalid node.

[0099] Step S1353: Analyze the manifestation of the new dynamic features, including the data type, structural form, and change pattern of the new dynamic features, and determine the appropriate extraction algorithm type and extraction parameter range.

[0100] In this embodiment, the new dynamic feature is the occlusion range feature in the environmental influence change features. Its data type is a two-dimensional region coordinate set, its structural shape is an irregular polygon, and its change pattern is dynamic change with the movement of the support arm. The adapted extraction algorithm type is a region extraction algorithm based on image segmentation. The extraction parameter range includes a segmentation threshold (set according to the grayscale difference between the occluded region and the background) and a region area threshold (to filter out noise regions that are too small).

[0101] Step S1354: Based on the adapted extraction algorithm type and extraction parameter range, add a new extraction node to the original feature extraction path. The position of the new extraction node corresponds to the stage where new dynamic features are generated, so that the new extraction node can accurately capture new dynamic features.

[0102] In this embodiment, an occlusion range extraction node is added after the original morphological change feature extraction node in the visual modality feature extraction path. This node employs a region extraction algorithm based on image segmentation, setting the segmentation threshold to t1 and the region area threshold to s1. The location of the new extraction node corresponds to the stage where the support arm and the surrounding environmental structure may occlude each other. By performing image segmentation on the visual frame, the coordinate set of the occlusion region is extracted, thereby accurately capturing the new dynamic feature of the occlusion range.

[0103] Step S1355: Build a data transmission channel for the newly added extraction node so that the feature data extracted by the new extraction node can be smoothly transmitted to the subsequent correlation analysis module. The transmission format of the data transmission channel is consistent with the input format of the subsequent module.

[0104] In this embodiment, the feature data output by the newly added occlusion range extraction node is a set of two-dimensional region coordinates, while the input format of the subsequent association analysis module is a feature vector. Therefore, when constructing the data transmission channel, the two-dimensional region coordinate set is converted into a fixed-length feature vector (by calculating parameters such as the area, perimeter, and center coordinates of the region) to ensure that the transmission format is consistent with the input format of the subsequent module, so that the occlusion range feature data can be smoothly transmitted to the association analysis module.

[0105] Step S1356: Identify extraction links in the original feature extraction path where the extracted feature types are not included in the dynamically changing feature set.

[0106] In this embodiment, there is a link in the original visual modality feature extraction path that extracts the texture features of the car body, but the dynamic change feature set does not contain texture features, so this extraction link is irrelevant to the current dynamic change.

[0107] Step S1357: Disconnect the irrelevant extraction link from other modules, delete all extraction nodes and data transmission channels in the irrelevant extraction link, and adjust the parameter settings of the original extraction nodes so that the extraction algorithm of the original extraction nodes can better adapt to the corresponding feature representation in the dynamically changing feature set.

[0108] In this embodiment, the link for extracting the car body texture features is severed from the subsequent modules, and the texture feature extraction node and corresponding data transmission channel in this link are deleted. For the retained visual modality morphological change feature extraction node, the size and number of convolutional neural network kernels are adjusted to enable it to better extract the dynamic change features such as the extension, contraction, and rotation of the support arm.

[0109] Step S1358: Reconfigure the data flow channel between the new extraction node and the original retained node, set the output order of each extraction node, and input the extracted feature data into the correlation analysis module.

[0110] In this embodiment, the newly added occlusion range extraction node is connected in parallel with the original retained morphological change feature extraction node and motion trajectory feature extraction node. Their output data is aggregated through the newly configured data flow channel and then input into the correlation analysis module in the order of morphological change feature, motion trajectory feature, and occlusion range feature, ensuring that the correlation analysis module can process each feature data in sequence.

[0111] Step S1359: By inputting dynamic data from the interactive mapping perception information, test the adjusted feature extraction path, verify whether all new dynamic features can be effectively extracted, and whether the original effective features can be obtained normally, so that the adjusted feature extraction path can fully adapt to the dynamically changing feature set.

[0112] In this embodiment, dynamic data of interactive mapping perception information including occlusion conditions is input into the adjusted feature extraction path. Test results show that the occlusion range features are effectively extracted, and the set of region coordinates accurately reflects the actual occlusion situation. Existing effective features such as morphological change features and motion trajectory features can also be obtained normally, and the extraction accuracy is improved compared to before the adjustment. This confirms that the adjusted feature extraction path is fully adapted to the dynamically changing feature set.

[0113] Step S136: To address the response bias in the association analysis method, based on the type of interaction relationship adjustment in the dynamically changing feature set, optimize the association analysis algorithm logic of the preset perception model to enhance the ability to identify and analyze new interaction relationships.

[0114] In this embodiment, the types of interaction relationships adjusted in the dynamically changing feature set include novel interaction relationships such as contact interaction between the support arm and the car chassis, connection interaction between the support arm and the hydraulic lifting column, and interaction relationships influenced by the surrounding environmental structure on the lift. The original correlation analysis algorithm logic only uses simple feature splicing, which cannot effectively identify these complex interaction relationships. Therefore, the correlation analysis algorithm logic is optimized by introducing a graph neural network model, treating each core element as a node and the interaction relationship as an edge. Different types of interaction relationships are identified and analyzed by learning node features and edge weights. For example, for the contact interaction between the support arm and the car chassis, the graph neural network enhances its ability to identify the contact state by learning the correlation weights between the positional and mechanical features of both.

[0115] Step S137: Extract the dynamic change features of the interactive mapping perception information at different time periods, and set the dynamic adjustment cycle of the response logic based on the degree of difference, so that the response logic can be updated in real time with the scene changes.

[0116] In this embodiment, the interactive mapping perception information is divided into multiple time periods, such as 0-100 seconds, 100-200 seconds, 200-300 seconds, etc. The degree of difference in dynamic change characteristics within each time period is calculated, and the degree of difference is measured by calculating the Euclidean distance between the feature data of adjacent time periods. If the Euclidean distance is large, it indicates that the scene changes drastically, and the dynamic adjustment cycle should be shortened; if the Euclidean distance is small, it indicates that the scene changes gently, and the dynamic adjustment cycle can be appropriately extended. For example, in the lift start-up phase (0-100 seconds), the degree of difference in dynamic change characteristics is large, and the dynamic adjustment cycle is set to 5 seconds; in the lift stabilization phase (200-300 seconds), the degree of difference is small, and the dynamic adjustment cycle is set to 15 seconds.

[0117] Step S138: Integrate the adjusted processing priority, feature extraction path, correlation analysis method and dynamic adjustment cycle to form a set of scene adaptation parameters.

[0118] In this embodiment, the reallocated processing priorities (position modality > mechanical modality > visual modality), the adjusted feature extraction paths (including newly added occlusion range extraction nodes and deleted irrelevant links), the optimized correlation analysis algorithm logic (graph neural network model), and the set dynamic adjustment period (5 seconds, 15 seconds, etc.) are integrated to form a scene adaptation parameter set. The parameters in this scene adaptation parameter set will be used to update the response logic of the preset perception model.

[0119] Step S139: Write the scene adaptation parameter set into the core processing module of the preset perception model, replace the original fixed parameters, and complete the adjustment of the perception model response logic.

[0120] In this embodiment, the core processing module of the preset perception model includes a parameter configuration file and an algorithm logic unit. Processing priority parameters, feature extraction path parameters, correlation analysis algorithm parameters, and dynamic adjustment cycle parameters from the scene adaptation parameter set are written into the parameter configuration file, overriding the original fixed parameters. Simultaneously, the optimized correlation analysis algorithm logic (graph neural network model) is integrated into the algorithm logic unit, replacing the original simple feature concatenation algorithm. Through this method, the adjustment of the perception model's response logic is completed.

[0121] Step S1310: Drive the adjusted perception model to run by interactively mapping the continuous dynamic data in the perception information, so that the adjusted perception model can further adapt to scene changes in actual operation, and finally obtain the scene-adapted perception model.

[0122] In this embodiment, continuous dynamic data (including visual, mechanical, and positional modal features across multiple time periods) from the interactive mapping perception information is input into the adjusted perception model to drive its operation. During model operation, the model's response output is periodically checked against the actual scene state according to a set dynamic adjustment cycle (e.g., 5 seconds). If a deviation exists, it is automatically fine-tuned according to the rules in the scene adaptation parameter set. After multiple dynamic adjustment cycles of operation and adaptation, the model can accurately respond to various dynamic changes in the work scene, ultimately resulting in a scene-adapted perception model.

[0123] Step S140: Input the interactive mapping perception information into the scene adaptation perception model, and the scene adaptation perception model performs multi-dimensional situational inference processing to generate a panoramic situational description that includes the associated running trajectories and potential change trends of each element in the work scene.

[0124] Step S141: Divide the interactive mapping perception information according to the core element type and timestamp to obtain the cross-modal feature data sequence of each core element at different time nodes.

[0125] In this embodiment, the interactive mapping perception information includes cross-modal feature data of core elements such as the lifting machine's execution components, the work-bearing object, and the surrounding environmental structure. These core element types are categorized into lifting machine execution component data, work-bearing object data, and surrounding environmental structure data. Under each core element type, corresponding cross-modal feature data is extracted in time-stamp order (e.g., one time node per second) to form a feature data sequence. For example, the cross-modal feature data sequence of support arm A in the lifting machine's execution components includes visual morphological changes, mechanical force distribution characteristics, and position coordinate distribution characteristics at time nodes t1, t2, t3, ...

[0126] Step S142: Through the element trajectory restoration module of the scene adaptation perception model, the cross-modal feature data sequence of each core element is integrated, and the actual running path of the core element is restored based on the motion trajectory features of the visual modality and the coordinate distribution features of the position modality.

[0127] In this embodiment, the element trajectory reconstruction module of the scene adaptation perception model receives the cross-modal feature data sequence of support arm A. The module first extracts the motion trajectory features of the visual modality (the sequence of position changes of key feature points) and the coordinate distribution features of the position modality (a three-dimensional coordinate sequence), and then performs time synchronization and spatial registration on these two features. The position data of the two modalities are fused using a fusion algorithm (such as Kalman filtering) to eliminate noise and errors, obtaining the continuous position points of support arm A in three-dimensional space. These continuous position points are connected in chronological order to reconstruct the actual running path of support arm A. This actual running path is a smooth spatial curve containing the complete motion trajectory of support arm A from its initial position to its extended position and then to its retracted position.

[0128] Step S143: Using the interaction relationship deduction module of the scene adaptation perception model, analyze the intersection points and time overlap segments of the operation paths of different core elements, and combine the force effect transmission characteristics of the mechanical mode to reconstruct the actual interaction process between core elements.

[0129] Step S1431: Obtain the actual running path data of each core element. The actual running path data includes the coordinate information, timestamp information, and morphological feature information of each point on the path and the corresponding visual frame.

[0130] In this embodiment, the actual running path data of support arm A, support arm B, and the car chassis are obtained from the element trajectory reconstruction module. The running path data of support arm A includes the three-dimensional coordinates (x_a1, y_a1, z_a1), (x_a2, y_a2, z_a2), ... of each point on the path, the corresponding timestamps t_a1, t_a2, ..., and the morphological feature information of support arm A in the visual frame corresponding to each timestamp (such as extension length and rotation angle). Similarly, the running path data of support arm B and the car chassis are obtained.

[0131] Step S1432: Import the actual running path data of all core elements into the interaction relationship inference module of the scene adaptation perception model, and identify the spatial intersection points between the running paths of different core elements through the path analysis algorithm in the interaction relationship inference module.

[0132] In this embodiment, the path analysis algorithm in the interaction relationship deduction module uses the spatial geometric intersection method to calculate the intersection of the running path curve of support arm A and the running path curve of the car chassis. When the distance between the two curves in space is less than a preset threshold (e.g., 5 cm), the point is determined to be a spatial intersection point. After calculation, the spatial intersection point P1 (x_p1, y_p1, z_p1) between support arm A and the car chassis, and the spatial intersection point P2 (x_p2, y_p2, z_p2) between support arm B and the car chassis are identified.

[0133] Step S1433: Extract the timestamp information corresponding to each spatial intersection point, compare the timestamps of different core elements arriving at the spatial intersection point, and determine the time overlap segment. The time overlap segment is the time period during which different core elements are simultaneously at the spatial intersection point.

[0134] In this embodiment, the timestamps t_a5-t_a10 (from the 5th second to the 10th second) of support arm A corresponding to the spatial intersection point P1 are extracted, as are the timestamps t_c5-t_c10 of the car chassis arriving at P1. A comparison reveals that the timestamp ranges are the same, therefore the time overlap segment is determined to be from the 5th second to the 10th second. Similarly, the time overlap segment between support arm B and the car chassis at the spatial intersection point P2 is determined to be from the 6th second to the 11th second.

[0135] Step S1434: Extract the mechanical modal perception information corresponding to the time overlap segment from the interactive mapping perception information, and analyze the force effect transmission characteristic data in the mechanical modal perception information, paying attention to the changes in the force distribution characteristics and reaction force feedback characteristics.

[0136] In this embodiment, the mechanical modal sensing information of support arm A during the overlapping time segment from the 5th to the 10th second is extracted, and the force distribution characteristic data (the vertical force value gradually increases from 0 to F_max, and the distribution area expands from point contact to surface contact) and reaction force feedback characteristic data (the magnitude of the reaction force is basically equal to the magnitude of the force, but the direction is opposite) are obtained. Similarly, the mechanical modal sensing information of support arm B during the overlapping time segment from the 6th to the 11th second is extracted, and its force transmission characteristic data are analyzed.

[0137] Step S1435: Combining the coordinates of the spatial intersection point with the duration of the time overlap segment, compare the mechanical modal force effect transmission characteristic data with the preset threshold to determine whether the contact mode between the core elements is direct or indirect contact.

[0138] In this embodiment, the preset force threshold for direct contact is F_th. In the force distribution characteristic data of support arm A during the time overlap segment, the vertical force value increases from 0 to F_max, where F_max > F_th, and the coordinates of the spatial intersection point P1 are consistent with the actual position coordinates of support arm A and the car chassis. The duration of the time overlap segment is 5 seconds (relatively long). Therefore, it is determined that the contact mode between support arm A and car chassis is direct contact.

[0139] Step S1436: Based on the changes in the contact method and force transmission characteristics, reconstruct the force generation process between the core elements, including the start time, peak time, and decay time of the force.

[0140] In this embodiment, the support arm A is in direct contact with the car chassis. In the force transmission characteristics, the vertical force value gradually increases from the 5th second (starting moment), reaches its maximum value F_max (peak moment) at the 7th second, and then remains stable until it begins to gradually decrease at the 10th second (decay moment). This reconstructs the force generation process of the support arm A on the car chassis: starting at the 5th second, reaching its peak at the 7th second, and beginning to decay at the 10th second.

[0141] Step S1437: Combine the morphological change features in the visual modal perception information, observe the morphological changes of the core elements in the time overlap segment, verify the rationality of the force generation process, and supplement the details of morphological changes in the interaction process.

[0142] In this embodiment, visual modal perception information shows that during the overlapping time segment from the 5th to the 10th second, the extension length of support arm A gradually increases, the contact area with the car chassis gradually expands, and the shape of the car chassis changes from slightly bending downwards from horizontal to gradually returning to horizontal. This is consistent with the generation process of force increasing from 0 to F_max and then remaining stable in the mechanical modality, verifying the rationality of the force generation process. Simultaneously, details of the morphological changes during the interaction process are supplemented: the rubber pad of support arm A undergoes elastic deformation under contact pressure, and a slight indentation appears in the support point area of ​​the car chassis.

[0143] Step S1438: Extract the positional modal spatial occupancy feature changes of the core elements during the interaction process, and restore the details of the relative position adjustment and spatial posture changes of the core elements during the interaction process.

[0144] In this embodiment, the position modal spatial occupancy feature shows that during the interaction process, the spatial attitude angle (pitch angle) of the support arm A is adjusted from α1 to α2 (α2>α1) to better fit the tilt angle of the car chassis; the spatial position z coordinate of the car chassis gradually rises from z1 to z2, indicating that under the force of the support arm A, the car chassis is gradually lifted. The details of the above relative position adjustment and spatial attitude change further enrich the reconstruction of the interaction process.

[0145] Step S1439: Integrate spatial intersection point information, time overlap segment information, force effect transmission characteristic changes, morphological change details, and spatial position adjustment information, and deduce the interaction process between core elements in chronological order.

[0146] In this embodiment, the above information is integrated in chronological order to deduce the interaction process between support arm A and the car chassis: At the 5th second, support arm A and the car chassis begin to contact at the spatial intersection point P1, and force begins to be generated. The extension length of the support arm increases, the contact area expands, and the z-coordinate of the chassis begins to rise; At the 7th second, the force reaches its peak value F_max, the elastic deformation of the support arm rubber pad is at its maximum, and the chassis shape slightly bends before stabilizing; From the 5th to the 10th second, the pitch angle of the support arm is adjusted, and the chassis continues to rise; At the 10th second, the force begins to decay, and the extension of the support arm stops.

[0147] Step S14310: Transform the derived interaction flow into a standardized interaction process description, including interaction start conditions, interaction methods, state changes during the interaction process, and interaction end markers, to fully restore the actual interaction process between core elements.

[0148] In this embodiment, the actual interaction process between support arm A and the car chassis is described as follows: the interaction starts when support arm A moves to the spatial intersection point P1, at the 5th second; the interaction method is direct contact support; the state changes during the interaction process include the support arm force increasing from 0 to F_max and then decreasing, the support arm extension length increasing, the contact area expanding, the rubber pad elastically deforming, the car chassis z-coordinate rising, the shape slightly bending and then stabilizing, and the support arm pitch angle adjusting; the interaction ends when the force begins to decrease, at the 10th second.

[0149] Step S144: Extract the change pattern of the running path of each core element, and combine it with the dynamic change characteristics in the interactive mapping perception information to deduce the possible subsequent running direction and state transformation trend of the core element.

[0150] In this embodiment, the running path change pattern of support arm A was extracted, revealing that its running path over the past 30 seconds exhibited a trend of first extending and then stabilizing. Combined with the running state transition characteristics of support arm A in the dynamic change feature set (the extended state has lasted for 20 seconds, approaching the maximum extended length), the possible subsequent running direction of support arm A is deduced to be stopping extension and maintaining the current length, with the state transition trend being from the extended state to the stable state. For the car chassis, its running path change pattern shows a continuous increase in the z-coordinate, and the dynamic change features indicate that its speed remains stable, suggesting that the subsequent running direction is to continue upward movement, with the state transition trend being from accelerated upward movement to uniform upward movement.

[0151] Step S145: Analyze the mutual influence of the potential operating trends of different core elements, and based on cross-modal interaction association rules, determine the new interactive relationships and effects that may be formed between elements in the future.

[0152] For example, step S1451: Obtain potential operating trend data for each core element, which includes possible future operating directions, operating speeds, and state transition nodes.

[0153] In this embodiment, the potential operating trend data for support arm A is: the operating direction is stationary, the operating speed is 0, and the state transition node is at second 150 (from the extended state to the stable state). The potential operating trend data for the car chassis is: the operating direction is vertically upward, the operating speed is 0.1 m / s, and the state transition node is at second 200 (from uniform ascent to stationary ascent). The potential operating trend data for the maintenance tool cabinet in the surrounding environmental structure is: the operating direction is stationary, the operating speed is 0, and there is no state transition node.

[0154] Step S1452: Based on the potential operational trend data of each core element, calculate its predicted position and predicted status at different time points in the future.

[0155] In this embodiment, the predicted position of the car chassis at a future time node t=160 seconds is calculated: the current z-coordinate is z_current, the running speed is 0.1 m / s, and the time interval from the current time t_current to t=160 seconds is Δt=160-t_current. Therefore, the predicted z-coordinate is z_current+0.1×Δt, and the predicted state is uniform upward movement. The predicted position of support arm A at t=160 seconds is its current extension length, and the predicted state is a stable state.

[0156] Step S1453: Based on cross-modal interaction association rules, set the judgment conditions for the formation of new interaction relationships between elements. The judgment conditions include spatial location coincidence conditions, state adaptation conditions, and force effect transmission possibility conditions.

[0157] In this embodiment, the spatial position coincidence condition is that the distance between the predicted positions of the two core elements is less than a preset distance threshold d_th; the state adaptation condition is that the predicted states of the two core elements can support interactive behavior (such as the support arm being in a stable state and the car being in an upward state); the force transmission possibility condition is that there is a force transmission path between the two core elements (such as direct contact or indirect contact through other objects).

[0158] Step S1454: Compare the predicted positions of different core elements at each future time point one by one, check whether the spatial position coincidence condition is met, and mark the time point and core element combination that meet the spatial position coincidence condition.

[0159] In this embodiment, the predicted positions of the maintenance lights above the car chassis and the surrounding environment are compared. At a future time node t=250 seconds, the distance between the predicted z-coordinate of the car chassis plus the vehicle height and the z-coordinate of the maintenance lights is less than d_th, which satisfies the condition of spatial position coincidence. This time node is marked as the combination of core elements (car chassis, maintenance lights).

[0160] Step S1455: For the combination of core elements marked, analyze whether the predicted state at the corresponding time node meets the state adaptation condition. If the combination of core elements meets the conditions of spatial location coincidence and state adaptation, further analyze whether there is a possibility of force effect transmission. Combine the force effect transmission characteristics of mechanical modes to determine whether there is a basis for forming action and reaction forces.

[0161] In this embodiment, the predicted state of the core elements (car chassis, maintenance lights) at t=250 seconds is: the car chassis is in a stopped rising state, and the maintenance lights are in a stationary state, satisfying the state adaptation condition (collision interaction may occur between stationary states). Further analysis shows that both the car chassis and the maintenance lights are solid objects; if they come into contact, there is a possibility of force transmission. Combining this with the force transmission characteristics of mechanical modes (solid contact generates elastic force), it is determined that the conditions for forming action and reaction forces are met.

[0162] Step S1456: For the combination of core elements that meet all the judgment conditions, determine the possible new interaction relationship types, which are contact interaction, non-contact influence, and force transmission interaction.

[0163] In this embodiment, the combination of core elements (car chassis, maintenance lights) satisfies the conditions of spatial position overlap, state adaptation and force transmission possibility, and the two will make direct contact at the predicted position. Therefore, the possible new interaction relationship type is determined to be contact interaction (collision).

[0164] Step S1457: Based on the potential operational trend data of the new interaction relationship type and core elements, deduce the effect after the interaction relationship occurs, including the change in the operational status of the core elements, the specific manifestation of force transmission, and the adjustment of subsequent trends.

[0165] In this embodiment, the new interaction relationship type is contact interaction (collision). Based on this, the effect is deduced as follows: the operating state of the car chassis will change from stopping to being obstructed, which may cause slight vibration; the maintenance lights are subjected to collision force, which may cause positional displacement or damage; the specific manifestation of force transmission is that the car chassis applies an impact force to the maintenance lights, and the maintenance lights generate a reaction force; the subsequent trend is that the car chassis can no longer rise, and the lifting operation may be forced to stop.

[0166] Step S1458: Sort all possible new interaction relationships and their effects according to the predicted time nodes, and generate a set of predicted future interaction relationships between elements.

[0167] In this embodiment, in addition to the contact-based interaction (car chassis, maintenance lights) (predicted time node t=250 seconds), there may be new interaction relationships between other core elements, such as the contact-based interaction between the support arm and the ground (if the support arm is overextended). A set of predicted future interaction relationships between elements is generated by sorting them according to the order of the predicted time nodes. This set of predicted future interaction relationships includes information such as the type of each new interaction relationship, the predicted time node, and its effect.

[0168] Step S146: Integrate the actual operation path, actual interaction process, potential operation trend, and future interaction relationship of all core elements to construct a multi-dimensional situational model of the lift operation scenario.

[0169] In this embodiment, the actual operating paths (spatial curves), actual interaction processes (such as the support interaction between the support arm and the chassis), potential operating trends (such as the car continuing to rise), and future interaction relationships (such as collisions with maintenance lights) of core elements such as the support arm, car chassis, and surrounding environmental structures are integrated into a three-dimensional spatiotemporal model. The model's time dimension covers from the start of the lifting operation to the predicted t=300 seconds, and the spatial dimension includes the three-dimensional coordinate range of the operation scenario. Each core element exists in the model as a dynamic entity, and its movement and interaction processes are visualized.

[0170] Step S147: Sort and interpolate the operation paths in the multidimensional situation model according to timestamps, so that the operation status and interaction process of each element are presented on a unified time axis; prioritize and classify the potential trends and possible interaction relationships of each element in the multidimensional situation model according to preset rules; transform the sorted operation paths, interaction processes, potential trends and future interaction relationships into text descriptions and graphical representations. The text descriptions include details of element operation, descriptions of interaction processes, and trend prediction content. The graphical representations include path diagrams, interaction relationship diagrams, and trend prediction diagrams.

[0171] In this embodiment, the operational paths of all core elements in the multi-dimensional situational model are sorted by timestamp from smallest to largest. Linear interpolation is performed on the paths between timestamps to make the path curves smoother, ensuring that the operational status and interaction process of each element are continuously presented on a unified time axis. For potential trends and possible interaction relationships, priority is ranked according to the urgency of the interaction (e.g., collision interaction has higher priority than non-impact interaction) and the scope of influence (interactions affecting the entire operation process have higher priority than local interactions), dividing them into high, medium, and low levels. Then, the sorted operational paths are converted into text descriptions, such as "Support arm A extends from its initial position to its maximum length from t=0-50 seconds, with the operational path being curve L1"; the interaction process is converted into text descriptions, such as "Support arm A has a direct contact support interaction with the car chassis from t=5-10 seconds, with the force increasing from 0 to F_max"; potential trends and future interaction relationships are converted into text descriptions, such as "The car chassis is expected to stop rising at t=200 seconds, posing a risk of collision with the maintenance lights." At the same time, draw path diagrams (showing the spatial trajectory of each element), interaction relationship diagrams (using directed edges to represent the interaction type and direction between elements), and trend prediction diagrams (using curves to represent the future state changes of elements).

[0172] Step S148: Integrate standardized text descriptions and graphical representations to supplement implicit situational information that is not directly presented in the multi-dimensional situational model but is derived through cross-modal feature correlation, forming a panoramic situational description that reflects the current state and future changes of the operational scenario.

[0173] In this embodiment, the aforementioned textual description and graphical representation are integrated to form a preliminary panoramic situational description. Then, implicit situational information is added. This information is not directly presented in the multi-dimensional situational model but is derived through cross-modal feature correlation. For example, based on the mechanical and positional characteristics of the support arm, the load condition of the lift's hydraulic system is derived; based on the morphological and positional characteristics of the car chassis, the changes in the car's center of gravity distribution are derived. Adding this implicit situational information to the panoramic situational description ultimately forms a comprehensive panoramic situational description that reflects the current state (position, state, and interaction of each element) and future changes (potential trends, possible new interactions, and effects) of the lift's operational scenario.

[0174] Step S150: Generate dynamically adapted lift operation control commands based on the panoramic situation description, and transmit the lift operation control commands to the lift's execution control unit.

[0175] In this embodiment, the panoramic situational awareness description indicates that the car chassis may collide with the maintenance lights at t=250 seconds. Based on this information, a dynamically adapted lift operation control command is generated: at t=200 seconds, the lift is controlled to stop rising to avoid a collision. The control command includes specific execution time, control parameters (such as adjusting the hydraulic system pressure value to 0), and information on the actuator (hydraulic lifting column). This control command is transmitted to the lift's execution control unit via an industrial bus. The execution control unit controls the lift to stop operation according to the command, ensuring operational safety.

[0176] Figure 2 The following is a schematic diagram of the hardware structure of a multimodal perception model-integrated lift operation scene understanding system 100 provided in an embodiment of the present invention for implementing the above-described multimodal perception model-integrated lift operation scene understanding method. Figure 2 As shown, the lift operation scene understanding system 100 that integrates multimodal perception models may include a processor 110, a machine-readable storage medium 120, a bus 130, and a communication unit 140.

[0177] Machine-readable storage medium 120 can store data and / or instructions. In some embodiments, machine-readable storage medium 120 can store data acquired from an external terminal. In some embodiments, machine-readable storage medium 120 can store data and / or instructions used by the lift operation scene understanding system 100 with fused multimodal perception model to execute or use in order to complete the exemplary methods described in this invention. In a specific implementation, one or more processors 110 execute the computer-executable instructions stored in machine-readable storage medium 120, enabling processor 110 to execute the lift operation scene understanding method with fused multimodal perception model as described in the above method embodiments. Processor 110, machine-readable storage medium 120, and communication unit 140 are connected via bus 130, and processor 110 can be used to control the transmission and reception actions of communication unit 140. The specific implementation process of processor 110 can be found in the various method embodiments executed by the lift operation scene understanding system 100 with fused multimodal perception model described above, and their implementation principles and technical effects are similar, so they will not be repeated here.

[0178] Furthermore, this embodiment of the invention also provides a readable storage medium containing computer-executable instructions. When the processor executes the computer-executable instructions, the above-mentioned method for understanding the lifting operation scenario by fusing a multimodal perception model is realized.

[0179] It should be noted that, in order to simplify the description of this invention and thus aid in the understanding of one or more embodiments, the foregoing description of the embodiments of this invention sometimes combines multiple features into a single embodiment, drawing, or description thereof. Similarly, it should be noted that, in order to simplify the description of this invention and thus aid in the understanding of one or more embodiments, the foregoing description of the embodiments of this invention sometimes combines multiple features into a single embodiment, drawing, or description thereof.

Claims

1. A method for understanding lift operation scenarios by integrating a multimodal perception model, characterized in that, The method includes: Scene-triggered perception information of the lifting operation scenario is collected to form a set of trigger perception information. The scene-triggered perception information includes visual modal perception information, mechanical modal perception information, and position modal perception information. All perception information is directly triggered and generated by dynamic changes in the operation scenario. Cross-modal interactive mapping processing is performed on the set of triggered sensing information. Different modal sensing information is compared, associated and integrated based on the real-time association of work scene elements to obtain interactive mapping sensing information. Scene adaptation calibration is performed based on interactive mapping perception information to keep the response logic of the perception model synchronized with the dynamic changes of the work scene, thus obtaining a scene-adapted perception model. The interactive mapping perception information is input into the scene adaptation perception model, which performs multi-dimensional situational inference processing to generate a panoramic situational description that includes the associated operation trajectory and potential change trend of each element in the operation scene. Based on the panoramic situation description, dynamically adapted lift operation control commands are generated and transmitted to the lift's execution control unit. The process of performing cross-modal interaction mapping on the trigger perception information set compares, associates, and integrates different modal perception information based on the real-time correlation of work scenario elements to obtain interaction mapping perception information, including: Based on the lifting machine operation scenario, core element types are defined, including lifting machine execution components, work-bearing objects, and surrounding environmental structures. Operational characteristics and interaction boundary parameters are extracted for each type of core element. Scene representation details of visual modal perception information are extracted from the trigger perception information set to form a visual modal representation set. The scene representation details include the morphological change features, color distribution features, and motion trajectory features of core elements. Extract the force effect transmission characteristics of the mechanical modal sensing information from the trigger sensing information set to form a mechanical modal characterization set. The force effect transmission characteristics include the force distribution characteristics, reaction force feedback characteristics, and force transmission attenuation characteristics between core elements. Spatial occupancy features of position modality perception information are extracted from the trigger perception information set to form a position modality representation set. The spatial occupancy features include the coordinate distribution features, relative distance features, and spatial posture features of the core elements. Establish cross-modal interaction association rules. The content of the cross-modal interaction association rules is that features describing the same core element in different modal representation sets must corroborate each other, and features describing the interaction of different core elements must be correlated with each other. Based on the cross-modal interaction association rules, the features of the same core element in the visual modal representation set and the mechanical modal representation set are compared and associated, and the visual morphological change features are used to explain the cause of the mechanical force distribution features. The characteristics of the interaction between different core elements in the mechanical modal characterization set and the position modal characterization set are linked and associated, and the dynamic change basis of the spatial occupancy characteristics is supplemented by the force effect transmission characteristics. The spatial posture features of the corresponding core elements in the location modality representation set and the visual modality representation set are complementary and correlated, and the spatial occupancy features are used to improve the spatial dimension information of the visual morphological change features. By integrating all the associated cross-modal features, a set of associated features containing mutual explanatory logic and supplementary evidence is formed. Based on the interaction strength of each modal feature in the set of associated features, the features are hierarchically organized according to the core element type and interaction relationship to construct a cross-modal interaction mapping framework. Each feature node in the cross-modal interaction mapping framework is accompanied by explanatory information and supplementary data of the associated modality, and finally the interaction mapping perception information is obtained.

2. The method for understanding lift operation scenarios by fusing multimodal perception models according to claim 1, characterized in that, The scene adaptation calibration process based on interactive mapping perception information ensures that the response logic of the perception model remains synchronized with the dynamic changes of the work scene, resulting in a scene-adapted perception model, including: The initial response logic of the preset perception model is analyzed, and the processing priority, feature extraction path, and correlation analysis method of each modality perception information in the initial response logic are derived. Dynamic change features of the scene are extracted from the interactive mapping perception information to form a dynamic change feature set. The dynamic change features of the scene include the operation state transformation features of core elements, the interaction relationship adjustment features, and the environmental impact change features. Based on the dynamic change feature set, the actual response output of the initial response logic under the driving of the dynamic change features of the scene is extracted. The actual response output is compared with the actual state of the scene reflected by the dynamic change feature set to identify the response deviation of the initial response logic in terms of processing priority, feature extraction path, and correlation analysis method. To address the response bias in processing priorities, the processing order of different modal sensing information is reallocated based on the frequency of operational state transitions of each core element in the dynamically changing feature set, so that the sensing information corresponding to elements whose operational state transitions occur continuously receives higher priority processing resources. To address the response bias in the feature extraction path, the feature extraction nodes of the preset perception model are adjusted according to the manifestation of features in the dynamically changing feature set. Extraction channels that adapt to new dynamic features are added, and extraction links that are irrelevant to the current dynamic changes are removed. To address the response bias in association analysis methods, the association analysis algorithm logic of the preset perception model is optimized based on the type of interaction relationship adjustment in the dynamically changing feature set, thereby enhancing the ability to identify and analyze new interaction relationships. Extract the dynamic change features of the interactive mapping perception information at different time periods, and set the dynamic adjustment cycle of the response logic based on the degree of difference, so that the response logic can be updated in real time with the scene changes. The adjusted processing priority, feature extraction path, correlation analysis method and dynamic adjustment cycle are integrated to form a set of scenario adaptation parameters. Write the scene adaptation parameter set into the core processing module of the preset perception model, replace the original fixed parameters, and complete the adjustment of the perception model's response logic. By driving the adjusted perception model through continuous dynamic data in the interactive mapping perception information, the adjusted perception model can be further adapted to scene changes in actual operation, and finally a scene-adapted perception model is obtained.

3. The method for understanding lift operation scenarios by fusing multimodal perception models according to claim 1, characterized in that, The step involves inputting the interactive mapping perception information into the scene adaptation perception model, which then performs multi-dimensional situational inference processing to generate a panoramic situational description containing the associated operational trajectories and potential change trends of various elements in the operational scene. This includes: The interactive mapping perception information is divided according to the core element type and timestamp to obtain the cross-modal feature data sequence of each core element at different time nodes; The element trajectory restoration module of the scene adaptation perception model integrates the cross-modal feature data sequence of each core element, and restores the actual running path of the core element based on the motion trajectory features of the visual modality and the coordinate distribution features of the position modality. By using the interaction relationship inference module of the scene adaptation perception model, the intersection points and time overlap segments of the operation paths of different core elements are analyzed. Combined with the force effect transmission characteristics of mechanical modes, the actual interaction process between core elements is restored. Extract the changing patterns of the operation path of each core element, and combine them with the dynamic change characteristics in the interactive mapping perception information to deduce the possible subsequent operation direction and state transition trend of the core elements. Analyze the mutual influence of the potential operational trends of different core elements, and judge the new interactive relationships and effects that may be formed between elements in the future based on cross-modal interaction association rules; By integrating the actual operation path, actual interaction process, potential operation trend, and future interaction relationship of all core elements, a multi-dimensional situational model of the lift operation scenario is constructed. The operational paths in the multidimensional situation model are sorted and interpolated according to timestamps, so that the operational status and interaction process of each element are presented on a unified time axis. The potential trends and possible interaction relationships of each element in the multidimensional situation model are prioritized and divided into levels according to preset rules. The sorted operational paths, interaction processes, potential trends and future interaction relationships are transformed into text descriptions and graphical representations. The text descriptions include details of element operation, descriptions of interaction processes, and trend predictions. The graphical representations include path diagrams, interaction relationship diagrams, and trend prediction diagrams. By integrating standardized text descriptions and graphical representations, and supplementing implicit situational information that is not directly presented in the multi-dimensional situational model but is derived through cross-modal feature correlation, a panoramic situational description reflecting the current state and future changes of the operational scenario is formed.

4. The method for understanding lift operation scenarios by fusing multimodal perception models according to claim 1, characterized in that, The extraction of scene representation details from the visual modal perception information set to form a visual modal representation set includes: The visual modal perception information in the trigger perception information set is processed by frame segmentation to obtain a continuous visual frame sequence, and each visual frame contains a complete picture of the work scene. For each visual frame, the contour morphology parameters of the lifting actuator are extracted, the contour morphology difference value of the actuator in adjacent visual frames is calculated, and the extension, contraction, and rotation morphology changes of the actuator are determined based on the difference value to form morphology change features. Analyze the color distribution of core elements in each visual frame, extract the main color tone information and color uniformity information of each core element, compare the color changes of the same core element in adjacent visual frames, record the changes in color depth and brightness, and form color distribution characteristics. In a continuous sequence of visual frames, key feature points of each core element are marked, the positional changes of key feature points in different visual frames are tracked, key feature points in each frame are connected to form motion trajectory segments of the core element, and information on the direction, curvature, and extension speed of the trajectory segments is extracted to form motion trajectory features. For each visual frame, the morphological change features, color distribution features, and motion trajectory features are labeled with correlations to define the correspondence between the morphological change features, color distribution features, and motion trajectory features of the same core element. The scene representation details of all visual frames are integrated in chronological order to form a visual feature sequence of each core element at different time points. The visual feature sequences of all core elements are classified according to the core element type to construct a visual feature classification set. Redundancy removal is performed on the feature data in the visual feature classification set. Visual modality labels and core element labels are added to the redundancy-removed feature data. All labeled feature data are then integrated to form a visual modality representation set.

5. The method for understanding lift operation scenarios by fusing multimodal perception models according to claim 2, characterized in that, The step of extracting dynamic change features of the scene from the interactive mapping perception information to form a dynamic change feature set includes: The cross-modal association data in the interactive mapping perception information is analyzed, and the current running state information of each core element is extracted. The current running state information includes motion speed, force magnitude, spatial position, and morphological state. By comparing the current operational status information of the same core element at different time points, the type of operational status transition is identified, and the starting conditions, transition duration, and post-transition status characteristics of the transition process are recorded to form operational status transition characteristics. Analyze the cross-modal correlation strength and type between different core elements, track the changes in correlation, record the adjustment process of the establishment, strengthening, weakening and termination of correlation, extract the key triggering factors and manifestations in the adjustment process, and form interactive relationship adjustment characteristics; Extract cross-modal feature data related to the structure of the working environment from the interactive mapping perception information, calculate the occlusion range and relative distance changes of the environmental structure on the lifting machine's execution components and the work-bearing objects, and form environmental impact change characteristics; The characteristics of operational state transition, interaction relationship adjustment, and environmental impact change are marked with time series, and the time nodes and duration periods corresponding to each type of characteristic are defined. Analyze the interaction relationships among operational state transition characteristics, interaction relationship adjustment characteristics, and environmental impact change characteristics; identify the triggering effect of operational state transition characteristics on interaction relationship adjustment characteristics and the impact of environmental impact change characteristics on operational state transition characteristics; and establish causal relationship links among the characteristics. All dynamic change features are classified and organized according to core element type and feature type to construct a dynamic change feature classification framework; The dynamic change features after classification are supplemented with details to improve the key information in the description of the dynamic change features, so that the features can accurately reflect the essence of scene changes. Add an associated modality identifier to each dynamically changing feature to define which modal interaction mapping information the feature data originates from; By integrating all the categorized and organized dynamic change features, causal links, and associated modal identifiers, a set of dynamic change features is formed.

6. The method for understanding lift operation scenarios by fusing multimodal perception models according to claim 3, characterized in that, The interaction relationship deduction module of the scene adaptation perception model analyzes the intersection points and time overlap segments of the operation paths of different core elements, and, combined with the force effect transmission characteristics of mechanical modes, reconstructs the actual interaction process between core elements, including: Obtain the actual running path data for each core element. The actual running path data includes the coordinate information, timestamp information, and morphological feature information of each point on the path and the corresponding visual frame. Import the actual operation path data of all core elements into the interaction relationship inference module of the scene adaptation perception model, and identify the spatial intersection points between the operation paths of different core elements through the path analysis algorithm in the interaction relationship inference module. Extract the timestamp information corresponding to each spatial intersection point, compare the timestamps of different core elements arriving at the spatial intersection point, and determine the time overlap segment. The time overlap segment is the time period when different core elements are simultaneously at the spatial intersection point. From the interactive mapping perception information, extract the mechanical modal perception information corresponding to the time coincidence segment, and analyze the force effect transmission characteristic data in the mechanical modal perception information, focusing on the changes in the force distribution characteristics and reaction force feedback characteristics; By combining the coordinates of the spatial intersection point with the duration of the time overlap segment, the mechanical modal force effect transmission characteristic data are compared with a preset threshold to determine whether the contact mode between the core elements is direct or indirect. Based on the changes in contact method and force transmission characteristics, the force generation process between core elements is reconstructed, including the start time, peak time, and decay time of the force. By combining the morphological change features in visual modal perception information, observe the morphological changes of core elements in the overlapping time segment, verify the rationality of the force generation process, and supplement the details of morphological changes in the interaction process. Extract the positional modal spatial occupancy features of core elements during the interaction process, and reconstruct the details of the relative position adjustment and spatial posture changes of core elements during the interaction process; Integrate information on spatial intersection points, time overlap segments, changes in force transmission characteristics, details of morphological changes, and spatial position adjustments, and deduce the interaction process between core elements in chronological order. The derived interaction flow is transformed into a standardized interaction process description, including interaction start conditions, interaction methods, state changes during the interaction process, and interaction end markers, fully restoring the actual interaction process between core elements.

7. The method for understanding lift operation scenarios by fusing multimodal perception models according to claim 1, characterized in that, The extraction of force effect transmission characteristics from the mechanical modal sensing information set in the trigger sensing information set forms a mechanical modal characterization set, including: Signal analysis is performed on the mechanical modal sensing information in the trigger sensing information set to extract force value change data, force direction data, and force application point data from the original mechanical signal; Based on the coordinates of the point of application of the force and the spatial coordinates of each core element, determine the core element and the position of application corresponding to each force value, establish the force transmission path between different core elements, and mark the direction of force transmission. Analyze the force value change data along each transmission path, extract the distribution of force values ​​on different core elements, including the magnitude distribution of force values, the shape characteristics of the distribution area, and the uniformity of distribution, to form the force distribution characteristics; For each action force, extract the force value data, action direction data, and point of application data of the reaction force, analyze the correspondence between the reaction force and the action force, including the proportional relationship of the force values, the degree of oppositeity of the action directions, and the corresponding position of the point of application, and form the reaction force feedback characteristics; By tracing the transmission process of force along the transmission path, comparing the force value changes at different core elements, calculating the attenuation of force during transmission, analyzing the relationship between attenuation and transmission distance and material properties of core elements, extracting the law of force transmission attenuation, and forming the characteristics of force transmission attenuation. For each force distribution feature, reaction force feedback feature, and force transmission attenuation feature, the transmission path is identified, and the force transmission path to which the feature belongs is defined. Mechanical characteristics are classified according to the interaction relationship between core elements, and force distribution characteristics, reaction force feedback characteristics, and force transmission attenuation characteristics belonging to the same interaction relationship are grouped together. Each set of mechanical features is sorted chronologically, and the feature data is organized according to the time sequence of force transmission to form a mechanical feature sequence for each interaction relationship. Modal identifiers and interaction relationship identifiers are added to each mechanical feature to define the perception mode to which the feature belongs and the corresponding core element interaction relationship. Integrate all labeled mechanical feature sequences to form a set of mechanical modal representations.

8. The method for understanding lift operation scenarios by fusing multimodal perception models according to claim 2, characterized in that, To address response deviations in the feature extraction path, the feature extraction nodes of the preset perception model are adjusted based on the representation of features in the dynamically changing feature set. This includes adding extraction channels adapted to new dynamic features and removing extraction links irrelevant to the current dynamic changes. Analyze the original feature extraction path of the preset perception model, and define the feature type, extraction algorithm, and output format corresponding to each extraction node in the original feature extraction path; Compare the feature types output by each extraction node in the original feature extraction path with the feature types contained in the dynamically changing feature set, and mark the extraction nodes that cannot output any feature type in the dynamically changing feature set as invalid nodes. Analyze the manifestation of new dynamic features, including the data type, structural form, and change pattern of the new dynamic features, and determine the appropriate extraction algorithm type and extraction parameter range; Based on the adapted extraction algorithm type and extraction parameter range, new extraction nodes are added to the original feature extraction path. The position of the new extraction node corresponds to the stage where new dynamic features are generated, so that the new extraction node can accurately capture new dynamic features. A data transmission channel is built for the newly added extraction node so that the feature data extracted by the new extraction node can be smoothly transmitted to the subsequent association analysis module. The transmission format of the data transmission channel is consistent with the input format of the subsequent module. Identify extraction links in the original feature extraction path where the extracted feature types are not included in the dynamically changing feature set; Disconnect the irrelevant extraction link from other modules, delete all extraction nodes and data transmission channels in the irrelevant extraction link, and adjust the parameter settings of the original extraction nodes to enable the extraction algorithm of the original extraction nodes to better adapt to the corresponding feature representation in the dynamically changing feature set. Reconfigure the data flow channel between the new extraction node and the original retained node, set the output order of each extraction node, and input the extracted feature data into the correlation analysis module; By inputting dynamic data from the interactive mapping perception information, the adjusted feature extraction path is tested to verify whether all new dynamic features can be effectively extracted and whether existing effective features can be obtained normally, so that the adjusted feature extraction path can fully adapt to the dynamically changing feature set.

9. A lifting operation scenario understanding system integrating a multimodal perception model, characterized in that, The lift operation scene understanding system integrating a multimodal perception model includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions or code, and the processor is used to run the programs, instructions or code in the memory to implement the lift operation scene understanding method integrating a multimodal perception model as described in any one of claims 1-8.