Multi-mode man-machine interaction intention recognition method, system and equipment for lifting console
By using a multimodal human-computer interaction intent recognition method, combined with ship motion and environmental data, and dynamically selecting the control mode, the problem of misoperation of the hoisting control console when the ship is swaying is solved, and the safe and efficient operation of hoisting operations is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-17
AI Technical Summary
When the ship is rocking violently, the single human-machine interaction mode of the hoisting control panel makes it difficult for the operator to accurately touch the screen, which can easily lead to misoperation and affect the safety and efficiency of the hoisting operation.
A multimodal human-computer interaction intent recognition method is adopted. By acquiring ship motion and environmental data, calculating sway intensity and operator preferences, the method dynamically selects voice, touch and hardware control modes, generates intent objects and calculates comprehensive weight values to ensure the accuracy and efficiency of intent recognition.
It improves the efficiency and accuracy of human-machine interaction on the hoisting control console, avoids misoperation caused by single modality or inappropriate intent recognition, and ensures the safety and efficiency of hoisting operations.
Smart Images

Figure CN121879563A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of shipboard launch control consoles, and in particular to a multimodal human-computer interaction intent recognition method, system and device for launch control consoles. Background Technology
[0002] In the field of marine operations, the hoisting control console plays a crucial role in equipment operation, and its accuracy and efficiency directly affect the safety and efficiency of ship operations. With the continuous development of marine technology, ships are becoming increasingly automated and intelligent, making the operations involved in hoisting control consoles increasingly complex.
[0003] Currently, some launching control consoles employ a single human-machine interface mode, such as touch control. Operators primarily input commands by touching the screen on the console. This method can meet basic operational needs when the ship is relatively stable and the environment is relatively quiet. It utilizes the intuitive display and convenient touch operation of the screen, allowing operators to directly control the relevant equipment; the operation process is relatively simple and straightforward.
[0004] However, when the ship is rocking, and the rocking is relatively violent, it is difficult for the operator to accurately touch the target area on the screen, which can easily lead to misoperation and affect the safety of the hoisting operation. Summary of the Invention
[0005] To reduce operator errors and ensure the safety of hoisting operations, this application provides a multimodal human-machine interaction intent recognition method, system, and device for hoisting control consoles.
[0006] In a first aspect, this application provides a multimodal human-computer interaction intent recognition method for a suspended control console, employing the following technical solution: A multimodal human-computer interaction intent recognition method for a suspended control console includes: The ship's motion data and environmental data of the ship's environment are acquired, and the sway intensity is synthesized based on the motion data; the motion data includes roll angle and pitch angle. Determine the operator's current interaction zone based on the distance and orientation between the operator's current location and the hoisting control console; Based on the environmental data, the operator's current interaction area, and operation preferences, calculate the suitability score for the voice control modality, and based on the shaking intensity, the operator's current interaction area, and operation preferences, calculate the suitability scores for the touch control modality and the hardware control modality, respectively. The control mode with the highest suitability score is selected as the preferred mode, and the control mode with the second highest suitability score is selected as the alternative mode. Initial confidence base values are configured for the hardware input channel, voice input channel and touch input channel. After detecting a valid operation signal, any input channel generates an intent object, which includes at least the source control mode and a signal quality score. The final confidence level for each intent object is calculated based on the signal quality score and the suitability score of the source manipulation mode in the intent object. All intent objects within an alignment window of length Z that point to the same device and have the same action or compatible logic are grouped into the same intent cluster. Calculate the overall weight value of each intent cluster based on the final confidence scores of all intent objects within the cluster and the initial confidence scores of the source modalities. Based on the comprehensive weight value, a valid intent cluster is determined, and the intents in the valid intent cluster are converted into specific control instructions.
[0007] By adopting the above technical solutions, the synthesized sway intensity can reflect the degree of swaying of the ship during motion, and the environmental data can reflect the external environment in which the ship is located. By acquiring the ship's motion data and environmental data, the actual conditions of the ship during operation can be taken into account. Considering the operator's operating preferences and determining the operator's interaction area relative to the hoisting control platform reflects attention to the operator's personalized needs and actual operating position, which can better adapt to the operator's usage habits and improve the comfort and efficiency of interaction. After detecting a valid operating signal, an intent object is generated and its final confidence is calculated. Combining the signal quality score and the suitability score of the source control mode, the system can more accurately evaluate the reliability of each intent object and avoid misjudgments caused by poor signal quality or unsuitable control modes. Establishing an alignment window and classifying intent objects into intent clusters can integrate similar or logically compatible intents, thereby reducing duplicate and redundant intents and improving the efficiency and accuracy of intent recognition. The calculation of the comprehensive weight value of intent clusters further filters and evaluates intents, enabling more accurate determination of valid intent clusters. Finally, the intents in the valid intent clusters are converted into specific control commands, realizing a seamless connection from the operator's operating intent to actual control. The entire process takes into account the actual conditions of ship operation, the personalized needs of operators, and the operating scenarios. Through multimodal interaction methods and intent recognition algorithms, it improves the efficiency and accuracy of human-machine interaction on the hoisting control console and avoids operational errors caused by single modality or inappropriate intent recognition.
[0008] Optionally, the step of synthesizing the sway intensity based on the motion data includes: The roll and pitch angular velocities are input into the constructed empirical formula for swaying, and the swaying intensity is output. The swaying intensity ranges from 0 to 1 and is used to characterize the effect of ship swaying on operation.
[0009] Optionally, the steps for calculating the suitability score of the voice control modality include: The noise attenuation factor is calculated based on the environmental noise in the environmental data, and the current interaction area of the operator is mapped to the position gain coefficient. The voice habit coefficient is also calculated based on the voice preference in the operation preference. The suitability score of the voice control modality is calculated based on the noise attenuation factor, the position gain coefficient, and the voice habit coefficient. The steps for calculating the touch control modality suitability score include: The shaking suppression factor is calculated based on the shaking intensity, and the touch habit coefficient is calculated based on the touch preference in the operation preference. The touch modality suitability score is calculated based on the shake suppression factor, the position gain coefficient, and the voice habit coefficient. The steps for calculating the hardware control modal suitability score include: The swaying influence factor is calculated based on the swaying intensity, and the hardware habit coefficient is calculated based on the hardware preference in the operation preference. The hardware control mode suitability score is calculated based on the swaying influence factor, the position gain coefficient, and the hardware habit coefficient.
[0010] Optionally, the steps for determining valid intent clusters based on the comprehensive weight value include: Select the intent clusters with the highest overall weight values to form a candidate set; If the number of intent clusters in the candidate set is 1, then the valid intent cluster is directly determined; If the number of intent clusters in the candidate set is at least 2, then the action conflict, modal priority and timestamp of all intent clusters in the candidate set are detected in sequence. If all the candidate intent clusters point to the same device and have the same action, they are considered as duplicate instructions, and any one intent cluster is directly selected as a valid intent cluster. If there are no duplicate instructions, the intent cluster with the highest modal suitability score and the source control modality being the preferred modality or alternative modality is determined as a valid intent cluster. If there are still ties, the intent cluster with the latest timestamp is directly determined as a valid intent cluster.
[0011] By adopting the above technical solution, when the number of intent clusters in the candidate set is at least two, the judgment process can be further refined by sequentially detecting action conflicts, modal priority, and timestamps, ensuring that the most reasonable and effective intent cluster is selected. For duplicate instructions pointing to the same device and performing the same actions, any intent cluster is directly selected as the effective intent cluster, avoiding the repeated execution of the same instructions, reducing resource waste, and ensuring the simplicity of system operation. In the absence of duplicate instructions, the intent cluster with the highest modal suitability score, whose source control modality is the preferred or alternative modality, is prioritized as the effective intent cluster. This is because the preferred or alternative modality usually has higher stability and reliability, and a high suitability score indicates that the modality is more suitable in the current scenario, improving the accuracy and safety of operation. If there are still ties, the intent cluster with the latest timestamp is determined as the effective intent cluster, reflecting the importance of the latest operation intent, better meeting the operator's real-time needs, and enabling timely response to the operator's latest instructions, keeping the operation of the hoisting control console synchronized with the operator's intent. Through a reasonable screening and meticulous judgment process, the accuracy and efficiency of intent recognition are improved, resource waste is reduced, and the system is able to respond to the operator's intent in a timely and accurate manner, providing a strong guarantee for the stable and efficient operation of the hoisting control console.
[0012] Optionally, the multimodal human-computer interaction intent recognition method further includes: If the number of intent clusters in the candidate set is 0, then determine whether the overall weight of all intent clusters is lower than the set conflict threshold. If so, it is determined that the intent confidence is insufficient, the suspended control console provides multimodal feedback and prompts "Instruction not recognized, please re-enter", and clears all intent objects in the current alignment window to wait for a new round of operation signals; otherwise, it is further determined whether the advantage difference between the comprehensive weights does not exceed the set minimum advantage difference. If so, the intent cluster with the highest overall weight is selected as the valid intent cluster, and an operation log warning "Modal conflict, executed according to weight priority" is generated; otherwise, a multimodal collaborative confirmation mechanism is activated: the question voice "Execute [action description]?" is played, and a confirmation button pops up on the touch interface. If no confirmation signal is received within the set time, the current batch of intent clusters is automatically abandoned.
[0013] By adopting the above technical solution, when the number of intent clusters in the candidate set is 0, it is first determined whether the comprehensive weight of all intent clusters is lower than the set conflict threshold. This effectively identifies whether the operator's operation intent is clear and reliable enough. If the comprehensive weights are all lower than the conflict threshold, it is determined that the intent confidence is insufficient. At this time, the hoisting control console provides multimodal feedback and prompts "Instruction not recognized, please re-enter," while clearing all intent objects in the current alignment window and waiting for a new round of operation signals. This avoids the system operating based on unreliable intents, prevents misoperation, and ensures the safety and accuracy of hoisting operations. If the comprehensive weights are not all lower than the conflict threshold, it is further determined whether the advantage difference between the comprehensive weights does not exceed the set minimum advantage difference. If the advantage difference does not exceed the minimum advantage difference, it indicates that the differences between the intent clusters are small, and the possibility of modal conflict is high. At this time, the intent cluster with the highest comprehensive weight is selected as the valid intent cluster, and an operation log warning "Modal conflict, executed according to weight priority" is generated. This ensures that the system can continue to execute operations while recording the modal conflict situation through operation log warnings, which helps to continuously optimize the system's intent recognition capability. If the advantage difference between the overall weights exceeds the set minimum advantage difference, a multimodal collaborative confirmation mechanism is activated, playing the voice prompt "Execute [action description]?" and displaying a confirmation button on the touch interface. This fully utilizes both voice and touch modes to confirm the operator's intention in a more intuitive and clear way, improving the accuracy and reliability of the operation. If no confirmation signal is received within the set time, the current batch of intent clusters is automatically abandoned, avoiding prolonged waiting for invalid signals, improving system efficiency, and ensuring that the system can respond to subsequent operation commands in a timely manner.
[0014] Optionally, the steps following the conversion of the intents in the valid intent cluster into specific control instructions include: Real-time acquisition of the status response data of the controlled equipment and determination of whether the actual response action deviates from the expected action based on the status response data; If so, a graded anomaly handling process is triggered based on the degree of deviation: when the deviation is less than the first threshold, it is only recorded in the operation log and marked as "slight drift"; when the deviation is between the first and second thresholds, a prompt message "Action deviation detected, recalibrate?" is pushed; when the deviation exceeds the second threshold, the subsequent instruction queue is immediately suspended and an alarm is triggered, and the second threshold is greater than the first threshold.
[0015] Optionally, the multimodal human-computer interaction intent recognition further includes: Regularly analyze the recognition success rate of each control mode over the past T hours; The performance score of the corresponding control mode is calculated based on the recognition success rate and environmental data. The initial confidence base value of the corresponding control mode is dynamically adjusted based on the performance score.
[0016] By adopting the above technical solution, periodically calculating the recognition success rate reflects the performance of each control modality in actual use, providing a quantitative understanding of its performance. Based on the recognition success rate and environmental data, the system calculates the effectiveness score of each control modality, comprehensively considering both the modality's own performance and the ship's actual environment. This allows the system to adjust the weight of each control modality in intent recognition according to its actual performance and environmental adaptability. Through periodic statistics, comprehensive evaluation, and dynamic correction, the multimodal human-machine interaction intent recognition system can better adapt to different environments and changes in the actual performance of each control modality, improving the accuracy and reliability of intent recognition and further optimizing the human-machine interaction experience of the hoisting control console.
[0017] Secondly, this application provides a multimodal human-computer interaction intent recognition system for a suspended control console, employing the following technical solution: A multimodal human-computer interaction intent recognition system for a suspended control console includes: The data acquisition module is used to acquire the ship's motion data and the environmental data of the ship's environment, and to synthesize the sway intensity based on the motion data; the motion data includes roll angle and pitch angle, as well as the operator's operating preferences. The location module is used to determine the operator's current interaction area based on the distance and orientation between the operator's current location and the hoisting control console; The data processing module is used to calculate the suitability score of the voice control mode based on the environmental data, the operator's current interaction area, and operation preferences; and to calculate the suitability score of the touch control mode and the hardware control mode based on the shaking intensity, the operator's current interaction area, and operation preferences, respectively; and to select the control mode with the highest suitability score as the preferred mode, the control mode with the second highest suitability score as the alternative mode, and to configure initial confidence base values for the hardware input channel, voice input channel, and touch input channel. The intent recognition and processing module is used to generate an intent object after any input channel detects a valid operation signal. The intent object includes at least the source control mode and a signal quality score. It is also used to calculate the final confidence of each intent object based on the signal quality score and the suitability score of the source control mode in the intent object. Furthermore, it is used to group all intent objects pointing to the same device and having the same action or logical compatibility within an alignment window of length Z into the same intent cluster. Finally, it is used to calculate the comprehensive weight value of each intent cluster based on the final confidence of all intent objects in the cluster and the initial confidence base value of the source mode, and to determine the valid intent cluster based on the comprehensive weight value. The instruction module is used to convert the intents in the valid intent cluster into specific control instructions and dispatch them.
[0018] Thirdly, this application provides a computer device that adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the multimodal human-computer interaction intent recognition method for a suspended control console as described in the first aspect.
[0019] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing the multimodal human-computer interaction intent recognition method for a suspended control console as described in the first aspect. Attached Figure Description
[0020] Figure 1 This is a first flowchart of an embodiment of the method of this application; Figure 2 This is a second flowchart of an embodiment of the method of this application; Figure 3 This is a third flowchart of an embodiment of the method of this application; Figure 4 This is the fourth flowchart of an embodiment of the method of this application. Detailed Implementation
[0021] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-4 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.
[0022] The first embodiment of this application discloses a method for multimodal human-computer interaction intent recognition of a suspended control console. (Refer to...) Figure 1 The multimodal human-computer interaction intent recognition method includes S110-S190: S110, acquire the ship's motion data and environmental data of the ship's environment, and synthesize the rolling intensity based on the motion data; the motion data includes roll angle and pitch angle; S120: Obtain the operator's operation preferences and determine the operator's current interaction area based on the distance and orientation between the operator's current location and the hoisting control console; S130: Based on environmental data, the operator's current interaction area and operation preferences, calculate the suitability score of the voice control mode, and based on the shaking intensity, the operator's current interaction area and operation preferences, calculate the suitability score of the touch control mode and the hardware control mode respectively. S140, the control mode with the highest suitability score is selected as the preferred mode, the control mode with the second highest suitability score is selected as the alternative mode, and initial confidence base values are configured for the hardware input channel, voice input channel and touch input channel, with the sum of the initial confidence base values being 1; S150, after any input channel detects a valid operation signal, it generates an intent object, which includes the source control modality, action description, timestamp, and signal quality score; S160, Calculate the final confidence level for each intent object based on the signal quality score and the suitability score of the source manipulation mode in the intent object; S170, with the current moment as the center, establish an alignment window of length Z, and group all intent objects within the alignment window that point to the same device and have the same action or logical compatibility into the same intent cluster; S180, Calculate the comprehensive weight value of each intent cluster based on the final confidence of all intent objects within the cluster and the initial confidence base value of the source modality; S190: Based on the comprehensive weight value, determine the valid intent clusters and convert the intents in the valid intent clusters into specific control instructions.
[0023] S110, the steps for synthesizing sway intensity based on motion data include: Specifically, the system collects real-time motion data of the ship using inertial measurement units (IMUs) sensors installed on the ship's deck. This motion data includes roll and pitch angles, and the roll and pitch angular velocities are obtained by calculating their time derivatives. Environmental sensors such as noise meters and anemometers are used to acquire environmental data such as noise levels and wind speed. The motion and environmental data are transmitted to the central processing unit via CAN bus or RS-485 protocol at a sampling rate of 10 times per second, where preprocessing such as filtering and noise reduction is performed to ensure the accuracy of the input data.
[0024] The system then inputs the aforementioned roll and pitch angular velocities into a preset empirical formula for swaying, and outputs the sway intensity to characterize the impact of ship swaying on operation. The empirical formula for swaying is as follows: Where, the sway intensity M∈[0,1]; k1 and k2 are weighting coefficients, which can be calibrated according to the ship type and deck operation area, k1+k2=1, for example, for large container ships with a high center of gravity, the pitching effect is greater, so k2>k1 can be set; Indicates the roll angular velocity. Indicates the pitch angular velocity, It is the preset maximum safe angular velocity threshold, which can be determined based on the safety threshold recommended by the International Maritime Organization or actual ship testing, and is usually set to about 5° / s.
[0025] In step S120, the system uses positioning technologies such as UWB ultra-wideband tags or visual tracking systems to determine the operator's current position. It then calculates the relative distance and azimuth between the operator and the control console based on their spatial coordinates, thereby dividing the system into multiple interactive areas, including the near-field area, mid-field area, far-field area, and non-facing area. For example, the near-field area is less than 2 meters, the mid-field area is 2 to 5 meters, the far-field area is greater than 5 meters, and the non-facing area represents an angle of more than 90° away from the control console.
[0026] The system retrieves personalized preference settings stored in the operator's profile, including usage tendencies for voice, touch, and hardware buttons, to form an initial operating style profile. It should be noted that these personalized preferences can be configured through questionnaires or automatically inferred from historical behavior clustering.
[0027] The steps for calculating the suitability score of the voice control modality include: The noise attenuation factor is calculated based on the environmental noise in the environmental data, and the operator's current interaction area is mapped to the position gain coefficient. The voice habit coefficient is calculated based on the voice preference in the operation preference. The suitability score of the voice control mode is calculated based on the noise attenuation factor, the position gain coefficient and the voice habit coefficient.
[0028] Specifically, based on the collected ambient noise level, a noise attenuation factor is calculated: Noise attenuation factor = max(0, 1 - (ambient noise - noise threshold at which speech clarity begins to decline) / (noise upper limit at which speech communication basically fails - noise threshold at which speech clarity begins to decline)). When the ambient noise is below the threshold of "speech clarity begins to decline," such as 65 dB, the attenuation factor is 1; as the noise rises to the upper limit of "speech communication basically fails," such as 85 dB, the factor linearly decreases to 0, reflecting the degradation trend of voice channel availability. The operator's interaction area is mapped to a position gain coefficient, characterizing the enhancement effect of physical proximity on the convenience of various input methods. For example, the near-field area is set to 1.0, the mid-field area to 0.8, the far-field area to 0.5, and the non-facing area to 0.3. Voice habit coefficient = The probability of using voice in the current environment, retrieved from operator historical data + Voice preferences in operation preferences, among which To control the balance between historical behavior and static preferences, the suitability score of the voice control modality is calculated. = Noise attenuation factor × position gain coefficient × speech habit coefficient.
[0029] The steps for calculating the touch control modality suitability score include: The sway suppression factor is calculated based on the sway intensity, and the touch habit coefficient is calculated based on the touch preference in the operation preferences; the touch modality suitability score is calculated based on the sway suppression factor, the position gain coefficient, and the voice habit coefficient.
[0030] Specifically, a shake suppression factor is calculated based on the shake intensity: Shake Suppression Factor = max(0, 1 - Shake Intensity / Shake Threshold for Significant Touch Operation Inaccuracy). That is, when M exceeds the shake threshold for significant touch operation inaccuracy (e.g., 0.6), the suppression factor drops to 0; otherwise, it linearly decreases to 1, indicating that increased finger positioning error under severe shaking leads to decreased touch reliability. Touch Habit Coefficient = The probability of using the touchscreen in the current environment, retrieved from operator historical data + Touch preference in operation preferences. Touch mode suitability score = shake suppression factor × position gain coefficient × touch habit coefficient, reflecting the impact of dynamic environment on touch screen operation.
[0031] The steps for calculating the hardware control modal suitability score include: The swaying influence factor is calculated based on the swaying intensity, and the hardware habit coefficient is calculated based on the hardware preference in the operation preference; the hardware control mode suitability score is calculated based on the swaying influence factor, the position gain coefficient, and the hardware habit coefficient.
[0032] Specifically, the swaying influence factor = 1 - (1, Shaking Intensity / Shaking Threshold for Significant Inaccuracy in Touch Operation), where p represents the degree to which hardware operation is affected by shaking. For example, if the knob is relatively stable, p = 0.3; if it's a handheld remote control, p = 0.7, etc. Hardware Habit Coefficient = The probability of hardware usage in the current environment, retrieved from operator historical data + Hardware preference in operational preferences. Hardware suitability score = Shaking influence factor × Position gain coefficient × Hardware habit coefficient.
[0033] The system compares the suitability scores of all modalities, selects the highest as the preferred modality, and the second highest as the alternative modality; the initial confidence base value is set based on expert rules, with a sum of 1, and is stored in memory, such as 0.4 for hardware, 0.3 for voice, and 0.3 for touch.
[0034] Reference Figure 2 In S190, the steps for determining valid intent clusters based on the comprehensive weight value include S210-S260: S210, select the intent clusters with the highest comprehensive weight values to form a candidate set; S220, if the number of intent clusters in the candidate set is 1, then the valid intent cluster is directly determined; S230, if the number of intent clusters in the candidate set is at least 2, then the action conflict, modal priority and timestamp of all intent clusters in the candidate set are detected in turn; if all the candidate intent clusters point to the same device and have the same action, they are considered as duplicate instructions, and any one intent cluster is directly selected as a valid intent cluster; if there are no duplicate instructions, the intent cluster with the highest modal suitability score and the source control modality being the preferred modality or alternative modality is determined as a valid intent cluster; if there are still ties, the intent cluster with the latest timestamp is directly determined as a valid intent cluster. S240, If the number of intent clusters in the candidate set is 0, then determine whether the overall weight of all intent clusters is lower than the set conflict threshold. S250, if so, it is determined that the intent confidence is insufficient, the suspended control console performs multimodal feedback and prompts "Instruction not recognized, please re-enter", and clears all intent objects in the current alignment window to wait for a new round of operation signals; otherwise, it is further determined whether the advantage difference between the comprehensive weights does not exceed the set minimum advantage difference. S260, if so, select the intent cluster with the highest comprehensive weight as the valid intent cluster and generate an operation log warning "Modal conflict, executed according to weight priority"; otherwise, start the multimodal collaborative confirmation mechanism: play the question voice "Execute [action description]?", and pop up a confirmation button on the touch interface. If no confirmation signal is received within the set time, the current batch of intent clusters will be automatically abandoned.
[0035] Specifically, in step S150, when an operation occurs, once any input channel detects a valid signal that conforms to semantic rules, such as voice keyword wake-up, touch gesture completion, or button press, it is encapsulated into a structured intent object, which includes source modality, action description, precise timestamp, and signal quality score. The signal quality score is given by the internal algorithm of each channel, such as the word accuracy of speech recognition, the smoothness of touch trajectory, etc., in the range [0, 1].
[0036] Final confidence score for each intent object = signal quality score c + (1 - c) The source mode corresponds to the suitability score; c represents the fusion weight coefficient, which is usually set to 0.6, indicating that more emphasis is placed on the quality of the current actual signal.
[0037] In step S170, the system establishes a sliding alignment window of length Z (e.g., 2 seconds). Centered on the current moment, it collects all intent objects generated within this time period and clusters them according to the criteria of "pointing to the same device" and "same action or logical compatibility," forming several intent clusters. Logical compatibility refers to cases where "ascend" and "accelerate ascend" are considered compatible. Each cluster represents a potential independent operation intent.
[0038] The system then calculates the overall weight value for each intent cluster. The overall weight is determined by the final confidence scores of all intents within the cluster and the initial weight base values of their modalities. ,in, Let i represent the overall weight of cluster k, and let i represent the i-th intent object in cluster k. The initial confidence baseline value configured for the control mode of intention i. This represents the final confidence level of the i-th intention.
[0039] The system selects several intent clusters with the highest overall weight to form a candidate set. If there is only one cluster, it is directly determined as a valid intent cluster. If there are multiple clusters, a conflict resolution mechanism is activated. First, it checks for duplicate instructions, i.e., multiple clusters pointing to the same device and performing completely identical actions. If so, any one of them can be selected. Otherwise, the cluster originating from the preferred modality with the highest modality suitability score is selected first. If there is no preferred modality, the cluster originating from the alternative modality with the highest modality suitability score is further selected. If there are still ties or no preferred or alternative modalities exist, the intent cluster with the latest timestamp is directly selected as the valid intent cluster. When there are no clusters in the candidate set, the system enters low-confidence processing. If the overall weight of all intent clusters is lower than the preset conflict threshold, it indicates that there is a lack of sufficiently reliable operation signals. The system determines this as "insufficient intent confidence" and triggers multimodal feedback. It announces "Command not recognized, please re-enter" via voice and flashes a prompt icon on the touchscreen. Then, it clears all intent objects in the current window and waits for a new round of signal input. Conversely, if there are several clusters with relatively high weights but slight differences between them, the system further determines whether their dominance differences are all less than the minimum dominance difference. If so, the highest-weighted cluster is adopted, and an operation log warning "Modal conflict, executed according to weight priority" is recorded, achieving transparent traceability. If not, it indicates a significant competitive intent, and the system initiates a multimodal collaborative confirmation mechanism, playing a voice prompt asking "Execute [action description]?" and displaying a confirmation / cancel button on the touch interface. A response time limit is set, such as 5 seconds. If a confirmation signal is received during this period, the intent is locked; otherwise, all intent clusters in this batch are automatically abandoned to avoid the risk of misoperation.
[0040] Once a valid intent cluster is identified, the system converts the intent into a specific sequence of control commands and sends it to the corresponding actuator, such as a crane controller or a servo drive module.
[0041] Reference Figure 3 S190, the steps following the conversion of the intent in the valid intent cluster into specific control instructions include S310-S330: S310, real-time acquisition of status response data of the controlled equipment; S320, determine whether the actual response action deviates from the expected action based on the status response data; S330, if so, triggers a graded anomaly handling process based on the degree of deviation: when the deviation is less than the first threshold, it is only recorded in the operation log and marked as "slight drift"; when the deviation is between the first threshold and the second threshold, a prompt message "Action deviation detected, recalibrate?" is pushed; when the deviation exceeds the second threshold, the subsequent instruction queue is immediately suspended and an alarm is triggered, and the second threshold is greater than the first threshold.
[0042] Specifically, the system collects real-time status response data of the controlled equipment through the device feedback interface, monitoring the consistency between the actual action trajectory, speed, and target command. The device feedback interface includes, but is not limited to, encoders, limit switches, and CAN bus status messages. If a deviation is found between the actual response and the expected action, the deviation analysis process is immediately initiated. Specifically, when the deviation is less than the first threshold, it is only recorded in the operation log and marked as "slight drift"; when the deviation is between the first and second thresholds, a pop-up window is pushed with the message "Action deviation detected, recalibrate?" for manual intervention; when the deviation exceeds the second threshold, the subsequent command queue is immediately suspended, power output is cut off, an audible and visual alarm is triggered, and the on-duty personnel are notified to investigate the fault.
[0043] Reference Figure 4 Multimodal human-computer interaction intent recognition also includes S410-S430: S410, periodically calculates the recognition success rate of each control mode over the past T hours; S420 calculates the performance score of the corresponding control mode based on the recognition success rate and environmental data; S430 dynamically adjusts the initial confidence baseline value of the corresponding control mode based on the performance score.
[0044] Specifically, the success rate of each control modality—that is, the proportion of successfully executed and error-free operations—is periodically calculated, such as hourly, over the past T hours (e.g., 24 hours). This is combined with environmental data and operational scenarios, including day / night or single / multi-person operations, to train a lightweight regression model or calculate the performance score of each modality using a sliding weighted method. Then, the initial confidence baseline value for the next period is dynamically adjusted using this performance score. The higher the performance score, the higher the initial confidence baseline value for that modality should be, and vice versa. The adjustment formula is: Initial confidence baseline value = Old baseline value × (1 + p × (Performance score - Threshold)), where p is the adjustment coefficient, such as 0.2, and the threshold is the minimum acceptable success rate for the modality.
[0045] Based on the above method embodiments, the second embodiment of this application discloses a multimodal human-computer interaction intent recognition system for a hoisting control console. The multimodal human-computer interaction intent recognition system for a hoisting control console of this application embodiment can implement any of the above-described methods for multimodal human-computer interaction intent recognition of a hoisting control console, and the specific working process of each module in the multimodal human-computer interaction intent recognition system for a hoisting control console can be referred to the corresponding process in the above method embodiments.
[0046] For ease of understanding, an example is provided below: A multimodal human-computer interaction intent recognition system for a suspended control console includes: The data acquisition module is used to acquire the ship's motion data and environmental data of the ship's environment, and to synthesize the rolling intensity based on the motion data, as well as to acquire the operator's operating preferences; the motion data includes roll angle and pitch angle; The location module is used to determine the operator's current interaction area based on the distance and orientation between the operator's current location and the hoisting control console; The data processing module is used to calculate the suitability score of the voice control mode based on environmental data, the operator's current interaction area, and operation preferences; and to calculate the suitability scores of the touch control mode and the hardware control mode based on the shaking intensity, the operator's current interaction area, and operation preferences, respectively; and to select the control mode with the highest suitability score as the preferred mode, the control mode with the second highest suitability score as the alternative mode, and to configure the initial confidence base value for the hardware input channel, voice input channel, and touch input channel. The intent recognition and processing module is used to generate an intent object after any input channel detects a valid operation signal. The intent object includes at least the source control mode and the signal quality score. It is also used to calculate the final confidence of each intent object based on the signal quality score and the suitability score of the source control mode in the intent object. Furthermore, it is used to group all intent objects that point to the same device and have the same action or are logically compatible within an alignment window of length Z into the same intent cluster. Finally, it is used to calculate the comprehensive weight value of each intent cluster based on the final confidence of all intent objects in the cluster and the initial confidence base value of the source mode, and to determine the valid intent cluster based on the comprehensive weight value. The instruction module is used to convert the intents in the valid intent cluster into specific control instructions and dispatch them.
[0047] The third embodiment of this application provides a computer device, which may include a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement a multimodal human-computer interaction intent recognition method for a suspended control console.
[0048] The memory can communicate with the processor via a communication bus, which can be an address bus, a data bus, a control bus, etc.
[0049] Additionally, the memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device.
[0050] Furthermore, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0051] The fourth embodiment of this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as a multimodal human-computer interaction intent recognition method for a suspended control console.
[0052] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0053] It should be noted that the computer device and storage medium in the embodiments of this application are respectively electronic devices and storage media that apply the above-described multimodal human-computer interaction intent recognition method for the suspended control console. Therefore, all embodiments of the above-described multimodal human-computer interaction intent recognition method for the suspended control console are applicable to the computer device and storage medium, and can achieve the same or similar beneficial effects. For the computer device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple; relevant details can be found in the descriptions of the method embodiments.
[0054] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.
[0055] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A method for recognizing multimodal human-computer interaction intents in the context of a suspended control console, characterized in that... include: The ship's motion data and environmental data of its surroundings are acquired, and the sway intensity is synthesized based on the motion data; the motion data includes roll angle and pitch angle. Determine the operator's current interaction zone based on the distance and orientation between the operator's current location and the hoisting control console; Based on the environmental data, the operator's current interaction area, and operation preferences, calculate the suitability score for the voice control modality, and based on the shaking intensity, the operator's current interaction area, and operation preferences, calculate the suitability scores for the touch control modality and the hardware control modality, respectively. The control mode with the highest suitability score is selected as the preferred mode, and the control mode with the second highest suitability score is selected as the alternative mode. Initial confidence base values are configured for the hardware input channel, voice input channel and touch input channel. After detecting a valid operation signal, any input channel generates an intent object, which includes at least the source control mode and a signal quality score. The final confidence level for each intent object is calculated based on the signal quality score and the suitability score of the source manipulation mode in the intent object. All intent objects within an alignment window of length Z that point to the same device and have the same action or compatible logic are grouped into the same intent cluster. Calculate the overall weight value of each intent cluster based on the final confidence scores of all intent objects within the cluster and the initial confidence scores of the source modalities. Based on the comprehensive weight value, a valid intent cluster is determined, and the intents in the valid intent cluster are converted into specific control instructions.
2. The multimodal human-computer interaction intent recognition method for a suspended control console according to claim 1, characterized in that, The steps for synthesizing sway intensity based on the motion data include: The roll and pitch angular velocities are input into the constructed empirical formula for swaying, and the swaying intensity is output. The swaying intensity ranges from 0 to 1 and is used to characterize the effect of ship swaying on operation.
3. The multimodal human-computer interaction intent recognition method for a suspended control console according to claim 1, characterized in that, The steps for calculating the suitability score of the voice control modality include: The noise attenuation factor is calculated based on the environmental noise in the environmental data, and the current interaction area of the operator is mapped to the position gain coefficient. The voice habit coefficient is also calculated based on the voice preference in the operation preference. The suitability score of the voice control modality is calculated based on the noise attenuation factor, the position gain coefficient, and the voice habit coefficient. The steps for calculating the touch control modality suitability score include: The shaking suppression factor is calculated based on the shaking intensity, and the touch habit coefficient is calculated based on the touch preference in the operation preference. The touch modality suitability score is calculated based on the shake suppression factor, the position gain coefficient, and the voice habit coefficient. The steps for calculating the hardware control modal suitability score include: The swaying influence factor is calculated based on the swaying intensity, and the hardware habit coefficient is calculated based on the hardware preference in the operation preference. The hardware control mode suitability score is calculated based on the sway influence factor, the position gain coefficient, and the hardware habit coefficient.
4. The multimodal human-computer interaction intent recognition method for a suspended control console according to claim 1, characterized in that, The steps for determining valid intent clusters based on the comprehensive weight value include: Select the intent clusters with the highest overall weight values to form a candidate set; If the number of intent clusters in the candidate set is 1, then the valid intent cluster is directly determined; If the number of intent clusters in the candidate set is at least 2, then the action conflict, modal priority and timestamp of all intent clusters in the candidate set are detected in sequence. If all the candidate intent clusters point to the same device and have the same action, they are considered as duplicate instructions, and any one intent cluster is directly selected as a valid intent cluster. If there are no duplicate instructions, the intent cluster with the highest modal suitability score and the source control modality being the preferred modality or alternative modality is determined as a valid intent cluster. If there are still ties, the intent cluster with the latest timestamp is directly determined as a valid intent cluster.
5. The multimodal human-computer interaction intent recognition method for a suspended control console according to claim 4, characterized in that, The multimodal human-computer interaction intent recognition method also includes: If the number of intent clusters in the candidate set is 0, then determine whether the overall weight of all intent clusters is lower than the set conflict threshold. If so, it is determined that the intent confidence is insufficient, the suspended control console provides multimodal feedback and prompts "Instruction not recognized, please re-enter", and clears all intent objects in the current alignment window to wait for a new round of operation signals; otherwise, it is further determined whether the advantage difference between the comprehensive weights does not exceed the set minimum advantage difference. If so, the intent cluster with the highest overall weight is selected as the valid intent cluster, and an operation log warning "Modal conflict, executed according to weight priority" is generated; otherwise, a multimodal collaborative confirmation mechanism is activated: the question voice "Execute [action description]?" is played, and a confirmation button pops up on the touch interface. If no confirmation signal is received within the set time, the current batch of intent clusters is automatically abandoned.
6. The multimodal human-computer interaction intent recognition method for a suspended control console according to claim 1, characterized in that, The steps following the conversion of the intents in the valid intent cluster into specific control commands include: Real-time acquisition of the status response data of the controlled equipment and determination of whether the actual response action deviates from the expected action based on the status response data; If so, a graded anomaly handling process is triggered based on the degree of deviation: when the deviation is less than the first threshold, it is only recorded in the operation log and marked as "slight drift"; when the deviation is between the first and second thresholds, a prompt message "Action deviation detected, recalibrate?" is pushed; when the deviation exceeds the second threshold, the subsequent instruction queue is immediately suspended and an alarm is triggered, and the second threshold is greater than the first threshold.
7. The multimodal human-computer interaction intent recognition method for a suspended control console according to claim 1, characterized in that, The multimodal human-computer interaction intent recognition also includes: Regularly analyze the recognition success rate of each control mode over the past T hours; The performance score of the corresponding control mode is calculated based on the recognition success rate and environmental data. The initial confidence base value of the corresponding control mode is dynamically adjusted based on the performance score.
8. A multimodal human-computer interaction intent recognition system for a suspended control console, characterized in that, The method for recognizing the multimodal human-computer interaction intent of the hoisting control console as described in any one of claims 1 to 7 includes: The data acquisition module is used to acquire the ship's motion data and the environmental data of the ship's environment, and to synthesize the sway intensity based on the motion data; the motion data includes roll angle and pitch angle, as well as the operator's operating preferences. The location module is used to determine the operator's current interaction area based on the distance and orientation between the operator's current location and the hoisting control console; The data processing module is used to calculate the suitability score of the voice control mode based on the environmental data, the operator's current interaction area, and operation preferences; and to calculate the suitability score of the touch control mode and the hardware control mode based on the shaking intensity, the operator's current interaction area, and operation preferences, respectively; and to select the control mode with the highest suitability score as the preferred mode, the control mode with the second highest suitability score as the alternative mode, and to configure initial confidence base values for the hardware input channel, voice input channel, and touch input channel. The intent recognition and processing module is used to generate an intent object after any input channel detects a valid operation signal. The intent object includes at least the source control mode and a signal quality score. It is also used to calculate the final confidence of each intent object based on the signal quality score and the suitability score of the source control mode in the intent object. Furthermore, it is used to group all intent objects pointing to the same device and having the same action or logical compatibility within an alignment window of length Z into the same intent cluster. Finally, it is used to calculate the comprehensive weight value of each intent cluster based on the final confidence of all intent objects in the cluster and the initial confidence base value of the source mode, and to determine the valid intent cluster based on the comprehensive weight value. The instruction module is used to convert the intents in the valid intent cluster into specific control instructions and dispatch them.
9. A computer device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the multimodal human-computer interaction intent recognition method for a suspended control console as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 7, for the multimodal human-computer interaction intent recognition method of the suspended control console.