Intelligent man-machine interaction method and system based on multi-modal large model

Through the multimodal large-scale intelligent human-computer interaction system, the problems of insufficient dynamic environment perception and real-time control delay in underwater operations are solved, efficient and accurate human-computer interaction is achieved, the safety and execution efficiency of underwater operation equipment are improved, and the use effect of the interactive interface is optimized through the interface self-check alarm unit.

CN120704534AActive Publication Date: 2025-09-26SUZHOU YUQIA TECH CO LTD

Patent Information

Application Number
CN202510875954.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Traditional human-computer interaction systems have problems in underwater operations, such as insufficient dynamic environment perception, rigid task planning, and real-time control delays. They are unable to achieve accurate intention understanding and efficient decision-making, and are also difficult to effectively monitor the operation of interactive display devices and issue abnormal alarms, affecting the safety and execution efficiency of the equipment.

Method used

An intelligent human-computer interaction system based on a multimodal large model is adopted. Through the modular design of multimodal input processing, intelligent analysis and decision-making, interface protocol management and visual human-computer interaction interface units, multimodal data collection, processing, fusion, decision support and instruction execution are realized. Combined with large model training and tuning and hardware adaptation, autonomous task planning and real-time intelligent control are provided, and the interface status is monitored through the interactive interface self-check alarm unit to generate abnormal alarm signals.

Benefits of technology

It improves the safety and execution efficiency of underwater operation equipment, enhances user experience and operational convenience, reduces the difficulty of managing interactive display equipment, and ensures efficient and stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704534A_ABST
    Figure CN120704534A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of man-machine interaction equipment, and particularly relates to an intelligent man-machine interaction method and system based on a multi-modal large model, and the system comprises a multi-modal input processing unit, an intelligent analysis decision unit, an interface protocol management unit and a visual man-machine interaction interface unit. Through modular design and tight data flow direction control, efficient and accurate man-machine interaction is achieved, all the modules cooperate with one another to jointly complete multi-modal data acquisition, processing, fusion, decision support, instruction execution and feedback tasks, and the system is high in practicability and easy to popularize. Decision-making auxiliary supports such as autonomous task planning, dynamic environment perception and real-time intelligent control are provided for underwater operation equipment, the safety and execution efficiency of the underwater operation equipment are improved, and the problems that a traditional man-machine interaction system is insufficient in perception, rigid in planning and delayed in control in complex scenes such as underwater operation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction equipment, and in particular to an intelligent human-computer interaction method and system based on a multimodal large model. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, intelligent human-computer interaction (HCI) systems have become key to improving device autonomy and operational efficiency. However, traditional HCI systems are often limited to single-modal data input, making it difficult to accurately understand intent and make efficient decisions in complex environments.

[0003] Especially in extreme environments such as underwater operations, traditional human-computer interaction systems face problems such as insufficient dynamic environmental perception, rigid task planning, and real-time control delays. These problems seriously restrict the safety and execution efficiency of underwater operation equipment. Furthermore, it is difficult to effectively monitor the operation of interactive display equipment and issue abnormal alarms, which is not conducive to ensuring interactive display performance and reducing its management difficulty. The system has a low level of intelligence.

[0004] In view of the above technical defects, a solution is now proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent human-computer interaction method and system based on a multimodal large model, which solves the problem that the existing technology is limited by a single modal data input, making it difficult to achieve accurate intention understanding and efficient decision-making in complex environments, seriously restricting the safety and execution efficiency of underwater operation equipment, and making it difficult to effectively monitor the operation of interactive display equipment and alarm for abnormalities, which is not conducive to ensuring interactive display performance and reducing its management difficulty.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] An intelligent human-computer interaction system based on a multimodal large model includes a multimodal input processing unit, an intelligent analysis and decision-making unit, an interface protocol management unit, and a visual human-computer interaction interface unit. The multimodal input processing unit is responsible for collecting raw data from multiple modalities and preprocessing it, and then outputting the preprocessed data to the intelligent analysis and decision-making unit.

[0008] The intelligent analysis and decision-making unit is responsible for the integration of multimodal data, intention understanding, decision planning and instruction generation, and sends the decision results and instructions to the visual human-computer interaction interface unit and the interface protocol management unit;

[0009] The interface protocol management unit receives the decision results and instructions of the intelligent analysis and decision unit, converts them into a protocol format understandable to external devices, and transmits them to the underwater operation equipment for execution. The interface protocol management unit is also responsible for receiving feedback data from external devices and transmitting it to the intelligent analysis and decision unit for status updates and decision adjustments.

[0010] Furthermore, the multimodal input processing module captures environmental sound sources through microphone components, performs noise suppression, echo cancellation and voice enhancement on the collected original audio, and relies on optical sensors to capture scene images, performs digital filtering, grayscale conversion, normalization and resolution adjustment on the original images, and integrates multi-source sensor data including sonar, pressure gauge and thermometer, and performs data fusion and preprocessing on the multi-source sensor data.

[0011] Furthermore, the intelligent analysis and decision-making unit includes a multimodal interaction engine module, an intelligent auxiliary decision-making module and an instruction processing and sequence conversion module; feature extraction and fusion are performed through the multimodal interaction engine module, dynamic strategies and risk predictions are generated using the intelligent auxiliary decision-making module, and intentions are converted into executable instructions through the instruction processing and sequence conversion module.

[0012] Furthermore, the interface protocol management unit supports multiple communication protocols including TCP / IP, Modbus, and CAN, enabling seamless communication between the system and external devices and adapting to the communication needs of different devices. It also provides a protocol conversion engine for data conversion and adaptation between different protocols, as well as real-time monitoring of errors in the protocol conversion process, providing automatic repair and manual intervention functions.

[0013] Furthermore, the visual human-computer interaction interface unit has multimodal input support, real-time data display and interactive operation control functions. Among them, multimodal input support is used to provide multiple input methods including voice, text and gestures. Real-time data display is used to display the sensor data, task progress and equipment status information of underwater operation equipment in real time in the form of charts and dashboards. Interactive operation control is used to provide interactive tools such as drag-and-drop task builders and parameter adjustment panels, supporting users to customize task steps and parameter settings.

[0014] Furthermore, the intelligent analysis and decision-making unit calls the large model training and tuning module during the inference operation process. During the system operation, the large model training and tuning module continuously receives model inference requests from the intelligent analysis and decision-making unit; the large model training and tuning module is responsible for the training, tuning and deployment of multimodal large models, and cleans, labels and vectorizes the received data to generate standardized data sets; adopts pre-training and fine-tuning strategies, uses large-scale data sets to train multimodal large models, and improves the performance of the model in specific business areas through LoRA training and tuning and adaptive hyperparameter optimization technology, optimizes the inference speed and resource consumption, and outputs the optimized model parameters to the intelligent analysis and decision-making unit.

[0015] Furthermore, the intelligent human-computer interaction unit and the large model training and tuning unit send computing power requirements to the hardware adaptation and computing power management unit. The hardware adaptation and computing power management unit provides heterogeneous computing power support and optimizes resource allocation. The specific operation process is as follows:

[0016] A domestic computing cluster is built based on the Kunpeng 920 CPU and Ascend 910 NPU to meet the computing power requirements of large-scale model inference. It also dynamically schedules GPU / NPU resources for multi-task parallel processing and low-latency response, improving overall system performance. It also dynamically adjusts resource allocation based on load, supports large-scale model inference and real-time control, and provides computing power support for large-scale model training and tuning units and intelligent human-computer interaction units.

[0017] Furthermore, the visual human-machine interaction interface unit is communicatively connected to the interactive interface self-checking alarm unit. The interactive interface self-checking alarm unit monitors the operation of the visual human-machine interaction interface unit and analyzes the operating status of the visual human-machine interaction interface unit in real time, thereby determining whether to generate a state abnormality alarm signal and triggering an alarm mechanism when an interface abnormality alarm signal is generated. The specific state analysis process is as follows:

[0018] The real-time interface temperature of the visual human-computer interaction interface unit is obtained and marked as the interface temperature measurement value, and the growth rate of the real-time interface temperature is marked as the interface temperature rise value. The interface temperature measurement value and the interface temperature rise value are numerically compared with the preset interface temperature measurement threshold value and the preset interface temperature rise threshold value respectively. If the interface temperature measurement value or the interface temperature rise value exceeds the corresponding preset threshold value, an abnormal state alarm signal is generated;

[0019] If the interface temperature measurement value and the interface temperature rise value do not exceed the corresponding preset threshold value, the display brightness of the visual human-computer interaction interface unit is collected, the display brightness is calculated and the difference between the display brightness and the set standard display brightness is calculated and the absolute value is taken to obtain the interface display value, and the interface display value is numerically compared with the preset interface display threshold value. If the interface display value exceeds the preset interface display threshold value, an abnormal status alarm signal is generated.

[0020] Furthermore, the interactive interface self-check alarm unit is further configured to set a detection period of duration L1. When the duration reaches L1, a comprehensive evaluation is performed on the operation optimization status of the visual human-computer interactive interface unit during the detection period to determine whether to generate an interface optimization alarm signal. When the interface optimization alarm signal is generated, an alarm mechanism is triggered. The specific evaluation and analysis process is as follows:

[0021] The generation moment of the corresponding state abnormality alarm signal is obtained and marked as the first characteristic moment, and the moment when the corresponding state abnormality alarm is released after regulation is collected and marked as the second characteristic moment, the time difference between the first characteristic moment and the second characteristic moment is calculated to obtain a characteristic duration value, and the characteristic duration value is numerically compared with a preset characteristic duration threshold value. If the characteristic duration value exceeds the corresponding preset characteristic duration threshold value, the corresponding characteristic duration value is marked as a characteristic abnormality value;

[0022] The number of characteristic abnormal values ​​during the detection period is obtained and marked as characteristic anomaly values, and the characteristic anomaly values ​​are numerically compared with the preset characteristic anomaly threshold. If the characteristic anomaly value exceeds the preset characteristic anomaly threshold, an interface optimization alarm signal is generated; if the characteristic anomaly value does not exceed the preset characteristic anomaly threshold, the number of black screens and freezes of the intelligent human-computer interaction interface unit during the detection period and the total duration are obtained and marked as unresponsive frequency and unresponsive duration, respectively. The unresponsive frequency and unresponsive duration are numerically compared with the preset unresponsive frequency threshold and the preset unresponsive duration threshold, respectively. If the unresponsive frequency or the unresponsive duration exceeds the corresponding preset threshold, an interface optimization alarm signal is generated;

[0023] If the unresponsiveness frequency and the unresponsiveness duration do not exceed the corresponding preset thresholds, then when the user performs a click operation on the intelligent human-computer interaction interface unit, the touch buffer duration is collected, and the average of all touch buffer durations within the detection period is calculated to obtain a touch buffer detection value; and when the user performs a drag operation on the intelligent human-computer interaction interface unit, the user's actual drag trajectory and the identified drag trajectory are obtained, the actual drag trajectory and the identified drag trajectory are overlapped and compared, and the length ratio of the non-overlapping trajectory is marked as the trajectory anomaly value, and the average of all trajectory anomalies within the detection period is calculated to obtain a drag anomaly value;

[0024] The interface optimization urgency coefficient is obtained by weighted summing up the characteristic anomaly value, unresponsiveness frequency, unresponsiveness duration, touch delay value and drag recognition value. The interface optimization urgency coefficient is numerically compared with the preset interface optimization urgency coefficient threshold. If the interface optimization urgency coefficient exceeds the preset interface optimization urgency coefficient threshold, an interface optimization alarm signal is generated.

[0025] Furthermore, the present invention also proposes an intelligent human-computer interaction method based on a multimodal large model, comprising the following steps:

[0026] Step 1: Collect raw data from multiple modalities and preprocess them;

[0027] Step 2: Perform multimodal data integration, intent understanding, decision planning, and instruction generation;

[0028] Step 3: Convert the decision results and instructions into a protocol format understandable to the external device and output it;

[0029] Step 4: Control the corresponding execution operation based on the output information.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] 1. The present invention achieves efficient and accurate human-machine interaction through modular design and tight data flow control. The modules collaborate with each other to complete multimodal data acquisition, processing, fusion, decision support, command execution, and feedback tasks. This provides underwater operating equipment with autonomous task planning, dynamic environment perception, and real-time intelligent control, among other decision-making support, thereby improving the safety and execution efficiency of underwater operating equipment.

[0032] 2. In the present invention, the operation of the visual human-computer interaction interface unit is monitored and the operating status of the visual human-computer interaction interface unit is analyzed in real time through the interactive interface self-check alarm unit. Corresponding adjustment measures are taken when an interface abnormality alarm signal is generated, and the operating optimization status of the visual human-computer interaction interface unit during the detection period is comprehensively evaluated. Corresponding improvement and optimization measures are taken when an interface optimization alarm signal is generated, which is beneficial to improving the use effect and operating stability of the visual human-computer interaction interface unit and significantly reducing the difficulty of its operation management. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings;

[0034] Figure 1 This is a system block diagram of Embodiment 1 of the present invention;

[0035] Figure 2 This is a system block diagram of Embodiment 2 of the present invention;

[0036] Figure 3 This is a flow chart of the method of embodiment 3 of the present invention. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0038] Example 1: Figure 1As shown, the present invention proposes an intelligent human-computer interaction system based on a multimodal large model, which includes a multimodal input processing unit, an intelligent analysis and decision-making unit, an interface protocol management unit and a visual human-computer interaction interface unit. It is mainly used to solve the problems of insufficient perception, rigid planning, and control delay existing in traditional human-computer interaction systems in complex scenarios such as underwater operations. By integrating multimodal data, in-depth analysis and contextual understanding of user intentions can be achieved, the safety and execution efficiency of underwater operation equipment can be improved, and the user experience and operation convenience can be improved through the user-side visual interaction interface.

[0039] The multimodal input processing unit is responsible for collecting and preprocessing raw data from multiple modalities, and then outputting the preprocessed data to the intelligent analysis and decision-making unit. For example, the microphone component captures the ambient sound source and performs noise suppression, echo cancellation, and voice enhancement on the collected raw audio to improve the audio signal quality and provide high-quality input for voice interaction.

[0040] It also relies on optical sensors to capture scene images, and performs digital filtering, grayscale conversion, normalization and resolution adjustment on the original images to improve image quality, provide high-quality input for visual interaction, and integrate multi-source sensor data including sonar, pressure gauge and thermometer, and perform data fusion and preprocessing on multi-source sensor data to provide the system with comprehensive environmental perception capabilities.

[0041] The intelligent analysis and decision-making unit is responsible for the integration of multimodal data, intent understanding, decision planning, and instruction generation, and sends the decision results and instructions to the visual human-computer interaction interface unit and the interface protocol management unit. It should be noted that the intelligent analysis and decision-making unit includes a multimodal interaction engine module, an intelligent auxiliary decision module, and an instruction processing and sequence conversion module.

[0042] Specifically, the multimodal interaction engine module performs feature extraction and fusion (integrating multimodal inputs such as voice, vision, and text to achieve in-depth analysis of user intent and contextual understanding. Analysis process: extracting and fusing features from multimodal data through a pre-trained large model, using the attention mechanism to capture key information, generating dynamic strategies and risk predictions, and providing input for the intelligent decision-making support module). The intelligent decision-making support module generates dynamic strategies and risk predictions (providing real-time decision recommendations and risk warnings, optimizing task operation processes). Analysis process: Based on data-driven and model-based reasoning, analyze business dynamics to generate optimization strategies, and combine knowledge bases for risk prediction).

[0043] And convert intentions into executable instructions through the instruction processing and sequence conversion module (convert voice instructions into structured control signals to achieve multi-step temporal reasoning and execution assistance for complex tasks, analysis process: through natural language understanding, parameter extraction and verification, multimodal instruction fusion and other technologies, the voice instructions are accurately decomposed into atomic operations and time-series and spatial control signals are generated).

[0044] The interface protocol management unit receives the decision results and instructions of the intelligent analysis and decision-making unit, converts them into a protocol format understandable by external devices, and transmits them to the underwater operation equipment for execution. The interface protocol management unit is also responsible for receiving feedback data from external devices and transmitting it to the intelligent analysis and decision-making unit for status update and decision adjustment;

[0045] It should be noted that the interface protocol management unit supports multiple communication protocols such as TCP / IP, Modbus, CAN, etc., enabling seamless communication between the system and external devices and adapting to the communication needs of different devices. It also provides a protocol conversion engine to realize data conversion and adaptation between different protocols and solve protocol compatibility issues; as well as real-time monitoring of errors in the protocol conversion process, providing automatic repair and manual intervention functions.

[0046] The visual human-computer interaction interface unit (touch screen) has multimodal input support, real-time data display and interactive operation control functions. Among them, multimodal input support is used to provide multiple input methods including voice, text and gestures. Real-time data display is used to display the sensor data, task progress and equipment status information of underwater operation equipment in real time in the form of charts and dashboards. Interactive operation control is used to provide interactive tools such as drag-and-drop task builders and parameter adjustment panels, supporting users to customize task steps and parameter settings.

[0047] In addition, the intelligent analysis and decision-making unit calls the large model training and tuning module during the inference process. During the system operation, the large model training and tuning module continuously receives model inference requests from the intelligent analysis and decision-making unit. Moreover, the intelligent human-computer interaction unit and the large model training and tuning unit send computing power requirements to the hardware adaptation and computing power management unit. The hardware adaptation and computing power management unit provides heterogeneous computing power support and optimizes resource allocation.

[0048] Among them, the large model training and tuning module is responsible for the training, tuning and deployment of multimodal large models, cleaning, labeling and vectorizing the received data to generate standardized data sets; adopting pre-training and fine-tuning strategies, using large-scale data sets to train multimodal large models, improving the model's generalization ability, and through LoRA training and tuning and adaptive hyperparameter optimization and other technologies, improving the model's performance in specific business areas, optimizing inference speed and resource consumption, and outputting the optimized model parameters to the intelligent analysis and decision-making unit.

[0049] The hardware adaptation and computing power management unit builds a domestic computing power cluster based on the Kunpeng 920 CPU and Ascend 910 NPU to meet the computing power requirements of large-scale model inference, and dynamically schedules GPU / NPU resources for multi-task parallel processing and low-latency response, improving the overall performance of the system. It also dynamically adjusts resource allocation according to the load, supports large-scale model inference and real-time control, ensures efficient and stable operation of the system, and provides computing power support for the large-scale model training and tuning unit and the intelligent human-computer interaction unit.

[0050] Example 2: Figure 2 As shown, the difference between this embodiment and the first embodiment is that the visual human-computer interaction interface unit is communicatively connected to the interactive interface self-checking alarm unit. The interactive interface self-checking alarm unit monitors the operation of the visual human-computer interaction interface unit and analyzes the operating status of the visual human-computer interaction interface unit in real time. Based on this, it is determined whether to generate an abnormal state alarm signal. When the abnormal interface alarm signal is generated, the alarm mechanism is triggered to automatically take corresponding adjustment measures to reduce the operating risk of the visual human-computer interaction interface unit and ensure its use effect. The specific state analysis process is as follows:

[0051] The real-time interface temperature of the visual human-computer interaction interface unit is obtained and marked as the interface temperature measurement value, and the growth rate of the real-time interface temperature is marked as the interface temperature rise value. The interface temperature measurement value and the interface temperature rise value are numerically compared with the preset interface temperature measurement threshold value and the preset interface temperature rise threshold value respectively. If the interface temperature measurement value or the interface temperature rise value exceeds the corresponding preset threshold value, it indicates that the temperature condition of the visual human-computer interaction interface unit is not good and there is a great safety risk, and a state abnormality alarm signal is generated;

[0052] If the interface temperature measurement value and the interface temperature rise value do not exceed the corresponding preset threshold value, indicating that the temperature risk of the visual human-computer interaction interface unit is small, the display brightness of the visual human-computer interaction interface unit is collected, the display brightness is calculated and the difference between the display brightness and the set standard display brightness is calculated and the absolute value is taken to obtain the interface display value, and the interface display value is numerically compared with the preset interface display threshold value. If the interface display value exceeds the preset interface display threshold value, indicating that the display performance of the visual human-computer interaction interface unit is poor, an abnormal status alarm signal is generated.

[0053] Furthermore, the interactive interface self-check alarm unit is further used to set a detection period of L1. When the time reaches L1, preferably, L1 is four hours and eight hours. The operation optimization status of the visual human-computer interaction interface unit during the detection period is comprehensively evaluated to determine whether to generate an interface optimization alarm signal. When the interface optimization alarm signal is generated, the alarm mechanism is triggered so that corresponding improvement and optimization measures can be taken in time, thereby further improving the use effect and operation stability of the visual human-computer interaction interface unit, significantly reducing its operation management difficulty, and having a high level of intelligence. The specific evaluation and analysis process is as follows:

[0054] The generation moment of the corresponding state abnormality alarm signal is obtained and marked as the first characteristic moment, and the moment when the corresponding state abnormality alarm is released after regulation is collected and marked as the second characteristic moment. The time difference between the first characteristic moment and the second characteristic moment is calculated to obtain a characteristic duration value, wherein the larger the value of the characteristic duration value is, the slower the efficiency of the processing and regulation of the corresponding state abnormality alarm signal is; the characteristic duration value is numerically compared with a preset characteristic duration threshold value. If the characteristic duration value exceeds the corresponding preset characteristic duration threshold value, the corresponding characteristic duration value is marked as a characteristic abnormality value;

[0055] The number of abnormal feature values ​​during the detection period is obtained and marked as a characteristic anomaly value. The characteristic anomaly value is numerically compared with a preset characteristic anomaly threshold. If the characteristic anomaly value exceeds the preset characteristic anomaly threshold, it indicates that the operation status of the visual human-computer interaction interface unit during the detection period is poor and the visual human-computer interaction interface unit needs to be optimized in time, and an interface optimization alarm signal is generated;

[0056] If the characteristic anomaly value does not exceed the preset characteristic anomaly threshold, the number of black screen and freeze times and the total duration of the intelligent human-computer interaction interface unit during the detection period are obtained and marked as unresponsive frequency and unresponsive duration respectively. The unresponsive frequency and unresponsive duration are numerically compared with the preset unresponsive frequency threshold and the preset unresponsive duration threshold respectively. If the unresponsive frequency or the unresponsive duration exceeds the corresponding preset threshold, an interface optimization alarm signal is generated;

[0057] If the unresponsiveness frequency and the unresponsiveness duration do not exceed the corresponding preset thresholds, then when the user performs a click operation on the intelligent human-computer interaction interface unit, the touch buffer duration is collected, and the average of all touch buffer durations within the detection period is calculated to obtain a touch buffer detection value; and when the user performs a drag operation on the intelligent human-computer interaction interface unit, the user's actual drag trajectory and the identified drag trajectory are obtained, the actual drag trajectory and the identified drag trajectory are overlapped and compared, and the length ratio of the non-overlapping trajectory is marked as the trajectory anomaly value, and the average of all trajectory anomalies within the detection period is calculated to obtain a drag anomaly value;

[0058] The interface optimization urgency coefficient is calculated by weighted summing the characteristic anomaly value, the unresponsiveness frequency, the unresponsiveness duration, the touch delay value, and the drag anomaly value. That is, the characteristic anomaly value, the unresponsiveness frequency, the unresponsiveness duration, the touch delay value, and the drag anomaly value are respectively assigned corresponding preset weight coefficients, and the characteristic anomaly value, the unresponsiveness frequency, the unresponsiveness duration, the touch delay value, and the drag anomaly value are respectively multiplied by the corresponding preset weight coefficients, and the sum of the five groups of product results is marked as the interface optimization urgency coefficient. It should be noted that the larger the value of the interface optimization urgency coefficient, the more timely it is necessary to optimize the visual human-computer interaction interface unit.

[0059] The interface optimization urgency coefficient is numerically compared with the preset interface optimization urgency coefficient threshold. If the interface optimization urgency coefficient exceeds the preset interface optimization urgency coefficient threshold, it indicates that the overall operating condition of the visual human-computer interaction interface unit during the detection period is poor, and the visual human-computer interaction interface unit needs to be optimized in time, then an interface optimization alarm signal is generated.

[0060] Example 3: Figure 3 As shown, the difference between this embodiment and the first and second embodiments is that the present invention proposes an intelligent human-computer interaction method based on a multimodal large model, comprising the following steps:

[0061] Step 1: Collect raw data from multiple modalities and preprocess them;

[0062] Step 2: Perform multimodal data integration, intent understanding, decision planning, and instruction generation;

[0063] Step 3: Convert the decision results and instructions into a protocol format understandable to the external device and output it;

[0064] Step 4: Control the corresponding execution operation based on the output information.

[0065] The working principle of the present invention is as follows: when in use, efficient and accurate human-computer interaction is achieved through modular design and tight data flow control, which is particularly suitable for complex scenarios such as underwater operations. The modules cooperate with each other to jointly complete tasks such as multimodal data collection, processing, fusion, decision support, instruction execution and feedback, and improve user experience and operational convenience through a visual human-computer interaction interface unit, providing underwater operation equipment with decision-making support such as autonomous task planning, dynamic environment perception and real-time intelligent control, thereby improving the safety and execution efficiency of underwater operation equipment, and solving the problems of insufficient perception, rigid planning and control delay in traditional human-computer interaction systems in complex scenarios such as underwater operations.

[0066] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The setting of the threshold in the technical solution is for result comparison and analysis in order to determine whether it is good or bad. As for the value of the threshold, it is set based on a combination of large-scale model analysis of sample data and manual experience to enter and store it. It can also be appropriately adjusted based on seasonal or common sense influencing conditions. The preferred embodiment does not describe all the details in detail, nor does it limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made based on the contents of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. An intelligent human-computer interaction system based on a multimodal large model, characterized in that: It includes a multimodal input processing unit, an intelligent analysis and decision-making unit, an interface protocol management unit, and a visual human-computer interaction interface unit; the multimodal input processing unit is responsible for collecting raw data from multiple modalities and preprocessing it, and outputting the preprocessed data to the intelligent analysis and decision-making unit; The intelligent analysis and decision-making unit is responsible for the integration of multimodal data, intention understanding, decision planning and instruction generation, and sends the decision results and instructions to the visual human-computer interaction interface unit and the interface protocol management unit; The interface protocol management unit receives the decision results and instructions of the intelligent analysis and decision unit and converts them into a protocol format that can be understood by external devices. The interface protocol management unit is also responsible for receiving feedback data from external devices and transmitting it to the intelligent analysis and decision unit for status updates and decision adjustments.

2. The intelligent human-computer interaction system based on a multimodal large model according to claim 1, characterized in that: The multimodal input processing module captures ambient sound sources through microphone components, performs noise suppression, echo cancellation and voice enhancement on the collected raw audio, relies on optical sensors to capture scene images, performs digital filtering, grayscale conversion, normalization and resolution adjustment on the raw images, integrates multi-source sensor data including sonar, pressure gauge and thermometer, and performs data fusion and preprocessing on the multi-source sensor data.

3. The intelligent human-computer interaction system based on a multimodal large model according to claim 1, characterized in that: The intelligent analysis and decision-making unit includes a multimodal interaction engine module, an intelligent decision-making support module, and an instruction processing and sequence conversion module; feature extraction and fusion are performed through the multimodal interaction engine module, dynamic strategies and risk predictions are generated using the intelligent decision-making support module, and intentions are converted into executable instructions through the instruction processing and sequence conversion module.

4. The intelligent human-computer interaction system based on a multimodal large model according to claim 1, characterized in that: The interface protocol management unit supports multiple communication protocols including TCP / IP, Modbus, and CAN, enabling seamless communication between the system and external devices. It also provides a protocol conversion engine, real-time monitoring of errors during the protocol conversion process, and automatic repair and manual intervention functions.

5. The intelligent human-computer interaction system based on a multimodal large model according to claim 1, characterized in that: The visual human-computer interaction interface unit has multi-modal input support, real-time data display and interactive operation control functions.

6. The intelligent human-computer interaction system based on a multimodal large model according to claim 1, characterized in that: The intelligent analysis and decision-making unit calls the large model training and tuning module during the inference process. During the system operation, the large model training and tuning module continuously receives model inference requests from the intelligent analysis and decision-making unit. The large model training and tuning module is responsible for the training, tuning and deployment of multimodal large models.

7. The intelligent human-computer interaction system based on a multimodal large model according to claim 6, characterized in that: The intelligent human-computer interaction unit and the large model training and tuning unit send computing power requirements to the hardware adaptation and computing power management unit. The hardware adaptation and computing power management unit provides heterogeneous computing power support and optimizes resource allocation. The specific operation process is as follows: A domestic computing cluster is built based on the Kunpeng 920 CPU and Ascend 910 NPU to meet the computing power requirements of large-scale model inference, dynamically schedule GPU / NPU resources, perform multi-task parallel processing and low-latency response, and dynamically adjust resource allocation according to load. It supports large-scale model inference and real-time control, and provides computing power support for large-scale model training and tuning units and intelligent human-computer interaction units.

8. The intelligent human-computer interaction system based on a multimodal large model according to claim 5, characterized in that: The visual human-computer interaction interface unit is communicatively connected to the interactive interface self-test alarm unit. The interactive interface self-test alarm unit analyzes the operating status of the visual human-computer interaction interface unit in real time and triggers the alarm mechanism when an interface abnormality alarm signal is generated. The specific status analysis process is: if the interface temperature measurement value or the interface temperature rise value exceeds the corresponding preset threshold, a status abnormality alarm signal is generated; if neither the interface temperature measurement value nor the interface temperature rise value exceeds the corresponding preset threshold, the interface display value is numerically compared with the preset interface display threshold. If the interface display value exceeds the preset interface display threshold, a status abnormality alarm signal is generated.

9. The intelligent human-computer interaction system based on a multimodal large model according to claim 8, characterized in that: The interactive interface self-check alarm unit is also used to comprehensively evaluate the operation optimization status of the visual human-computer interactive interface unit during the detection period, and trigger the alarm mechanism when generating an interface optimization alarm signal. The specific evaluation and analysis process is as follows: the number of feature abnormal values ​​during the detection period is obtained and marked as feature anomaly values. If the feature anomaly value exceeds the preset feature anomaly threshold, an interface optimization alarm signal is generated. If the characteristic anomaly value does not exceed the preset characteristic anomaly threshold, the number of black screen and freeze times and the total duration of the intelligent human-computer interaction interface unit during the detection period are obtained and marked as unresponsive frequency and unresponsive duration respectively. If the unresponsive frequency or unresponsive duration exceeds the corresponding preset threshold, an interface optimization alarm signal is generated; if the unresponsive frequency and unresponsive duration do not exceed the corresponding preset threshold, the interface optimization urgency coefficient is obtained by weighted summing the characteristic anomaly value, unresponsive frequency, unresponsive duration, touch delay value and drag recognition value. If the interface optimization urgency coefficient exceeds the preset interface optimization urgency coefficient threshold, an interface optimization alarm signal is generated.

10. An intelligent human-computer interaction method based on a multimodal large model, characterized in that: The intelligent human-computer interaction method adopts the intelligent human-computer interaction system based on the multimodal large model as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Portable electronic device and touch operation method thereof

    CN105791667A

  • Data processing method for realizing multi-modal interaction and multi-modal interaction system

    CN105843381A

  • Intelligent man-machine interaction device and method

    CN118607641A

  • Robot behavior control system capable of dynamically adapting to environment

    CN118876067A

  • Intelligent operation equipment operation monitoring management and control system based on artificial intelligence

    CN119046844A

Cited By

  • Semantic communication protocol-based end-cloud cooperation method and underwater vehicle-mounted intelligent operation command system

    CN122027691A