Intelligent human-computer interaction method and system based on multi-modal large model
By using an intelligent human-machine interaction system based on a multimodal large model, the problems of insufficient perception and control delay in traditional human-machine interaction systems during underwater operations have been solved. This system enables efficient and accurate multimodal data processing and real-time decision-making, improving equipment safety and execution efficiency while reducing management difficulty.
Patent Information
- Application Number
- CN202510875954.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Traditional human-computer interaction systems suffer from insufficient dynamic environment perception, rigid task planning, and real-time control delays in extreme environments such as underwater operations. This makes it difficult to achieve accurate intent understanding and efficient decision-making, and also makes it difficult to effectively monitor the operation of interactive display devices and issue alarms for abnormalities, thus affecting the safety and execution efficiency of the equipment.
An intelligent human-computer interaction system based on a multimodal large model is adopted, including multimodal input processing, intelligent parsing and decision-making, interface protocol management and visualization human-computer interaction interface units. Through multimodal data fusion, intent understanding, decision planning and real-time monitoring, it provides autonomous task planning, dynamic environment perception and real-time control support, and monitors the operation status through the interactive interface self-check alarm unit to generate abnormal alarm signals.
It achieves efficient and accurate human-computer interaction, improves the safety and execution efficiency of underwater operation equipment, reduces the management difficulty of interactive display devices, and enhances user experience and ease of operation.
Smart Images

Figure CN120704534B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction devices, in particular to an intelligent human-computer interaction method and system based on a multi-modal large model. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, intelligent human-computer interaction systems have become the key to improving the autonomy and operation efficiency of devices. However, traditional human-computer interaction systems are often limited by single-modal data input, making it difficult to achieve accurate intent understanding and efficient decision-making in complex environments.
[0003] In particular, in extreme environments such as underwater operations, traditional human-computer interaction systems face problems such as insufficient dynamic environment perception, rigid task planning, and real-time control delay, which seriously restrict the safety and execution efficiency of underwater operation devices. Moreover, it is difficult to effectively monitor and alarm the operation of the interactive display device, which is not conducive to ensuring the performance of the interactive display and reducing its management difficulty, and the level of intelligence is low.
[0004] In view of the above technical defects, a solution is proposed. SUMMARY
[0005] The present application aims to provide an intelligent human-computer interaction method and system based on a multi-modal large model, which solves the problem that the prior art is limited by single-modal data input, making it difficult to achieve accurate intent understanding and efficient decision-making in complex environments, which seriously restricts the safety and execution efficiency of underwater operation devices, and it is difficult to effectively monitor and alarm the operation of the interactive display device, which is not conducive to ensuring the performance of the interactive display and reducing its management difficulty.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical scheme:
[0007] An intelligent human-computer interaction system based on a multi-modal large model, comprising a multi-modal input processing unit, an intelligent analysis and decision unit, an interface protocol management unit, and a visual human-computer interaction interface unit; the multi-modal input processing unit is responsible for collecting raw data from multiple modalities and preprocessing it, and outputs the preprocessed data to the intelligent analysis and decision unit;
[0008] The intelligent analysis and decision unit is responsible for the fusion of multi-modal data, intent understanding, decision planning, and instruction generation, and sends the decision results and instructions to the visual human-computer interaction interface unit and the interface protocol management unit;
[0009] The interface protocol management unit receives the decision results and instructions of the intelligent analysis and decision unit, converts them into a protocol format that can be understood by external devices, and transmits them to underwater operation devices for execution. The interface protocol management unit is also responsible for receiving feedback data from external devices and transmitting it to the intelligent analysis and decision unit for state updating and decision adjustment.
[0010] Further, the multi-modal input processing module captures environmental sound sources through the microphone assembly, performs noise suppression, echo cancellation, and speech enhancement processing on the collected raw audio, and relies on optical sensors to capture scene images, performs digital filtering, grayscale conversion, normalization, and resolution adjustment processing on the raw images, and integrates multi-source sensor data including sonar, pressure gauges, and thermometers, and performs data fusion and preprocessing on the multi-source sensor data.
[0011] Further, the intelligent analysis and decision unit includes a multi-modal interaction engine module, an intelligent auxiliary decision module, and an instruction processing and sequence conversion module; feature extraction and fusion are performed through the multi-modal interaction engine module, dynamic strategies and risk prediction are generated using the intelligent auxiliary decision module, and the intent is converted into executable instructions through the instruction processing and sequence conversion module.
[0012] Further, the interface protocol management unit supports multiple communication protocols including TCP / IP, Modbus, and CAN, enabling seamless communication between the system and external devices, adapting to the communication needs of different devices, and providing a protocol conversion engine for data conversion and adaptation between different protocols, as well as real-time monitoring of errors during protocol conversion, providing automatic repair and manual intervention functions.
[0013] Further, the visual human-computer interaction interface unit has multi-modal input support, real-time data display, and interactive operation control functions, wherein the multi-modal input support provides multiple input methods including voice, text, and gestures, the real-time data display displays sensor data, task progress, and device status information of underwater operation equipment in the form of charts and dashboards, and the interactive operation control provides interactive tools such as drag-and-drop task builders and parameter adjustment panels, supporting user-defined task steps and parameter settings.
[0014] Further, the intelligent analysis and decision unit calls the large model training and optimization module during inference and operation, which continuously receives model inference requests from the intelligent analysis and decision unit during system operation; the large model training and optimization module is responsible for the training, optimization, and deployment of multi-modal large models, performs cleaning, labeling, and vectorization processing on received data to generate standardized data sets; adopts pre-training and fine-tuning strategies, trains multi-modal large models using large-scale data sets, and improves model performance in specific business domains through LoRA training and optimization and adaptive hyperparameter optimization techniques, optimizes inference speed and resource consumption, and outputs optimized model parameters to the intelligent analysis and decision unit.
[0015] Further, the intelligent human-computer interaction unit and the large model training and optimization unit send the computing power demand to the hardware adaptation and computing power management unit, the hardware adaptation and computing power management unit provides heterogeneous computing power support, optimizes resource allocation, and the specific operation process is as follows:
[0016] Based on the Kunpeng 920 CPU and the Ascend 910 NPU, a domestic algorithm cluster is constructed to meet the computing power demand of large model inference, dynamically schedule GPU / NPU resources, perform multi-task parallel processing and low-latency response, improve the overall performance of the system, and dynamically adjust resource allocation according to the load, support large model inference and real-time control, and provide computing power support for the large model training and optimization unit and the intelligent human-computer interaction unit.
[0017] Further, the visual human-computer interaction interface unit is communicatively connected to the interactive interface self-checking alarm unit, the interactive interface self-checking alarm unit monitors the operation of the visual human-computer interaction interface unit, analyzes the running state of the visual human-computer interaction interface unit in real time, and judges whether to generate a state abnormal alarm signal according to the running state, and triggers the alarm mechanism when the interface abnormal alarm signal is generated; the specific state analysis process is as follows:
[0018] The real-time interface temperature of the visual human-computer interaction interface unit is obtained and marked as an interface temperature measurement value, and the growth rate of the real-time interface temperature is marked as an interface temperature rise value, the interface temperature measurement value and the interface temperature rise value are compared with the preset interface temperature measurement threshold and the preset interface temperature rise threshold respectively, if the interface temperature measurement value or the interface temperature rise value exceeds the corresponding preset threshold, a state abnormal alarm signal is generated;
[0019] If the interface temperature measurement value and the interface temperature rise value do not exceed the corresponding preset threshold, the display brightness of the visual human-computer interaction interface unit is collected, the display brightness is compared with the standard display brightness, and the absolute value is obtained to obtain the interface display value, the interface display value is compared with the preset interface display threshold, if the interface display value exceeds the preset interface display threshold, a state abnormal alarm signal is generated.
[0020] Further, the interactive interface self-checking alarm unit is also used to set a detection period with a time length of L1, when the time length reaches L1, the running optimization condition of the visual human-computer interaction interface unit in the detection period is comprehensively evaluated, and it is judged whether to generate an interface optimization alarm signal according to the running optimization condition, and the alarm mechanism is triggered when the interface optimization alarm signal is generated; the specific evaluation and analysis process is as follows:
[0021] The generation time of the corresponding state abnormal alarm signal is acquired and marked as a first feature time, and the time when the corresponding state abnormal alarm is removed after being regulated is acquired and marked as a second feature time. The first feature time and the second feature time are calculated to obtain a feature duration value, and the feature duration value is compared with a preset feature duration threshold value. If the feature duration value exceeds the corresponding preset feature duration threshold value, the corresponding feature duration value is marked as a feature abnormal value.
[0022] The number of feature abnormal values in the detection period is acquired and marked as a feature abnormal value, and the feature abnormal value is compared with a preset feature abnormal value threshold value. If the feature abnormal value exceeds the preset feature abnormal value threshold value, an interface optimization alarm signal is generated. If the feature abnormal value does not exceed the preset feature abnormal value threshold value, the number of times and the total time length of the black screen and the lag of the intelligent man-machine interaction interface unit in the detection period are acquired and marked as a non-response frequency and a non-response time length, respectively. The non-response frequency and the non-response time length are compared with a preset non-response frequency threshold value and a preset non-response time length threshold value, respectively. If the non-response frequency or the non-response time length exceeds the corresponding preset threshold value, an interface optimization alarm signal is generated.
[0023] If the non-response frequency and the non-response time length do not exceed the corresponding preset threshold value, when the user performs a click operation on the intelligent man-machine interaction interface unit, the touch buffer time length is collected. The average value of all touch buffer time lengths in the detection period is calculated to obtain a touch buffer value. When the user performs a drag operation on the intelligent man-machine interaction interface unit, the actual drag trajectory and the recognized drag trajectory of the user are acquired. The actual drag trajectory and the recognized drag trajectory are compared, and the length proportion of the non-coincidence trajectory is marked as a trajectory abnormal occupation value. The average value of all trajectory abnormal occupation values in the detection period is calculated to obtain a drag recognition abnormal value.
[0024] The feature abnormal value, the non-response frequency, the non-response time length, the touch buffer value and the drag recognition abnormal value are weighted and summed to obtain an interface optimization urgency coefficient. The interface optimization urgency coefficient is compared with a preset interface optimization urgency coefficient threshold value. If the interface optimization urgency coefficient exceeds the preset interface optimization urgency coefficient threshold value, an interface optimization alarm signal is generated.
[0025] Further, the application also provides an intelligent man-machine interaction method based on a multi-modal large model, which comprises the following steps:
[0026] Step 1: Collecting original data from multiple modes and preprocessing the original data;
[0027] Step 2: Fusion, intent understanding, decision planning and instruction generation of multi-modal data;
[0028] Step three, converting the decision result and the instruction into a protocol format that can be understood by the external device and outputting;
[0029] Step four, controlling the corresponding execution operation based on the output information.
[0030] Compared with the prior art, the beneficial effects of the present application are:
[0031] 1、In the present application, through the modular design and close data flow control, efficient and accurate human-computer interaction is realized, and each module cooperates with each other to complete the multi-modal data acquisition, processing, fusion, decision support, instruction execution and feedback tasks, to provide autonomous task planning, dynamic environment perception and real-time intelligent control decision auxiliary support for underwater operation equipment, and improve the safety and execution efficiency of the underwater operation equipment.
[0032] 2、In the present application, the operation of the visual human-computer interaction interface unit is monitored by the interactive interface self-checking alarm unit, and the running state of the visual human-computer interaction interface unit is analyzed in real time, corresponding adjustment measures are taken when the interface exception alarm signal is generated, and the running optimization condition of the visual human-computer interaction interface unit in the detection period is comprehensively evaluated, corresponding improvement optimization measures are taken when the interface optimization alarm signal is generated, which is beneficial to improve the use effect and running stability of the visual human-computer interaction interface unit, and significantly reduce the operation and management difficulty. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to facilitate those skilled in the art to understand, the present application will be further described below in conjunction with the drawings;
[0034] Figure 1 The system block diagram of the first embodiment in the present application is shown in the figure;
[0035] Figure 2 The system block diagram of the second embodiment in the present application is shown in the figure;
[0036] Figure 3 The method flow chart of the third embodiment in the present application is shown in the figure. DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0038] Embodiment one: as Figure 1As shown, the intelligent man-machine interaction system based on a multi-modal large model is provided, which comprises a multi-modal input processing unit, an intelligent analysis and decision unit, an interface protocol management unit and a visual man-machine interaction interface unit, and is mainly used for solving the problems of insufficient perception, rigid planning and control delay of the traditional man-machine interaction system in complex scenes such as underwater operation, realizing deep analysis of user intention and context understanding by integrating multi-modal data, improving the safety and execution efficiency of underwater operation equipment, and improving user experience and operation convenience through the visual interaction interface of the user end.
[0039] The multi-modal input processing unit is responsible for collecting raw data from multiple modalities and pre-processing the raw data, and outputs the pre-processed data to the intelligent analysis and decision unit; for example, an environmental sound source is captured through a microphone assembly, and the collected raw audio is processed for noise suppression, echo cancellation and voice enhancement, etc., to improve the quality of the audio signal and provide high-quality input for voice interaction;
[0040] And relying on an optical sensor to capture scene pictures, the original images are processed for digital filtering, grayscale conversion, normalization and resolution adjustment, etc., to improve the quality of the images and provide high-quality input for visual interaction, and multi-source sensor data including sonar, pressure gauge and thermometer, etc. are integrated, and the multi-source sensor data is processed for data fusion and pre-processing to provide comprehensive environmental perception capability for the system.
[0041] The intelligent analysis and decision unit is responsible for the fusion of multi-modal data, intention understanding, decision planning and instruction generation, and sends the decision results and instructions to the visual man-machine interaction interface unit and the interface protocol management unit; it should be noted that the intelligent analysis and decision unit comprises a multi-modal interaction engine module, an intelligent auxiliary decision module and an instruction processing and sequence conversion module.
[0042] Specifically, feature extraction and fusion are performed through the multi-modal interaction engine module (integrating voice, vision, text and other multi-modal inputs, realizing deep analysis of user intention and context understanding, and the analysis process: through a pre-trained large model, multi-modal data are extracted and fused, key information is captured by using an attention mechanism, dynamic strategies and risk prediction are generated, and the intelligent auxiliary decision module is provided with input), and dynamic strategies and risk prediction are generated by using the intelligent auxiliary decision module (providing real-time decision suggestions and risk warnings, optimizing task operation processes, and the analysis process: based on data-driven and model inference, optimized strategies are generated by analyzing business dynamics, and risk prediction is performed in combination with a knowledge base).
[0043] and converting the intention into executable instructions through the instruction processing and sequence conversion module (converting voice instructions into structured control signals, multi-step timing reasoning and execution assistance for complex tasks, analysis process: through natural language understanding, parameter extraction and verification, multi-modal instruction fusion, etc. Technology, accurately disassembles voice instructions into atomic operations, and generates timing and spatial control signals).
[0044] The interface protocol management unit receives the decision results and instructions of the intelligent analysis and decision unit, converts them into a protocol format that can be understood by external devices, and transmits them to underwater operation equipment for execution. The interface protocol management unit is also responsible for receiving feedback data from external devices and transmitting it to the intelligent analysis and decision unit for state updating and decision adjustment.
[0045] It should be noted that the interface protocol management unit supports multiple communication protocols such as TCP / IP, Modbus, and CAN, enabling seamless communication between the system and external devices, adapting to the communication needs of different devices, and providing a protocol conversion engine to achieve data conversion and adaptation between different protocols, solving protocol compatibility problems. Real-time monitoring of errors during protocol conversion, providing automatic repair and manual intervention functions.
[0046] The visual human-computer interaction interface unit (touch display) has multi-modal input support, real-time data display, and interactive operation control functions. Multi-modal input support provides multiple input methods including voice, text, and gestures. Real-time data display displays sensor data, task progress, and device status information of underwater operation equipment in real time through charts and dashboards. Interactive operation control provides interactive tools such as drag-and-drop task builders and parameter adjustment panels, allowing users to customize task steps and parameter settings.
[0047] In addition, the intelligent analysis and decision unit calls the large model training and optimization module during reasoning and execution. The large model training and optimization module continuously receives model reasoning requests from the intelligent analysis and decision unit during system operation. The intelligent human-computer interaction unit and the large model training and optimization unit send computing power requirements to the hardware adaptation and computing power management unit. The hardware adaptation and computing power management unit provides heterogeneous computing power support and optimizes resource allocation.
[0048] The large model training and optimization module is responsible for training, optimizing, and deploying multi-modal large models. It cleans, labels, and vectorizes the received data to generate standardized data sets. It uses pre-training and fine-tuning strategies to train multi-modal large models using large-scale data sets to improve model generalization. Through LoRA training optimization and adaptive hyperparameter optimization, the model's performance in specific business domains is improved, and the inference speed and resource consumption are optimized. The optimized model parameters are output to the intelligent analysis and decision unit.
[0049] The hardware adaptation and computing power management unit is built on the Kunpeng 920 CPU and Ascend 910 NPU to create a domestic computing power cluster that meets the computing power requirements of large model inference. It dynamically schedules GPU / NPU resources to perform multi-task parallel processing and low-latency response, improving the overall system performance. It also dynamically adjusts resource allocation according to the load, supports large model inference and real-time control, and ensures the efficient and stable operation of the system. It provides computing power support for the large model training and optimization unit and the intelligent human-computer interaction unit.
[0050] Example 2: Figure 2 As shown, the difference between this embodiment and Embodiment 1 is that the visual human-computer interaction interface unit is communicatively connected to the interface self-test alarm unit. The interface self-test alarm unit monitors the operation of the visual human-computer interaction interface unit, analyzes its operating status in real time, and determines whether to generate an abnormal status alarm signal. When an abnormal status alarm signal is generated, an alarm mechanism is triggered to automatically take corresponding adjustment measures, reduce the operational risk of the visual human-computer interaction interface unit, and ensure its effectiveness. The specific status analysis process is as follows:
[0051] The real-time interface temperature of the visual human-computer interaction interface unit is obtained and marked as the interface temperature measurement value. The growth rate of the real-time interface temperature is marked as the interface temperature rise value. The interface temperature measurement value and the interface temperature rise value are compared with the preset interface temperature measurement threshold and the preset interface temperature rise threshold respectively. If the interface temperature measurement value or the interface temperature rise value exceeds the corresponding preset threshold, it indicates that the temperature condition of the visual human-computer interaction interface unit is not good and there is a significant safety risk. Then, an abnormal status alarm signal is generated.
[0052] If the interface temperature measurement value and the interface temperature rise value do not exceed the corresponding preset threshold, it indicates that the temperature risk of the visual human-computer interaction interface unit is small. Then, the display brightness of the visual human-computer interaction interface unit is collected, and the difference between the display brightness and the set standard display brightness is calculated and the absolute value is taken to obtain the interface display value. The interface display value is compared with the preset interface display threshold. If the interface display value exceeds the preset interface display threshold, it indicates that the display performance of the visual human-computer interaction interface unit is poor, and an abnormal status alarm signal is generated.
[0053] Furthermore, the interactive interface self-test alarm unit is also used to set a detection period of duration L1. Preferably, L1 is four hours (eight hours). The unit comprehensively evaluates the operational optimization status of the visual human-computer interaction interface unit during the detection period to determine whether to generate an interface optimization alarm signal. When an interface optimization alarm signal is generated, an alarm mechanism is triggered to promptly implement corresponding improvement and optimization measures. This further enhances the usability and operational stability of the visual human-computer interaction interface unit, significantly reduces its operational management difficulty, and demonstrates a high level of intelligence. The specific evaluation and analysis process is as follows:
[0054] The generation time of the corresponding state abnormal alarm signal is obtained and marked as the first feature time, and the time when the corresponding state abnormal alarm is released after being regulated is collected and marked as the second feature time. The first feature time and the second feature time are time difference calculated to obtain the feature duration value. The greater the value of the feature duration value, the slower the processing and regulation efficiency for the corresponding state abnormal alarm signal. The feature duration value is compared with the preset feature duration threshold value. If the feature duration value exceeds the corresponding preset feature duration threshold value, the corresponding feature duration value is marked as a feature abnormal value.
[0055] The number of feature abnormal values in the detection period is obtained and marked as a feature anomaly value. The feature anomaly value is compared with the preset feature anomaly threshold value. If the feature anomaly value exceeds the preset feature anomaly threshold value, it indicates that the running state of the visual human-computer interaction interface unit in the detection period is poor, and the visual human-computer interaction interface unit needs to be optimized in a timely manner. An interface optimization alarm signal is generated.
[0056] If the feature anomaly value does not exceed the preset feature anomaly threshold value, the number of times and the total time length of the black screen and the lag of the intelligent human-computer interaction interface unit in the detection period are obtained and marked as the non-response frequency and the non-response time length, respectively. The non-response frequency and the non-response time length are compared with the preset non-response frequency threshold value and the preset non-response time length threshold value, respectively. If the non-response frequency or the non-response time length exceeds the corresponding preset threshold value, an interface optimization alarm signal is generated.
[0057] If the non-response frequency and the non-response time length do not exceed the corresponding preset threshold value, the touch buffer time length of the user when performing a click operation on the intelligent human-computer interaction interface unit is collected. The average value of all touch buffer time lengths in the detection period is calculated to obtain the touch buffer value. When the user performs a drag operation on the intelligent human-computer interaction interface unit, the actual drag trajectory of the user and the recognized drag trajectory are obtained. The actual drag trajectory and the recognized drag trajectory are compared, and the length proportion of the non-coincidence trajectory is marked as a trajectory anomaly value. The average value of all trajectory anomaly values in the detection period is calculated to obtain a drag recognition anomaly value.
[0058] The interface optimization urgency coefficient is calculated by weighted sum of the feature trace difference value, the non-response frequency, the non-response duration, the touch delay value and the drag recognition difference value, that is, the feature trace difference value, the non-response frequency, the non-response duration, the touch delay value and the drag recognition difference value are respectively assigned with corresponding preset weight coefficients, and the feature trace difference value, the non-response frequency, the non-response duration, the touch delay value and the drag recognition difference value are respectively multiplied by the corresponding preset weight coefficients, and the sum of the five groups of product results is marked as the interface optimization urgency coefficient; it should be noted that the greater the value of the interface optimization urgency coefficient, the more the need to optimize the visual human-computer interaction interface unit in time;
[0059] The interface optimization urgency coefficient is compared with the preset interface optimization urgency coefficient threshold value, if the interface optimization urgency coefficient exceeds the preset interface optimization urgency coefficient threshold value, it indicates that the running condition of the visual human-computer interaction interface unit in the detection period is poor in general, and the visual human-computer interaction interface unit needs to be optimized in time at present, and an interface optimization alarm signal is generated.
[0060] Embodiment three: as shown in the embodiment, the embodiment is different from embodiment one and embodiment two, and the intelligent human-computer interaction method based on a multi-modal large model comprises the following steps: Figure 3
[0061] Step one, collecting original data from multiple modes and pre-processing the original data;
[0062] Step two, fusion of multi-modal data, intent understanding, decision planning and instruction generation;
[0063] Step three, converting the decision result and the instruction into a protocol format that can be understood by an external device and outputting;
[0064] Step four, controlling the corresponding execution operation based on the output information.
[0065] The working principle of the application: in use, through modular design and close data flow control, efficient and accurate human-computer interaction is realized, which is especially suitable for complex scenes such as underwater operation, and each module cooperates with each other to complete tasks such as multi-modal data acquisition, processing, fusion, decision support, instruction execution and feedback, and through the visual human-computer interaction interface unit, user experience and operation convenience are improved, autonomous task planning, dynamic environment perception and real-time intelligent control decision auxiliary support are provided for underwater operation equipment, the safety and execution efficiency of the underwater operation equipment are improved, and the problems of insufficient perception, rigid planning and delayed control of the traditional human-computer interaction system in complex scenes such as underwater operation are solved.
[0066] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The setting of the threshold in the technical solution is for result comparison analysis, so as to determine whether it is good or not. The size of the threshold is set according to the large model analysis of sample data and the combination of artificial experience to enter storage. It can also be appropriately adjusted through seasonal or reasonable influence conditions. The preferred embodiments do not describe all the details and are not limited to the specific implementation. Obviously, according to the content of the specification, many modifications and changes can be made. The specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is limited by the claims and their entire scope and equivalents.
Claims
1. An intelligent human-computer interaction system based on a multi-modal large model, characterized in that, The system comprises a multi-modal input processing unit, an intelligent analysis and decision unit, an interface protocol management unit, and a visual human-computer interaction interface unit. The multi-modal input processing unit is responsible for collecting raw data from multiple modalities and pre-processing the data, and outputs the pre-processed data to the intelligent analysis and decision unit. The intelligent analysis and decision unit is responsible for multi-modal data fusion, intent understanding, decision planning, and instruction generation, and sends the decision results and instructions to the visual human-computer interaction interface unit and the interface protocol management unit. The interface protocol management unit receives the decision results and instructions from the intelligent analysis and decision unit, converts them into a protocol format that can be understood by external devices, and is also responsible for receiving feedback data from external devices and transmitting it to the intelligent analysis and decision unit for state updating and decision adjustment. The visual human-computer interaction interface unit is connected to an interactive interface self-checking alarm unit, which monitors the operation of the visual human-computer interaction interface unit, analyzes the running state of the visual human-computer interaction interface unit in real time, and judges whether to generate a state abnormal alarm signal based on this. When the interface abnormal alarm signal is generated, the alarm mechanism is triggered. The specific state analysis process is as follows: The real-time interface temperature of the visual human-computer interaction interface unit is obtained and marked as the interface temperature measurement value, and the growth rate of the real-time interface temperature is marked as the interface temperature rise value. The interface temperature measurement value and the interface temperature rise value are compared with the preset interface temperature measurement threshold and the preset interface temperature rise threshold, respectively. If the interface temperature measurement value or the interface temperature rise value exceeds the corresponding preset threshold, a state abnormal alarm signal is generated. If the interface temperature measurement value and the interface temperature rise value do not exceed the corresponding preset threshold, the display brightness of the visual human-computer interaction interface unit is collected, the difference between the display brightness and the standard display brightness is calculated and the absolute value is taken to obtain the interface display value, and the interface display value is compared with the preset interface display threshold. If the interface display value exceeds the preset interface display threshold, a state abnormal alarm signal is generated. The interactive interface self-checking alarm unit is also used to set a detection period with a time length of L1. When the time length reaches L1, the running optimization status of the visual human-computer interaction interface unit during the detection period is comprehensively evaluated, and it is judged whether to generate an interface optimization alarm signal based on this. When the interface optimization alarm signal is generated, the alarm mechanism is triggered. The specific evaluation and analysis process is as follows: The generation time of the corresponding state abnormal alarm signal is obtained and marked as the first feature time, and the time when the corresponding state abnormal alarm is removed after regulation is collected and marked as the second feature time. The first feature time and the second feature time are calculated to obtain the feature duration value, and the feature duration value is compared with the preset feature duration threshold. If the feature duration value exceeds the corresponding preset feature duration threshold, the corresponding feature duration value is marked as a feature abnormal value. The number of feature abnormal values during the detection period is obtained and marked as a feature abnormal value. The feature abnormal value is compared with the preset feature abnormal threshold. If the feature abnormal value exceeds the preset feature abnormal threshold, an interface optimization alarm signal is generated. If the feature trace value does not exceed the preset feature trace threshold, the number of times and the total duration of the black screen and the lag of the intelligent human-computer interaction interface unit during the detection period are obtained and are marked as the non-response frequency and the non-response duration respectively, and the non-response frequency and the non-response duration are compared with the preset non-response frequency threshold and the preset non-response duration threshold respectively, if the non-response frequency or the non-response duration exceeds the corresponding preset threshold, an interface optimization alarm signal is generated; If the non-response frequency and the non-response duration do not exceed the corresponding preset threshold, when the user performs a click operation on the intelligent human-computer interaction interface unit, the touch buffer duration is collected, the average of all touch buffer durations during the detection period is calculated, and the touch buffer value is obtained accordingly; and when the user performs a drag operation on the intelligent human-computer interaction interface unit, the actual drag trajectory and the recognized drag trajectory of the user are obtained, the actual drag trajectory and the recognized drag trajectory are compared, the length proportion of the non-overlapping trajectory is marked as the trajectory occupation value, and the average of all trajectory occupation values during the detection period is calculated to obtain the drag recognition difference value; The interface optimization urgency coefficient is calculated by weighted summation of the feature trace value, the non-response frequency, the non-response duration, the touch buffer value and the drag recognition difference value, and the interface optimization urgency coefficient is compared with the preset interface optimization urgency coefficient threshold, if the interface optimization urgency coefficient exceeds the preset interface optimization urgency coefficient threshold, an interface optimization alarm signal is generated.
2. The intelligent human-computer interaction system based on a multi-modal large model according to claim 1, characterized in that, The multi-modal input processing module captures environmental sound sources through the microphone assembly, performs noise suppression, echo cancellation and speech enhancement processing on the collected raw audio, and relies on the optical sensor to capture scene pictures, performs digital filtering, grayscale conversion, normalization and resolution adjustment processing on the raw images, and integrates multi-source sensor data including sonar, pressure gauge and thermometer, and performs data fusion and preprocessing on the multi-source sensor data.
3. The intelligent human-computer interaction system based on a multi-modal large model according to claim 1, characterized in that, The intelligent analysis and decision unit includes a multi-modal interaction engine module, an intelligent auxiliary decision module and an instruction processing and sequence conversion module; feature extraction and fusion are performed through the multi-modal interaction engine module, dynamic strategies and risk prediction are generated using the intelligent auxiliary decision module, and the intention is converted into executable instructions through the instruction processing and sequence conversion module.
4. The intelligent human-computer interaction system based on a multi-modal large model according to claim 1, characterized in that, The interface protocol management unit supports multiple communication protocols including TCP / IP, Modbus and CAN, enabling seamless communication between the system and external devices, and provides a protocol conversion engine and real-time monitoring of errors during protocol conversion, as well as automatic repair and manual intervention functions.
5. The intelligent human-computer interaction system based on a multi-modal large model according to claim 1, characterized in that, The visual human-computer interaction interface unit has multi-modal input support, real-time data display and interactive operation control functions.
6. The intelligent human-computer interaction system based on a multi-modal large model according to claim 1, characterized in that, The intelligent analysis and decision unit calls the large model training and optimization module during inference and operation, the large model training and optimization module continuously receives model inference requests from the intelligent analysis and decision unit during system operation, and the large model training and optimization module is responsible for the training, optimization and deployment of multi-modal large models.
7. The intelligent human-computer interaction system based on a multi-modal large model according to claim 6, characterized in that, The intelligent human-computer interaction unit and the large model training and optimization unit send computing power requirements to the hardware adaptation and computing power management unit, the hardware adaptation and computing power management unit provides heterogeneous computing power support, optimizes resource allocation, and the specific operation process is as follows: Based on Kunpeng 920 CPU and Ascend 910 NPU, a domestic computing power cluster is constructed to meet the computing power requirements of large model inference. GPU / NPU resources are dynamically scheduled for multi-task parallel processing and low-latency response. According to the load, the resource allocation is dynamically adjusted to support large model inference and real-time control, providing computing power support for the large model training and optimization unit and the intelligent human-computer interaction unit.
8. An intelligent human-computer interaction method based on a multi-modal large model, characterized in that, The intelligent human-computer interaction method adopts the intelligent human-computer interaction system based on the multi-modal large model as claimed in any one of claims 1-7.
Citation Information
Patent Citations
Intelligent man-machine interaction device and method
CN118607641A
Robot behavior control system capable of dynamically adapting to environment
CN118876067A