Human-computer interaction method, system and equipment based on inverse reinforcement learning and medium

By adopting an inverse reinforcement learning training machine learning model in the human-computer interaction system and combining the user's feedback information to adjust the model, the problem that the existing system cannot meet user needs is solved, and the accuracy of the detection results and user experience are improved.

CN120124770APending Publication Date: 2025-06-10LINGYANGE SEMICONDUCTOR, INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510017563.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing human-computer interaction system is difficult to meet the actual needs of users, and users cannot actively select and adjust machine learning models to meet the needs of specific tasks.

Method used

Using a human-computer interaction method based on inverse reinforcement learning, the system displays multiple machine learning models trained by inverse reinforcement learning and their performance and parameters. Users can filter and select appropriate models and parameters according to their needs. The system adjusts the model based on user feedback information to obtain the final detection result.

Benefits of technology

It realizes that users actively select and adjust machine learning models to meet the needs of specific tasks, and improves the accuracy and user experience of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124770A_ABST
    Figure CN120124770A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of man-machine interaction, in particular to a man-machine interaction method, system and device based on inverse reinforcement learning and a medium. After a user inputs task data and a detection task into the man-machine interaction system, the system analyzes each machine learning model matched with the task data and the detection task according to the task data and the detection task. A user screens out a pre-selected model capable of meeting the actual demand from each machine learning model according to the actual demand and the performance of the machine learning model. The system processes task data based on the pre-selected model loaded with the model parameters to preliminarily obtain a detection result of the detection task, a user makes corresponding feedback information based on the detection result, and the system can adjust the model and the parameters thereof according to the feedback information to obtain a final detection result. According to the model, in the process of outputting the detection result, the user interacts with the system for many times to improve the final detection result, so that the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and in particular, to a human-computer interaction method, system, device and medium based on inverse reinforcement learning. Background Art

[0002] A machine learning model obtained based on reinforcement learning is implanted into a human-computer interaction system. When a user inputs task data and a detection task for the task data into the human-computer interaction system, the human-computer interaction system uses the machine learning model inside it to process the task data to complete the detection task and obtain a detection result of the detection task. Different machine learning models have different performances, and there are certain differences in the results obtained by different performance machine learning models when processing task data. For example, the accuracy and speed of different performance machine learning models when processing task data are also different. However, the existing human-computer interaction system only allows users to passively use the machine learning model inside the system to process task data.

[0003] In summary, the existing human-computer interaction system is difficult to meet the actual needs of users.

[0004] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a human-computer interaction method, system, device and medium based on inverse reinforcement learning, which solves the problem that the existing human-computer interaction system is difficult to meet the actual needs of users.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides a human-computer interaction method based on inverse reinforcement learning, which includes:

[0008] Obtain task data and a detection task input by a user terminal, and display the machine learning model required to complete the detection task, the performance of the machine learning model, and the model parameters. The machine learning model is a model after being trained by inverse reinforcement learning;

[0009] Obtain the preselected model input by the user terminal and the model parameters of the preselected model, apply the preselected model after loading the model parameters to the task data to obtain a detection result of the detection task. The preselected model is a model selected by the user terminal from the machine learning models according to the performance of the machine learning model;

[0010] Obtain feedback information of the user terminal for the detection result, and obtain a final detection result based on the feedback information.

[0011] In one implementation, the machine learning model required for the display to complete the detection task, and the performance and model parameters of the machine learning model include:

[0012] Split the detection task into subtasks, and determine the machine learning submodels required to complete the subtasks. The machine learning submodels are models trained using inverse reinforcement learning.

[0013] Use the machine learning submodels as the machine learning model, and display the performance and the corresponding machine learning submodels and model parameters of the machine learning submodels in the order of the quality of the performance.

[0014] In one implementation, applying the preselected model after loading the model parameters to the task data to obtain the detection result of the detection task includes:

[0015] Determine various modal data in the task data, and obtain the processing duration required for processing each piece of modal data.

[0016] According to the processing duration, divide the modal data into a first data group and a second data group, where the processing duration of the first data group is greater than that of the second data group.

[0017] Apply the preselected model after loading the model parameters to each piece of modal data in the first data group, and process various pieces of modal data in the first data group in parallel to obtain the first processing result of the first data group.

[0018] Within a set duration before the end of the parallel processing, apply the preselected model after loading the model parameters to each piece of modal data in the second data group, and process various pieces of modal data in the second data group serially to obtain the second processing result of the second data group.

[0019] Obtain the detection result of the detection task according to the first processing result and the second processing result.

[0020] In one implementation, obtaining the final detection result according to the feedback information includes:

[0021] When the feedback information is negative feedback, replace the model parameters, load the replaced model parameters into the preselected model, and reapply the preselected model to the task data to adjust the detection result to obtain the final detection result.

[0022] In one implementation, obtaining the final detection result according to the feedback information includes:

[0023] When the feedback information is negative feedback, real-time training samples are obtained by retrieving an enhanced generation algorithm;

[0024] The preselected model is retrained using the real-time training samples;

[0025] The preselected model after retraining is applied to the task data to obtain the final detection result.

[0026] In one implementation, the learning method of the machine learning model includes:

[0027] Determine each machine learning submodel in the machine learning model;

[0028] Apply inverse reinforcement learning to each machine learning submodel to obtain the reward subfunction of each machine learning submodel;

[0029] Obtain a global reward function based on the reward subfunction of each machine learning submodel;

[0030] Apply reinforcement learning to the machine learning model based on the global reward function.

[0031] In one implementation, the task data is a human electrocardiogram signal, and the detection task is to detect the human heart rhythm; or, the task data is the vibration signal of a device and the operation record text of the device, and the detection task is to detect the fault type of the device.

[0032] In a second aspect, an embodiment of the present invention further provides a learning method for a machine learning model, including:

[0033] Determine each machine learning submodel in the machine learning model;

[0034] Apply inverse reinforcement learning to each machine learning submodel to obtain the reward subfunction of each machine learning submodel;

[0035] Obtain a global reward function based on the reward subfunction of each machine learning submodel;

[0036] Use the global reward function as the loss function of the machine learning model, and apply reinforcement learning to the machine learning model based on the loss function to obtain the machine learning model after reinforcement learning. The machine learning model after reinforcement learning is used to implement the above-mentioned human-computer interaction method based on inverse reinforcement learning

[0037] In a third aspect, an embodiment of the present invention further provides a human-computer interaction system based on inverse reinforcement learning, where the system includes the following components:

[0038] A display module, configured to obtain task data and a detection task input by a user terminal, and display a machine learning model required for completing the detection task, the performance of the machine learning model, and model parameters, where the machine learning model is a model after being trained by inverse reinforcement learning;

[0039] A detection module, configured to obtain a preselected model input by the user terminal and model parameters of the preselected model, apply the preselected model after loading the model parameters to the task data, and obtain a detection result of the detection task, where the preselected model is a model selected by the user terminal from the machine learning models according to the performance of the machine learning model;

[0040] An interaction module, configured to obtain feedback information of the user terminal on the detection result, and obtain a final detection result according to the feedback information.

[0041] In a fourth aspect, an embodiment of the present invention further provides a terminal device, where the terminal device includes a memory, a processor, and a human-computer interaction program based on inverse reinforcement learning stored in the memory and executable on the processor. When the processor executes the human-computer interaction program based on inverse reinforcement learning, the steps of the above-mentioned human-computer interaction method based on inverse reinforcement learning are implemented.

[0042] In a fifth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where a human-computer interaction program based on inverse reinforcement learning is stored on the computer-readable storage medium. When the human-computer interaction program based on inverse reinforcement learning is executed by a processor, the steps of the above-mentioned human-computer interaction method based on inverse reinforcement learning are implemented.

[0043] Beneficial effects: In the human-machine system of the present invention, several machine learning models after inverse reinforcement learning training are pre-saved. When the user inputs task data and a detection task into the human-machine interaction system, the system analyzes each machine learning model that matches the task data and the detection task according to the task data and the detection task, and displays each matched machine learning model, as well as the performance and model parameters of each machine learning model. The user selects a preselected model that can meet the actual needs from each machine learning model according to the actual needs and the performance of the machine learning model, and further selects the model parameters of the preselected model that can meet the actual needs. The system processes the task data based on the preselected model loaded with the model parameters to initially obtain the detection result of the detection task. The user makes corresponding feedback information based on the detection result. The system can adjust the model and its parameters according to the feedback information to obtain the final detection result. From the above analysis, it can be seen that in the present application, both the selection of the model and the selection of the model parameters fully consider the actual needs of the user. Not only can the final detection result meet the actual needs of the user, but also the user interacts with the system multiple times during the process of outputting the detection result to improve the final detection result, thereby enhancing the user experience. Description of the Drawings

[0044] Figure 1 It is the overall flowchart of the present invention;

[0045] Figure 2 It is the structural diagram of the human-machine interaction system based on inverse reinforcement learning provided by the present invention;

[0046] Figure 3 It is the internal structure principle block diagram of the terminal device provided by the embodiment of the present invention. Detailed Embodiments

[0047] The following combines the embodiments and the drawings of the specification to clearly and completely describe the technical solutions in the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0048] It has been found through research that implanting a machine learning model obtained through reinforcement learning into a human-machine interaction system. When the user inputs task data and a detection task for the task data into the human-machine interaction system, the human-machine interaction system uses the internal machine learning model to process the task data to complete the detection task and obtain the detection result of the detection task. Different machine learning models have different performances, and there are certain differences in the results obtained by machine learning models with different performances when processing task data. For example, the accuracy and speed of different machine learning models in processing task data are also different. However, the existing human-machine interaction system only allows users to passively use the internal machine learning model of the system to process task data.

[0049] To solve the above technical problems, the present invention provides a human-computer interaction method, system, device and medium based on inverse reinforcement learning, which solves the problem that the existing human-computer interaction system is difficult to meet the actual needs of users. Specifically in implementation, the human-computer interaction system first obtains the task data and detection task input by the user terminal. The human-computer interaction system displays the machine learning model required to complete the detection task, the performance of the machine learning model and the model parameters. The machine learning model is the model after being trained by inverse reinforcement learning; the human-computer interaction system obtains the preselected model input by the user terminal and the model parameters of the preselected model, applies the preselected model after loading the model parameters to the task data to obtain the detection result of the detection task. The preselected model is the model selected by the user terminal from the machine learning models according to the performance of the machine learning model; the human-computer interaction system obtains the feedback information of the user terminal for the detection result, and obtains the final detection result based on the feedback information.

[0050] For example, taking the detection of the failure type of mechanical equipment using a human-computer interaction system as an example, the tester inputs the detection task and multimodal task data to the human-computer interaction system through a display screen (user terminal). The multimodal task data includes the vibration signal of the machine equipment and the text of the equipment operation record. The vibration signal is a signal that changes with time for the mechanical equipment, that is, the vibration signal is a time series signal. The detection task is to detect the failure type of the machine equipment. The human-computer interaction system has pre-stored various trained machine learning models in advance. The human-computer interaction system selects several machine learning models that can complete the detection task, their performance and model parameters from various machine learning models according to the detection task. The GUI interface of the display screen displays the selected several machine learning models, their performance and model parameters. The tester selects the preselected model that can meet his own needs according to the performance of the machine learning model (the performance includes the accuracy of the model and the time required to run the model). If the tester has a high requirement for the accuracy of the detection result, then a machine learning model with high accuracy will be selected; if the tester has a low requirement for the detection accuracy, but needs to quickly obtain the detection process, then a machine learning model with a short running time will be selected.

[0051] After the tester selects a preselected model on the GUI interface, the human-computer interaction system loads the model parameters corresponding to the preselected model (the parameters corresponding to the parameter model after training is completed) into the preselected model. Then, the preselected model applies the Fourier transform (the Fourier transform is the FFT transform) to the vibration signal to obtain the spectrogram of the vibration signal. The preselected model also identifies the illegal operation records in the device operation record text. The preselected model outputs the fault type of the machine device (the fault type is the detection result) based on the abnormal frequency on the spectrogram and the illegal operation records. The GUI interface displays the detection result to the user, and the user inputs the feedback information to the human-computer interaction system through the GUI interface. When the feedback information is negative feedback, that is, when the detection result has a large deviation, the human-computer interaction system replaces the model parameters. That is, the model parameters used to obtain the above detection result are not suitable for detecting the fault type of the machine device. Therefore, it is necessary to try to load the replaced model parameters into the preselected model to re-detect the fault type to complete the detection task. The model parameters before replacement and the model parameters after replacement are different model parameters generated by training the machine learning model with different training samples. When the feedback information is positive feedback, the detection result is saved and displayed to the user.

[0052] The human-computer interaction method based on inverse reinforcement learning in this embodiment can be applied to a terminal device. The terminal device can be a terminal product with data processing functions, such as a human-computer interaction system, etc. In this embodiment, as Figure 1 shown, the human-computer interaction method based on inverse reinforcement learning specifically includes the following steps:

[0053] S100, obtain the task data and detection task input by the user terminal, and display the machine learning model required to complete the detection task, the performance of the machine learning model, and the model parameters. The machine learning model is a model trained by using inverse reinforcement learning;

[0054] The reward function of inverse reinforcement learning is not a pre-defined reward function, but continuously adjusts the reward function according to the feedback information of the user on the model decision during the model learning process, that is, reversely adjusts or selects the reward function through the feedback information of the user on the model decision. Applying inverse reinforcement learning to the learning of the machine learning model in the human-computer interaction system can make the machine learning model better meet the user's needs.

[0055] S200, obtain the preselected model input by the user terminal and the model parameters of the preselected model, apply the preselected model after loading the model parameters to the task data, and obtain the detection result of the detection task. The preselected model is a model selected by the user terminal from the machine learning models according to the performance of the machine learning model;

[0056] S300. Obtain the feedback information of the user terminal for the detection result, and based on the feedback information, obtain the final detection result.

[0057] In the first embodiment, the machine learning model in steps S100 to S300 includes several machine learning sub-models. The several machine learning sub-models work together to complete the detection task. The several machine learning sub-models adopt the joint learning method, that is, each machine learning sub-model is trained by the joint training method. The specific steps are as follows: Apply inverse reinforcement learning to each machine learning sub-model to obtain the reward sub-function of each machine learning sub-model; Based on the reward sub-function of each machine learning sub-model, obtain the global reward function; Based on the global reward function, apply reinforcement learning to the machine learning model.

[0058] Use the sample detection task to train the machine learning model. The sample detection task includes sample detection sub-tasks. Obtain the example detection results of the sample detection sub-tasks given by the expert examples, and then apply inverse reinforcement learning to the sample detection sub-tasks and the example detection results to infer the reward sub-function required from the sample detection sub-tasks to the example detection results. The reward sub-functions of each machine learning sub-model are jointly calculated. The joint calculation includes weighted calculation to obtain the global reward function, and then apply reinforcement learning to the global reward function to complete the training of the machine learning model.

[0059] When the sample detection task is to detect whether the heart rhythm is abnormal, the detection sub-tasks include electrocardiogram (ECG) signal denoising, ECG signal feature extraction, and abnormality detection. The purpose of ECG signal denoising is to enhance the signal clarity. The reward function R of ECG signal denoising 1 :

[0060]

[0061] The purpose of ECG signal feature extraction is to accurately locate the R-wave peak. The reward function R of ECG signal feature extraction 2 :

[0062]

[0063] Precision is the accuracy of the model used for extracting ECG signal features, and recall is the recall rate of the model for extracting ECG signal features.

[0064] The purpose of abnormality detection is to identify whether the heart rhythm is abnormal. The reward function R of abnormality detection 3 :

[0065]

[0066] The number of correct samples is the number of samples correctly classified by the model used for abnormality detection.

[0067] Global reward function R:

[0068] R = w 1 R 1 + w 2 R 2 + w 3 R 3 ;

[0069] Wherein, w 1 , w 2 , w 3 are all weights, and w 1 , w 2 , w 3 can be dynamically adjusted according to the detection task.

[0070] In this embodiment, the reward function R can also be quantified by the final detection result:

[0071] R = α·TPR - β·FPR;

[0072] Both α and β are weights that can be adjusted according to the requirements of the detection task. TPR is the true positive rate, and FPR is the false positive rate. Among them, a true positive means that the actual situation is arrhythmia and the detection result is also arrhythmia; a false positive means that the actual situation is arrhythmia but the detection result is normal heart rhythm.

[0073] For example, the sample detection data is the sample vibration signal and sample operation record text of a machine device, and the sample detection task is to detect the device failure type. The sample detection task is split into extracting signal features, identifying illegal operations, and generating detection results. The first machine learning sub-model is used to extract signal features, the second machine learning sub-model is used to identify illegal operations, and the third machine learning sub-model is used to generate detection results.

[0074] Obtain the example signal features of the sample vibration signal given by the expert example, obtain the example illegal operations of the sample operation record text identified by the expert example, and obtain the example detection results given by the expert example.

[0075] Apply the first machine learning sub-model to the sample vibration signal to obtain the observed signal features, and apply inverse reinforcement learning to the example signal features and the observed signal features to obtain the first reward function of the first machine learning sub-model.

[0076] Apply the second machine learning sub-model to the sample operation record text to obtain the observed illegal operations, and apply inverse reinforcement learning to the observed illegal operations and the example illegal operations to obtain the second reward function of the second machine learning sub-model.

[0077] Apply the third machine learning sub - model to the characteristics of the observation signal and the illegal observation operations to obtain the observation detection result. Apply inverse reinforcement learning to the observation detection result and the example detection result to obtain the third reward function of the third machine learning sub - model.

[0078] Perform weighted calculation on the first reward function, the second reward function, and the third reward function to obtain the global reward function. Based on the global reward function, apply reinforcement learning to the machine learning model to complete the training of the machine learning model, that is, use the global reward function obtained by inverse reinforcement learning as the loss function of the machine learning model in reinforcement learning, and perform reinforcement learning on the machine learning model.

[0079] The detection task includes several subtasks. The reward function can quantify the completion degree of each subtask. For example, in electrocardiogram signal analysis, the reward function can represent the accuracy of R - wave peak detection. The definition of accuracy is the difference between the number of detected R - waves and the number of true R - waves. The reward function is also used to characterize the completion quality of the subtask. For example, in signal denoising, the reward function can be the correlation coefficient between the signal and the reference signal or the signal - to - noise ratio (SNR), which reflects the denoising effect.

[0080] Embodiment 2, based on Embodiment 1, the display in step S100 of the machine learning model required to complete the detection task and the performance and model parameters of the machine learning model includes: splitting the detection task into subtasks, determining the machine learning sub - models required to complete the subtasks, and the machine learning sub - models are models after being trained using inverse reinforcement learning; using the machine learning sub - models as the machine learning model, and displaying the performance and the machine learning sub - models corresponding to the performance and the model parameters of the machine learning sub - models in the order of the quality of the performance.

[0081] This embodiment displays the machine learning sub - models in the order of the quality of the performance, which is convenient for users to select the machine learning sub - models with corresponding performance according to their needs.

[0082] Embodiment 3, based on Embodiment 2 or Embodiment 1, the step S200 in this embodiment includes the following specific steps S201 to S205:

[0083] S201, determine various modal data in the task data, and obtain the processing duration required for processing each type of modal data.

[0084] S202, according to the processing duration, divide the modal data into a first data group and a second data group, and the processing duration of the first data group is greater than that of the second data group.

[0085] S203. Apply the preselected model after loading the model parameters to each piece of the modality data in the first data group, and process various pieces of the modality data in the first data group in parallel to obtain a first processing result of the first data group.

[0086] S204. Within a set duration before the end of the parallel processing, apply the preselected model after loading the model parameters to each piece of the modality data in the second data group, and process various pieces of the modality data in the second data group serially to obtain a second processing result of the second data group.

[0087] S205. Obtain a detection result of the detection task according to the first processing result and the second processing result.

[0088] The processing times of different modality data are different. If all the modality data are processed simultaneously, that is, processed in parallel, it will cause a huge pressure on the memory of the human-machine system; if all the modality data are processed serially, the detection duration will be increased, thus reducing the detection speed. In this embodiment, the modality data with a relatively large required processing duration is processed in parallel, and the modality data with a relatively small required processing duration is processed serially before the end of the parallel processing, so as to both relieve the memory pressure of the human-machine system and ensure the detection efficiency.

[0089] Embodiment 4. Based on Embodiment 3 or Embodiment 2 or Embodiment 1, step S300 in this embodiment includes the following specific steps: when the feedback information is negative feedback, replace the model parameters, load the replaced model parameters into the preselected model, and apply the preselected model to the task data again to adjust the detection result so as to obtain a final detection result.

[0090] When the user gives feedback information, it means that the user is not satisfied with the detection result given by the human-machine interaction system. Therefore, a new set of model parameters is replaced, and the preselected model after replacing the model parameters is used to detect the task data again to obtain a final detection result.

[0091] Replacing a new set of model parameters is an optimization strategy. A strategy is a set of actions taken by the human-machine interaction system in the current state to complete the detection task. A strategy is a function, and this function is used to map the relationship between the state and the action:

[0092] π(s) = a

[0093] Where \(s\) is the current state and \(a\) is the action. When the detection task is to detect whether the heart rhythm is abnormal, the current state includes the input data state, the user feedback state, the model state, the dynamic threshold adjustment, and the feedback iteration. Among them, the input data state is the current processing stage of the electrocardiogram (ECG) signal, including the original signal, the noise-reduced signal, and the feature-extracted signal. The user feedback state is the feedback result of the user on the previous detection result. The model state is the currently selected machine learning model and its loaded parameters. The detection progress state is the progress of the task, including whether noise reduction or feature extraction has been completed.

[0094] The actions include data preprocessing actions, feature extraction actions, model update actions, dynamic threshold adjustment actions, and feedback iteration actions. Among them, the data preprocessing action is to perform noise reduction, high-pass filtering, or low-pass filtering on the ECG signal. The feature extraction action is to select a feature extraction algorithm to extract the R-wave peak. The model update action is to change the structure of the machine learning model or adjust the model parameters according to the user feedback. The dynamic threshold adjustment action is to optimize the filter threshold used in the feature extraction process. The feedback iteration action is to retrain the model or optimize the detection strategy based on the feedback input by the user.

[0095] The strategies include task selection strategies, feature extraction strategies, optimization strategies, and detection result display strategies. Among them, the task selection strategy is to decide which subtask to process first. The feature extraction strategy is to select the optimal feature extraction algorithm and parameters. For example, choosing Fourier transform or wavelet transform in ECG signal analysis is a strategy selection. The optimization strategy is to dynamically adjust the threshold of the filter. The threshold is a boundary when the filter processes the signal, used to distinguish the effective part and the invalid part of the signal. The threshold determines which frequency components or intensities in the signal are retained and which are weakened. A larger threshold will ignore the signal components with lower amplitudes (this signal component is noise), thus improving the signal quality. However, if the threshold is too large, the effective information of the original signal will be weakened. A smaller threshold may result in more signal details being retained, but it may also retain noise. A smaller threshold may not be able to effectively suppress noise.

[0096] The threshold of the filter controls the trade-off between the signal and the noise, and its adjustment needs to be balanced between noise suppression and signal fidelity. Among them, a larger threshold is suitable for tasks that are sensitive to noise and have a larger signal amplitude, but attention should be paid to avoiding losing the weak signal features. A smaller threshold is suitable for scenarios that require high signal detail retention, but other methods (such as subsequent feature extraction strategies) need to be combined to process the noise.

[0097] For example, in ECG signal analysis, a larger threshold is used to exclude noise when detecting the QRS complex. A smaller threshold is used to retain weak signals when analyzing the characteristics of the P wave and the T wave.

[0098] The filter is used to remove the noise from the electrocardiogram (ECG) signal. The detection result display strategy determines how to present the detection results to the user in the most intuitive way.

[0099] When the detection task is to detect whether the heart rhythm is abnormal, the strategy update mechanism includes dynamically updating the strategy based on user input, the output of the human-computer interaction system, and environmental changes; or deriving the reward function through inverse reinforcement learning, and iteratively optimizing the strategy through the reward function to maximize the reward value; or when the input task data changes significantly, trying a new machine learning model; or dynamically adjusting the memory resource weights assigned to each detection task according to the priority and complexity of the detection task that the user needs to solve.

[0100] Embodiment 5, based on Embodiment 3 or Embodiment 2 or Embodiment 1, the step S300 in this embodiment includes the following specific steps: when the feedback information is negative feedback, obtain real-time training samples through the retrieval-augmented generation algorithm; use the real-time training samples to retrain the preselected model; apply the retrained preselected model to the task data to obtain the final detection result.

[0101] When the feedback information is negative feedback, it means that the user is not satisfied with the detection result. Then it is necessary to use the retrieval-augmented generation algorithm RAG to obtain the latest real-time training samples, retrain the preselected model according to the user's latest detection task with the real-time training samples, and use the retrained preselected model to process the task data again to obtain the final detection result.

[0102] The human-computer interaction system based on any one of Embodiments 1 to 5 can be applied to ECG signal detection to determine whether the user's heart rhythm is abnormal, including preprocessing the ECG signal, feature extraction, and abnormality detection.

[0103] Among them, preprocessing the ECG signal includes: using high-pass filtering on the ECG signal (the cut-off frequency of the high-pass filter is 0.5 Hz, and the high-pass filter includes FIR filter and IIR filter, the FIR filter is the finite impulse response filter, and the IIR filter is the recursive filter) to remove the DC component of the ECG signal, so as to achieve the technical effect of removing DC offset; then using low-pass filtering on the ECG signal (the cut-off frequency of the low-pass filter is 150 Hz) to remove the high-frequency noise generated by muscle activity, so as to achieve the technical effect of removing baseline drift.

[0104] Feature extraction includes: detecting the R-wave peaks of the preprocessed electrocardiogram signal using the Pan-Tompkins algorithm (the positive upward peaks that appear on the electrocardiogram of the R-wave peaks), calculating the time interval between two adjacent R-wave peaks, calculating the average heart rate and HRV index of the user based on the electrocardiogram signal, generating a spectrogram of the electrocardiogram signal, extracting the low-frequency power and high-frequency power from the spectrogram, and calculating the ratio of the low-frequency power to the high-frequency power.

[0105] Determining whether the user's heart rhythm is abnormal includes: using the time interval between two adjacent R-wave peaks as the input of the classifier to obtain the detection result of whether the heart rhythm is abnormal, where the classifier is trained based on the reward function obtained by inverse reinforcement learning, and the calculation method of the reward function is to calculate the difference Diff between the standard classification result and the actual classification result of the classifier, then calculate the absolute value of the difference Diff|, and finally take the reverse score of the absolute value as the reward function R.

[0106] The noise reduction and feature extraction steps of the electrocardiogram signal are independent of each other and can be optimized separately. Use inverse reinforcement learning to optimize the classification strategy, dynamically adjust parameters, and improve the accuracy of heart rhythm abnormality detection.

[0107] The human-computer interaction system based on any one of Embodiment 1 to Embodiment 5 can be applied to seismic signal analysis, including: the user uploads the time series data collected by the seismic monitor (the time series data includes the task data of the seismic waveform) to the human-computer interaction system, and sets the detection task to identify the epicenter location and magnitude. The human-computer interaction system performs a Fourier transform on the seismic waveform to extract the characteristics of the frequency change of the seismic signal over time, identifies the P-band and S-band of the seismic waveform, and needs to adjust the window length of the Fourier transform to observe the characteristics of the frequency change of different frequency resolution rates over time in real time. Finally, the human-computer interaction system combines the change characteristics and geographical information data to obtain the epicenter location prediction result and the magnitude prediction result.

[0108] In summary, the present invention improves the processing efficiency of complex tasks and the user operation experience. The multi-modal feature fusion improves the accuracy of data processing. The reward inference and optimization significantly enhance the learning ability of complex tasks. The dynamic knowledge retrieval ensures the real-time nature of the collected training samples.

[0109] This embodiment also provides a human-computer interaction system based on inverse reinforcement learning, as Figure 2 shown, the system includes the following components:

[0110] A display module 01 is configured to obtain task data and a detection task input by a user terminal, and display a machine learning model required to complete the detection task, the performance of the machine learning model, and model parameters. The machine learning model is a model after being trained by inverse reinforcement learning;

[0111] A detection module 02 is configured to obtain a preselected model input by the user terminal and the model parameters of the preselected model, apply the preselected model after loading the model parameters to the task data, and obtain a detection result of the detection task. The preselected model is a model selected by the user terminal from the machine learning models according to the performance of the machine learning model;

[0112] An interaction module 03 is configured to obtain feedback information of the user terminal on the detection result, and obtain a final detection result according to the feedback information.

[0113] Based on the above embodiments, the present invention further provides a terminal device, and its principle block diagram can be as Figure 3 shown. The terminal device includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the terminal device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a human-computer interaction method based on inverse reinforcement learning. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen.

[0114] Those skilled in the art can understand that Figure 3 the principle block diagram shown in

[0115] merely shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0116] Obtain task data and a detection task input by a user terminal, and display a machine learning model required to complete the detection task, the performance of the machine learning model, and model parameters. The machine learning model is a model after being trained by inverse reinforcement learning;

[0117] Obtain the preselected model input by the user terminal and the model parameters of the preselected model, apply the preselected model after loading the model parameters to the task data to obtain the detection result of the detection task, where the preselected model is a model selected by the user terminal from the machine learning models according to the performance of the machine learning model;

[0118] Obtain the feedback information of the user terminal for the detection result, and obtain the final detection result based on the feedback information.

[0119] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A human-computer interaction method based on inverse reinforcement learning, characterized in that: include: Obtaining task data and detection tasks input by a user terminal, and displaying the machine learning model required to complete the detection task and the performance and model parameters of the machine learning model, wherein the machine learning model is a model trained by inverse reinforcement learning; Obtaining a preselected model input by the user terminal and model parameters of the preselected model, applying the preselected model after loading the model parameters to the task data, and obtaining a detection result of the detection task, wherein the preselected model is a model selected by the user terminal from the machine learning model according to the performance of the machine learning model; Acquire feedback information of the user terminal regarding the detection result, and obtain a final detection result based on the feedback information.

2. The human-computer interaction method based on inverse reinforcement learning according to claim 1, characterized in that: The display of the machine learning model required to complete the detection task and the performance and model parameters of the machine learning model include: Splitting the detection task into subtasks, and determining a machine learning submodel required to complete the subtasks, wherein the machine learning submodel is a model trained by inverse reinforcement learning; The machine learning sub-model is used as the machine learning model, and the performance and the machine learning sub-model and model parameters of the machine learning sub-model corresponding to the performance are displayed in order of the performance.

3. The human-computer interaction method based on inverse reinforcement learning according to claim 1, characterized in that: Applying the preselected model after loading the model parameters to the task data to obtain a detection result of the detection task includes: Determine various modal data in the task data, and obtain the processing time required to process each of the modal data; According to the processing time, the modal data is divided into a first data group and a second data group, wherein the processing time of the first data group is longer than the processing time of the second data group; Applying the preselected model after loading the model parameters to each of the modal data in the first data group, processing the various modal data in the first data group in parallel, and obtaining a first processing result of the first data group; Within a set time period before the parallel processing ends, applying the preselected model after loading the model parameters to each of the modal data in the second data group, serially processing the various modal data in the second data group, and obtaining a second processing result of the second data group; A detection result of the detection task is obtained according to the first processing result and the second processing result.

4. The human-computer interaction method based on inverse reinforcement learning according to claim 1, characterized in that: According to the feedback information, the final detection result is obtained, including: When the feedback information is negative feedback, the model parameters are replaced, and the replaced model parameters are loaded into the preselected model, and the preselected model is reapplied to the task data to adjust the detection result to obtain the final detection result.

5. The human-computer interaction method based on inverse reinforcement learning according to claim 1, characterized in that: According to the feedback information, the final detection result is obtained, including: When the feedback information is negative feedback, obtaining real-time training samples through a retrieval enhancement generation algorithm; Retraining the preselected model using the real-time training samples; The retrained preselected model is applied to the task data to obtain a final detection result.

6. The human-computer interaction method based on inverse reinforcement learning according to any one of claims 1 to 5, characterized in that: The learning method of the machine learning model includes: Determining each machine learning sub-model in the machine learning model; Applying inverse reinforcement learning to each of the machine learning sub-models to obtain a reward sub-function of each of the machine learning sub-models; Obtaining a global reward function according to the reward sub-function of each of the machine learning sub-models; Reinforcement learning is applied to the machine learning model according to the global reward function.

7. The human-computer interaction method based on inverse reinforcement learning according to any one of claims 1 to 5, characterized in that: The task data is a human electrocardiogram signal, and the detection task is to detect a human heart rhythm; or, the task data is a vibration signal of a device and an operation record text of the device, and the detection task is to detect a fault type of the device.

8. A learning method for a machine learning model, characterized in that: include: Determining each machine learning sub-model in the machine learning model; Applying inverse reinforcement learning to each of the machine learning sub-models to obtain a reward sub-function of each of the machine learning sub-models; Obtaining a global reward function according to the reward sub-function of each of the machine learning sub-models; The global reward function is used as the loss function of the machine learning model, and reinforcement learning is applied to the machine learning model based on the loss function to obtain the machine learning model after reinforcement learning. The machine learning model after reinforcement learning is used to implement the human-computer interaction method based on inverse reinforcement learning as described in claim 1.

9. A human-computer interaction system based on inverse reinforcement learning, characterized in that: The system comprises the following components: A display module is used to obtain task data and detection tasks input by a user terminal, and to display the machine learning model required to complete the detection task and the performance and model parameters of the machine learning model, wherein the machine learning model is a model trained by inverse reinforcement learning; a detection module, configured to obtain a preselected model input by the user terminal and model parameters of the preselected model, apply the preselected model after loading the model parameters to the task data, and obtain a detection result of the detection task, wherein the preselected model is a model selected by the user terminal from the machine learning model according to the performance of the machine learning model; The interaction module is used to obtain feedback information from the user terminal regarding the detection result, and obtain the final detection result based on the feedback information.

10. A terminal device, characterized in that: The terminal device includes a memory, a processor, and a human-computer interaction program based on inverse reinforcement learning stored in the memory and executable on the processor. When the processor executes the human-computer interaction program based on inverse reinforcement learning, the steps of the human-computer interaction method based on inverse reinforcement learning as described in any one of claims 1 to 7 are implemented.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a human-computer interaction program based on inverse reinforcement learning. When the human-computer interaction program based on inverse reinforcement learning is executed by the processor, the steps of the human-computer interaction method based on inverse reinforcement learning as described in any one of claims 1 to 7 are implemented.