Intention recognition method and system based on combination of electroencephalogram signals and images
By combining EEG signals and image data, feature extraction and recognition are performed, and adjustments are made with external stimulation data, the problem of low intention recognition accuracy in the prior art is solved, and higher intention recognition accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510167726.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the recognition accuracy of intention recognition is low, making it difficult to effectively understand the user's intention and emotional state.
The intention recognition method based on the combination of EEG signals and images is adopted. By collecting the user's EEG signal data and image information, the EEG features and image features are extracted, and the identification and fusion are finally adjusted with external stimulation data to improve the accuracy of intention recognition.
By integrating multiple data sources, users’ intentions can be fully captured from multiple levels, the accuracy of identification can be improved, and the identification effect can be further improved through adaptive adjustment mechanisms.
Smart Images

Figure CN119989061A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and more specifically to an intention recognition method and system based on the combination of electroencephalogram signals and images. Background Art
[0002] At present, intention recognition is to understand the user's potential intentions semantically in order to better answer the user's questions or provide relevant services, which is of great significance for realizing the intelligent and personalized human-computer interaction. User intention recognition has many application scenarios, especially in the field of simulated automatic dressing. User intention recognition can provide richer and more accurate information to help the system understand the user's intentions and emotional state.
[0003] However, the process of intent recognition in the prior art has the problem of low recognition accuracy.
[0004] Therefore, how to provide an intention recognition method that can solve the above problems is an issue that those skilled in the art urgently need to solve. Summary of the invention
[0005] In view of this, the present invention provides an intention recognition method and system based on the combination of EEG signals and images, which improves the accuracy of user intention recognition.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] An intention recognition method based on the combination of EEG signals and images comprises the following steps:
[0008] Collect the user's EEG signal data and corresponding image information;
[0009] Extracting features from the EEG signal data and the image information respectively, and obtaining corresponding EEG features and image features;
[0010] Identify the EEG features and the image features to obtain corresponding preliminary intention recognition results;
[0011] The external stimulus data of the user's environment within a preset time period is acquired in real time, and the preliminary intention recognition result is adjusted in combination with the external stimulus data to obtain a final intention recognition result.
[0012] Preferably, the specific processing process of obtaining the corresponding preliminary intention recognition result includes:
[0013] Respectively performing dimensionality reduction and normalization processing on the image features and the EEG features, and performing feature fusion on the image features and the EEG features after the above processing to obtain corresponding feature fusion results;
[0014] A recognition model is constructed, and the feature fusion result is input into the recognition model for processing to obtain a corresponding preliminary intent recognition result.
[0015] Preferably, the specific processing of obtaining the final intention recognition result includes:
[0016] Extracting the image information to obtain corresponding prior environment data and prior user action data;
[0017] Acquire in real time external stimulus data of the user's environment within a preset period of time, wherein the external stimulus data includes real-time image data;
[0018] Extracting the real-time image data to obtain corresponding subsequent environment data and subsequent user action data;
[0019] Performing similarity judgment on the previous environment data and the subsequent environment data to obtain a first judgment result, and performing similarity judgment on the previous user action data and the subsequent user action data to obtain a second judgment result;
[0020] The preliminary intention recognition result is adjusted according to the first judgment result and the second judgment result to obtain a final intention recognition result.
[0021] Preferably, the specific processing of adjusting the preliminary intention recognition result according to the first judgment result and the second judgment result includes:
[0022] Respectively comparing the first judgment result and the second judgment result with a preset similarity threshold to obtain a corresponding first comparison result and a corresponding second comparison result;
[0023] When the first comparison result meets the threshold requirement and the second comparison result does not meet the threshold requirement, extracting features from the subsequent user action data to obtain corresponding subsequent action features;
[0024] The subsequent action feature and the feature fusion result are again subjected to feature fusion to obtain a corresponding first fusion feature, and the first fusion feature is input into the recognition model for recognition to obtain a final intention recognition result;
[0025] When the first comparison result does not meet the threshold requirement and the second comparison result meets the threshold requirement, extracting features from the subsequent environment data to obtain corresponding subsequent environment features;
[0026] The post-environment feature and the feature fusion result are fused again to obtain a corresponding second fused feature, and the second fused feature is input into the recognition model for recognition to obtain a final intention recognition result.
[0027] Preferably, the specific processing of adjusting the preliminary intention recognition result according to the first judgment result and the second judgment result further includes:
[0028] When both the first judgment result and the second judgment result do not meet the threshold requirements, the subsequent action features, the subsequent environment features and the feature fusion results are respectively fused to obtain corresponding third fused features, and the third fused features are input into the recognition model for recognition to obtain the final intention recognition result.
[0029] Preferably, when both the first judgment result and the second judgment result meet the threshold requirement, the specific processing process further includes:
[0030] Acquire the ambient audio data of the user's environment within a preset time period in real time, and process the ambient audio data to obtain corresponding audio processing results;
[0031] The audio processing result and the feature fusion result are input into the recognition model for recognition to obtain the final intent recognition result.
[0032] The present invention also provides a recognition system for an intention recognition method based on a combination of EEG signals and images, comprising:
[0033] A data acquisition module, used to collect the user's EEG signal data and corresponding image information;
[0034] A feature extraction module, used to extract features from the EEG signal data and the image information respectively, and obtain corresponding EEG features and image features;
[0035] A recognition module, used to recognize the EEG features and the image features to obtain corresponding preliminary intention recognition results;
[0036] The adjustment module is used to obtain the external stimulus data of the user's environment within a preset time period in real time, and adjust the preliminary intention recognition result in combination with the external stimulus data to obtain the final intention recognition result.
[0037] Preferably, the identification module includes:
[0038] A feature fusion unit, used to perform dimensionality reduction and normalization processing on the image features and the EEG features respectively, and perform feature fusion on the image features and the EEG features after the above processing to obtain corresponding feature fusion results;
[0039] The model building and recognition unit is used to build a recognition model and input the feature fusion result into the recognition model for processing to obtain a corresponding preliminary intention recognition result.
[0040] Preferably, the adjustment module includes:
[0041] A first extraction unit, configured to extract the image information to obtain corresponding prior environment data and prior user action data;
[0042] A collection unit, used for acquiring in real time external stimulus data of the user's environment within a preset period of time, wherein the external stimulus data includes real-time image data;
[0043] A second extraction unit, used to extract the real-time image data to obtain corresponding subsequent environment data and subsequent user action data;
[0044] An adjustment unit, wherein the adjustment unit is connected to the model building and identification unit, and is used to perform a similarity judgment on the prior environmental data and the subsequent environmental data to obtain a first judgment result, and simultaneously perform a similarity judgment on the prior user action data and the subsequent user action data to obtain a second judgment result, and adjust the preliminary intention recognition result according to the first judgment result and the second judgment result to obtain a final intention recognition result.
[0045] Preferably, the adjustment unit comprises:
[0046] A comparing unit, used to compare the first judgment result and the second judgment result with a preset similarity threshold value respectively, to obtain a corresponding first comparison result and a corresponding second comparison result;
[0047] A first processing unit, the first processing unit is connected to the model building and recognition unit, and is used for, when the first comparison result meets the threshold requirement and the second comparison result does not meet the threshold requirement, performing feature extraction on the subsequent user action data to obtain a corresponding subsequent action feature, performing feature fusion on the subsequent action feature and the feature fusion result again to obtain a corresponding first fusion feature, and inputting the first fusion feature into the model building and recognition unit for recognition to obtain a final intention recognition result;
[0048] A second processing unit, the second processing unit is connected to the model building and recognition unit, and is used for, when the first comparison result does not meet the threshold requirement and the second comparison result meets the threshold requirement, extracting features from the subsequent environment data to obtain corresponding subsequent environment features, performing feature fusion on the subsequent environment features and the feature fusion result again to obtain corresponding second fused features, and inputting the second fused features into the model building and recognition unit for recognition to obtain a final intention recognition result;
[0049] a third processing unit, the third processing unit being connected to the first processing unit, the second processing unit and the model building and recognition unit, and being used for, when the first judgment result and the second judgment result do not meet the threshold requirement, respectively fusing the subsequent action feature, the subsequent environment feature and the feature fusion result to obtain a corresponding third fused feature, and inputting the third fused feature into the model building and recognition unit for recognition to obtain a final intention recognition result;
[0050] a fourth processing unit, the fourth processing unit being connected to the first processing unit, the second processing unit and the model building and recognition unit, and being used for acquiring in real time the ambient audio data of the user's environment within a preset time period when both the first judgment result and the second judgment result meet the threshold requirements, and processing the ambient audio data to obtain a corresponding audio processing result, and inputting the audio processing result and the feature fusion result into the model building and recognition unit for recognition to obtain a final intention recognition result;
[0051] It can be seen from the above technical solution that, compared with the prior art, the present invention discloses an intention recognition method and system based on the combination of EEG signals and images, which collects the user's EEG signal data and corresponding image information, and obtains corresponding EEG features and image features; identifies the EEG features and the image features to obtain corresponding preliminary intention recognition results; obtains in real time the external stimulus data of the user's environment within a preset time period, and adjusts the preliminary intention recognition results in combination with the external stimulus data to obtain the final intention recognition results; the method provided by the present invention integrates multiple data sources, and different types of data each contain different features, which can comprehensively capture the user's intentions from multiple levels, thereby improving the accuracy of recognition.
[0052] At the same time, the present invention can adjust the preliminary intention recognition results in combination with external stimulus data. This mechanism enables the system to perform adaptive adjustments according to different environmental conditions, further improving the accuracy of recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0054] Figure 1 The overall flow chart of the intention recognition method based on the combination of EEG signals and images provided by the present invention;
[0055] Figure 2 A structural principle block diagram of an intention recognition system based on the combination of EEG signals and images provided by the present invention;
[0056] Figure 3 This is a structural principle block diagram of the adjustment unit provided by the present invention. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] See also Figure 1 As shown, the embodiment of the present invention discloses a method for intention recognition based on the combination of EEG signals and images, comprising the following steps:
[0059] Collect the user's EEG signal data and corresponding image information;
[0060] Extract features from EEG signal data and image information respectively, and obtain corresponding EEG features and image features;
[0061] Identify EEG features and image features to obtain corresponding preliminary intention recognition results;
[0062] The external stimulus data of the user's environment within a preset time period is obtained in real time, and the preliminary intention recognition result is adjusted in combination with the external stimulus data to obtain the final intention recognition result. This embodiment can also make personalized recommendations based on the final intention recognition result.
[0063] Specifically, when the intention recognition method provided by the present embodiment is applied to the field of motion imagination, the recognized intention at this time is the user's action intention, and the present embodiment can also generate a corresponding recommended route based on the final intention recognition result; when the intention recognition method provided by the present embodiment is applied to the field of simulated automatic changing of clothes, the recognized intention is the user's changing of clothes intention, and the present embodiment can also generate a corresponding personalized clothing recommendation result based on the final intention recognition result to ensure that it meets the user's current situation and needs.
[0064] In a specific embodiment, the specific processing process of obtaining the corresponding preliminary intention recognition result includes:
[0065] Performing dimensionality reduction and normalization processing on the image features and EEG features respectively, and performing feature fusion on the processed image features and EEG features to obtain corresponding feature fusion results;
[0066] Construct a recognition model and input the feature fusion results into the recognition model for processing to obtain the corresponding preliminary intent recognition results.
[0067] Specifically, the recognition model may include a CNN network and an LSTM recurrent neural network connected in sequence, which may realize joint processing of feature fusion results and improve recognition accuracy.
[0068] In a specific embodiment, the specific process of obtaining the final intention recognition result includes:
[0069] Extracting image information to obtain corresponding prior environment data and prior user action data;
[0070] Acquire in real time external stimulus data of the user's environment within a preset period of time, wherein the external stimulus data includes real-time image data;
[0071] Extracting real-time image data to obtain corresponding subsequent environment data and subsequent user action data;
[0072] Performing similarity judgment on the previous environment data and the subsequent environment data to obtain a first judgment result, and performing similarity judgment on the previous user action data and the subsequent user action data to obtain a second judgment result;
[0073] The preliminary intention recognition result is adjusted according to the first judgment result and the second judgment result to obtain a final intention recognition result.
[0074] In a specific embodiment, the specific processing of adjusting the preliminary intention recognition result according to the first judgment result and the second judgment result includes:
[0075] Respectively comparing the first judgment result and the second judgment result with a preset similarity threshold to obtain a corresponding first comparison result and a corresponding second comparison result;
[0076] When the first comparison result meets the threshold requirement and the second comparison result does not meet the threshold requirement, feature extraction is performed on the subsequent user action data to obtain corresponding subsequent action features;
[0077] The subsequent action features and feature fusion results are fused again to obtain the corresponding first fusion features, and the first fusion features are input into the recognition model for recognition to obtain the final intention recognition result;
[0078] When the first comparison result does not meet the threshold requirement and the second comparison result meets the threshold requirement, feature extraction is performed on the post-environment data to obtain corresponding post-environment features;
[0079] The post-environment features and feature fusion results are fused again to obtain the corresponding second fusion features, and the second fusion features are input into the recognition model for recognition to obtain the final intention recognition result.
[0080] In a specific embodiment, the specific processing of adjusting the preliminary intention recognition result according to the first judgment result and the second judgment result further includes:
[0081] When both the first judgment result and the second judgment result do not meet the threshold requirements, feature fusion is performed on the subsequent action features, the subsequent environment features and the feature fusion results respectively to obtain the corresponding third fusion features, and the third fusion features are input into the recognition model for recognition to obtain the final intention recognition result.
[0082] In a specific embodiment, when both the first judgment result and the second judgment result meet the threshold requirement, the specific processing process further includes:
[0083] Acquire the ambient audio data of the user's environment within a preset period in real time, and process the ambient audio data to obtain corresponding audio processing results;
[0084] The audio processing results and feature fusion results are input into the recognition model for recognition to obtain the final intent recognition result.
[0085] See also Figure 2 As shown, an embodiment of the present invention further provides a recognition system using the intention recognition method based on the combination of EEG signals and images according to any one of the above embodiments, comprising:
[0086] A data acquisition module, used to collect the user's EEG signal data and corresponding image information;
[0087] A feature extraction module is used to extract features from EEG signal data and image information respectively, and obtain corresponding EEG features and image features;
[0088] The recognition module is used to recognize EEG features and image features to obtain the corresponding preliminary intention recognition results;
[0089] The adjustment module is used to obtain the external stimulus data of the user's environment within a preset time period in real time, and adjust the preliminary intention recognition result based on the external stimulus data to obtain the final intention recognition result.
[0090] In a specific embodiment, the identification module includes:
[0091] The feature fusion unit is used to perform dimensionality reduction and normalization processing on the image features and the EEG features respectively, and perform feature fusion on the processed image features and the EEG features to obtain corresponding feature fusion results;
[0092] The model building and recognition unit is used to build a recognition model and input the feature fusion results into the recognition model for processing to obtain the corresponding preliminary intention recognition results.
[0093] In a specific embodiment, the adjustment module includes:
[0094] A first extraction unit, used to extract image information to obtain corresponding prior environment data and prior user action data;
[0095] A collection unit, used to obtain in real time external stimulus data of the user's environment within a preset period of time, wherein the external stimulus data includes real-time image data;
[0096] A second extraction unit is used to extract the real-time image data to obtain corresponding subsequent environment data and subsequent user action data;
[0097] The adjustment unit is connected to the model building and identification unit, and is used to make a similarity judgment on the prior environmental data and the subsequent environmental data to obtain a first judgment result, and at the same time make a similarity judgment on the prior user action data and the subsequent user action data to obtain a second judgment result, and adjust the preliminary intention recognition result according to the first judgment result and the second judgment result to obtain a final intention recognition result.
[0098] See also Figure 3 As shown, in a specific embodiment, the adjustment unit includes:
[0099] A comparison unit, used to compare the first judgment result and the second judgment result with a preset similarity threshold value to obtain a corresponding first comparison result and a corresponding second comparison result;
[0100] A first processing unit, the first processing unit is connected to the model building and recognition unit, and is used for, when the first comparison result meets the threshold requirement and the second comparison result does not meet the threshold requirement, performing feature extraction on the subsequent user action data to obtain the corresponding subsequent action feature, performing feature fusion on the subsequent action feature and the feature fusion result again to obtain the corresponding first fusion feature, and inputting the first fusion feature into the model building and recognition unit for recognition to obtain a final intention recognition result;
[0101] The second processing unit is connected to the model building and recognition unit, and is used for, when the first comparison result does not meet the threshold requirement and the second comparison result meets the threshold requirement, performing feature extraction on the post-environment data to obtain corresponding post-environment features, performing feature fusion again on the post-environment features and feature fusion results to obtain corresponding second fusion features, and inputting the second fusion features into the model building and recognition unit for recognition to obtain a final intention recognition result;
[0102] The third processing unit is connected to the first processing unit, the second processing unit and the model building and recognition unit, and is used for, when the first judgment result and the second judgment result do not meet the threshold requirement, respectively fusing the subsequent action feature, the subsequent environment feature and the feature fusion result to obtain a corresponding third fusion feature, and inputting the third fusion feature into the model building and recognition unit for recognition to obtain a final intention recognition result;
[0103] The fourth processing unit is connected to the first processing unit, the second processing unit and the model building and recognition unit, and is used to obtain the ambient audio data of the user's environment within a preset time period in real time when the first judgment result and the second judgment result both meet the threshold requirements, and process the ambient audio data to obtain the corresponding audio processing result, and input the audio processing result and the feature fusion result into the model building and recognition unit for recognition to obtain the final intention recognition result.
[0104] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0105] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for intention recognition based on the combination of EEG signals and images, characterized in that: The following steps are involved: Collect the user's EEG signal data and corresponding image information; Extracting features from the EEG signal data and the image information respectively, and obtaining corresponding EEG features and image features; Identify the EEG features and the image features to obtain corresponding preliminary intention recognition results; The external stimulus data of the user's environment within a preset time period is acquired in real time, and the preliminary intention recognition result is adjusted in combination with the external stimulus data to obtain a final intention recognition result.
2. The method for intention recognition based on the combination of EEG signals and images according to claim 1, characterized in that: The specific process of obtaining the corresponding preliminary intention recognition results includes: Respectively performing dimensionality reduction and normalization processing on the image features and the EEG features, and performing feature fusion on the image features and the EEG features after the above processing to obtain corresponding feature fusion results; A recognition model is constructed, and the feature fusion result is input into the recognition model for processing to obtain a corresponding preliminary intent recognition result.
3. The method for intention recognition based on the combination of EEG signals and images according to claim 2, characterized in that: The specific processing to obtain the final intent recognition result includes: Extracting the image information to obtain corresponding prior environment data and prior user action data; Acquire in real time external stimulus data of the user's environment within a preset period of time, wherein the external stimulus data includes real-time image data; Extracting the real-time image data to obtain corresponding subsequent environment data and subsequent user action data; Performing similarity judgment on the previous environment data and the subsequent environment data to obtain a first judgment result, and performing similarity judgment on the previous user action data and the subsequent user action data to obtain a second judgment result; The preliminary intention recognition result is adjusted according to the first judgment result and the second judgment result to obtain a final intention recognition result.
4. The method for intention recognition based on the combination of EEG signals and images according to claim 3, characterized in that: The specific processing process of adjusting the preliminary intention recognition result according to the first judgment result and the second judgment result includes: Respectively comparing the first judgment result and the second judgment result with a preset similarity threshold to obtain a corresponding first comparison result and a corresponding second comparison result; When the first comparison result meets the threshold requirement and the second comparison result does not meet the threshold requirement, extracting features from the subsequent user action data to obtain corresponding subsequent action features; The subsequent action feature and the feature fusion result are again subjected to feature fusion to obtain a corresponding first fusion feature, and the first fusion feature is input into the recognition model for recognition to obtain a final intention recognition result; When the first comparison result does not meet the threshold requirement and the second comparison result meets the threshold requirement, extracting features from the subsequent environment data to obtain corresponding subsequent environment features; The post-environment feature and the feature fusion result are fused again to obtain a corresponding second fused feature, and the second fused feature is input into the recognition model for recognition to obtain a final intention recognition result.
5. The method for intention recognition based on the combination of EEG signals and images according to claim 4, characterized in that: The specific processing process of adjusting the preliminary intention recognition result according to the first judgment result and the second judgment result also includes: When both the first judgment result and the second judgment result do not meet the threshold requirements, the subsequent action features, the subsequent environment features and the feature fusion results are respectively fused to obtain corresponding third fused features, and the third fused features are input into the recognition model for recognition to obtain the final intention recognition result.
6. The method for intention recognition based on the combination of EEG signals and images according to claim 5, characterized in that: When both the first judgment result and the second judgment result meet the threshold requirement, the specific processing process further includes: Acquire the ambient audio data of the user's environment within a preset time period in real time, and process the ambient audio data to obtain corresponding audio processing results; The audio processing result and the feature fusion result are input into the recognition model for recognition to obtain the final intent recognition result.
7. A recognition system using the intention recognition method based on the combination of EEG signals and images as described in any one of claims 1 to 6, characterized in that: include: A data acquisition module, used to collect the user's EEG signal data and corresponding image information; A feature extraction module, used to extract features from the EEG signal data and the image information respectively, and obtain corresponding EEG features and image features; A recognition module, used to recognize the EEG features and the image features to obtain corresponding preliminary intention recognition results; The adjustment module is used to obtain the external stimulus data of the user's environment within a preset time period in real time, and adjust the preliminary intention recognition result in combination with the external stimulus data to obtain the final intention recognition result.
8. The identification system according to claim 7, characterized in that: The identification module comprises: A feature fusion unit, used to perform dimensionality reduction and normalization processing on the image features and the EEG features respectively, and perform feature fusion on the image features and the EEG features after the above processing to obtain corresponding feature fusion results; The model building and recognition unit is used to build a recognition model and input the feature fusion result into the recognition model for processing to obtain a corresponding preliminary intention recognition result.
9. The identification system according to claim 8, characterized in that: The adjustment module comprises: A first extraction unit, configured to extract the image information to obtain corresponding prior environment data and prior user action data; A collection unit, used for acquiring in real time external stimulus data of the user's environment within a preset period of time, wherein the external stimulus data includes real-time image data; A second extraction unit, used to extract the real-time image data to obtain corresponding subsequent environment data and subsequent user action data; An adjustment unit, wherein the adjustment unit is connected to the model building and identification unit, and is used to perform a similarity judgment on the prior environmental data and the subsequent environmental data to obtain a first judgment result, and simultaneously perform a similarity judgment on the prior user action data and the subsequent user action data to obtain a second judgment result, and adjust the preliminary intention recognition result according to the first judgment result and the second judgment result to obtain a final intention recognition result.
10. The identification system according to claim 9, characterized in that: The adjustment unit comprises: A comparing unit, used to compare the first judgment result and the second judgment result with a preset similarity threshold value respectively, to obtain a corresponding first comparison result and a corresponding second comparison result; A first processing unit, the first processing unit is connected to the model building and recognition unit, and is used for, when the first comparison result meets the threshold requirement and the second comparison result does not meet the threshold requirement, performing feature extraction on the subsequent user action data to obtain a corresponding subsequent action feature, performing feature fusion on the subsequent action feature and the feature fusion result again to obtain a corresponding first fusion feature, and inputting the first fusion feature into the model building and recognition unit for recognition to obtain a final intention recognition result; A second processing unit, the second processing unit is connected to the model building and recognition unit, and is used for, when the first comparison result does not meet the threshold requirement and the second comparison result meets the threshold requirement, extracting features from the subsequent environment data to obtain corresponding subsequent environment features, performing feature fusion on the subsequent environment features and the feature fusion result again to obtain corresponding second fused features, and inputting the second fused features into the model building and recognition unit for recognition to obtain a final intention recognition result; a third processing unit, the third processing unit being connected to the first processing unit, the second processing unit and the model building and recognition unit, and being used for, when the first judgment result and the second judgment result do not meet the threshold requirement, respectively fusing the subsequent action feature, the subsequent environment feature and the feature fusion result to obtain a corresponding third fused feature, and inputting the third fused feature into the model building and recognition unit for recognition to obtain a final intention recognition result; A fourth processing unit, which is connected to the first processing unit, the second processing unit and the model building and recognition unit, and is used to obtain in real time the ambient audio data of the user's environment within a preset time period when the first judgment result and the second judgment result both meet the threshold requirements, and process the ambient audio data to obtain a corresponding audio processing result, and input the audio processing result and the feature fusion result into the model building and recognition unit for recognition to obtain a final intention recognition result.