A multi-modal intent recognition method based on a text core and related equipment
By constructing unidirectional connection weights and eliminating spurious associations, the problem of ignoring causal dependencies in multimodal intent recognition is solved, achieving more accurate intent recognition results.
Patent Information
- Application Number
- CN202511569550.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-10-30
AI Technical Summary
In existing multimodal user intent recognition methods, traditional fusion methods ignore causal dependencies, leading to false associations and interference from non-dominant modalities. Furthermore, fixed weights cannot be adaptively adjusted, resulting in inaccurate intent recognition.
By constructing non-textual modal structure equations and textual modal structure equations, unidirectional connection weights are obtained, spurious associations are eliminated, pure features are obtained, and feature fusion is performed to finally perform intent recognition.
It improves the accuracy of intent recognition by preserving the parts of data features that are relevant to other modalities and enhancing the information relevance of pure features, thus achieving more accurate intent recognition.
Smart Images

Figure CN121030247B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intention recognition, in particular to a multi-modal intention recognition method based on a text core and related equipment. BACKGROUND
[0002] In multi-modal user intention recognition, traditional fusion methods such as early fusion directly concatenate multi-modal features, but ignore causal dependencies. Attention mechanisms lack causal constraints and are easily disturbed by non-dominant modalities. Existing causal discovery methods assume symmetric modal dependencies and fail to highlight the role of the text core. These methods generally have key problems such as false associations, non-dominant modality interference, and static fusion strategies. Specifically, statistical correlation does not equal causality, leading to learning of spurious causal relationships, redundant information from non-text modalities such as vision and audio interfering, and fixed connection weights failing to adaptively adjust modality interaction strength.
[0003] In multi-modal user intention recognition tasks, the core challenge is how to effectively fuse multi-modal information such as text, speech, and vision. Traditional methods have significant shortcomings: early fusion directly concatenates features but ignores causal dependencies between modalities, leading to the possibility of learning false associations; attention mechanisms lack causal constraints and are easily disturbed by non-dominant modalities; existing methods typically use static fusion strategies, which have fixed weights that cannot adapt to dynamic scenario requirements, and generally fail to highlight the core driving role of the text modality in intention understanding. These limitations make it difficult for existing models to distinguish between statistical correlation and true causality, leading to inaccurate intention recognition. SUMMARY
[0004] The present application provides a multi-modal intention recognition method based on a text core and related equipment, which can solve the problem of inaccurate intention recognition.
[0005] In a first aspect, the present application provides a multi-modal intention recognition method based on a text core, which includes:
[0006] Obtaining multi-modal interaction data when a target object interacts with an environment, and performing feature extraction on each interaction data to obtain data features of each modality; the multiple modalities include a text modality and multiple non-text modalities;
[0007] Obtaining a one-way connection weight between each two modalities, and constructing a non-text modality structural equation and a text modality structural equation based on all connection weights; the connection weight is used to describe the correlation between two modalities, the non-text modality structural equation is used to convert the interaction data of the text modality into the data features of the non-text modality, and the text modality structural equation is used to convert the interaction data of all non-text modalities into the data features of the text modality;
[0008] According to the non-textual modal structural equation and the textual modal structural equation, false elimination is performed on data characteristics of each modality to obtain pure characteristics of each modality; the pure characteristics of each modality are parts of data characteristics of the modality that are related to data characteristics of other modalities;
[0009] Feature fusion is performed on all pure characteristics to obtain fusion characteristics of the target object;
[0010] According to the fusion characteristics, intent recognition is performed on the target object to obtain an intent recognition result of the target object.
[0011] Optionally, the multiple non-textual modalities are a visual modality and an audio modality;
[0012] The non-textual modal structural equation is:
[0013] ;
[0014] ;
[0015] wherein, denotes data characteristics of the visual modality converted from data characteristics of the textual modality and data characteristics of the audio modality, denotes data characteristics of the audio modality converted from data characteristics of the textual modality and data characteristics of the visual modality, denotes data characteristics of the textual modality, denotes data characteristics of the audio modality, denotes data characteristics of the visual modality, denotes a strong causal function from the textual modality to the visual modality, denotes a weak causal function from the audio modality to the visual modality, denotes a strong causal function from the textual modality to the audio modality, denotes a weak causal function from the visual modality to the audio modality, denotes inherent noise of the visual modality, denotes inherent noise of the audio modality, denotes a variance of a Gaussian distribution, and all denote dynamic gating weights:
[0016] ;
[0017] ;
[0018] wherein, denotes a Sigmoid function, denotes a connection weight from the visual modality to the audio modality, denotes a connection weight from the audio modality to the visual modality, denotes an indicator function conditioned on denotes the inter-modal local similarity.
[0019] Optionally, the text modality structural equation is:
[0020] ;
[0021] wherein, denotes the data feature of the text modality converted from the data feature of the visual modality and the data feature of the audio modality, denotes the intrinsic noise of the text modality, denotes a strong causal function from the audio modality to the text modality, denotes a strong causal function from the visual modality to the text modality.
[0022] Optionally, according to the non-text modality structural equation and the text modality structural equation, the data feature of each modality is de-noised to obtain the pure feature of each modality, including:
[0023] According to the non-text modality structural equation and the text modality structural equation, the false correlation component of the data feature of each modality is calculated.
[0024] For each modality, the false correlation component is removed from the data feature of the modality to obtain the pure feature of the modality.
[0025] Optionally, according to the non-text modality structural equation and the text modality structural equation, the false correlation component of the data feature of each modality is calculated, including:
[0026] Through the formula:
[0027] ;
[0028] ;
[0029] ;
[0030] The false correlation component of the data feature of the visual modality , the false correlation component of the data feature of the audio modality , and the false correlation component of the data feature of the text modality ;
[0031] wherein, denotes the correlation representation between the data features, denotes the average expected influence.
[0032] Optionally, the false correlation component is removed from the data feature of the modality to obtain the pure feature of the modality, including:
[0033] Through the formula:
[0034] ;
[0035] Calculate modes Pure characteristics ;
[0036] in, Indicates the intensity of elimination. Representing modes Data characteristics Representing modes spurious correlation components, ,when At that time, mode For text modality, when At that time, mode For visual modality, when At that time, mode It is an audio modal.
[0037] Optionally, feature fusion is performed on all pure features to obtain the fused features of the target object, including:
[0038] The attention weights for the visual modality and the audio modality are calculated based on the data features of all modalities.
[0039] The data features of all modalities are fused based on all attention weights, and the pure features of all modalities are combined to obtain the fused features of the target object.
[0040] Optionally, attention weights for the visual modality and attention weights for the audio modality are calculated based on data features from all modalities, including:
[0041] Through the formula:
[0042] ;
[0043] Calculate modes Attention weight ;
[0044] in, Representing modes Data characteristics Representing feature dimension, Represents text modality to modality The connection weight, when At that time, mode For visual modality, when At that time, mode For audio modality, This represents the Softmax activation function. represents a Sigmoid function;
[0045] According to all attention weights, data features of all modalities are fused, and pure features of all modalities are combined to obtain fusion features of the target object, including:
[0046] Through the formula:
[0047] ;
[0048] Calculate the fusion features ;
[0049] wherein, represents a batch normalization layer, represents a multi-layer perception.
[0050] In a second aspect, the present application provides a multi-modal intent recognition device based on a text core, comprising:
[0051] A feature extraction module is configured to obtain a plurality of modal interaction data of a target object when interacting with an environment, and extract features of each interaction data to obtain data features of each modality; the plurality of modalities include a text modality and a plurality of non-text modalities;
[0052] A construction module is configured to obtain a one-way connection weight between each two modalities, and construct a non-text modality structural equation and a text modality structural equation based on all connection weights; the connection weight is used to describe the correlation between two modalities, the non-text modality structural equation is used to convert the interaction data of the text modality into the data features of the non-text modality, and the text modality structural equation is used to convert the interaction data of all non-text modalities into the data features of the text modality;
[0053] A false elimination module is configured to eliminate false according to the non-text modality structural equation and the text modality structural equation, and obtain pure features of each modality; the pure features of each modality are the parts related to the data features of other modalities in the data features corresponding to the modality;
[0054] A feature fusion module is configured to fuse all pure features to obtain fusion features of the target object;
[0055] An intent recognition module is configured to recognize the intent of the target object according to the fusion features to obtain an intent recognition result of the target object.
[0056] In a third aspect, the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned computer program to realize the above-mentioned multi-modal intent recognition method based on a text core.
[0057] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the multi-modal intention recognition method based on a text core.
[0058] The above scheme of the present application has the following advantages:
[0059] In some embodiments of the present application, by obtaining a plurality of modal interaction data of a target object when interacting with an environment, and performing feature extraction on each interaction data, data features of each modality are obtained. Then, a one-way connection weight between each two modalities is obtained, and a non-text modality structural equation and a text modality structural equation are constructed based on all connection weights. Then, according to the non-text modality structural equation and the text modality structural equation, false elimination is performed on the data features of each modality to obtain pure features of each modality. Then, feature fusion is performed on all pure features to obtain fusion features of the target object. Finally, intention recognition is performed on the target object according to the fusion features to obtain an intention recognition result of the target object. Wherein, the false elimination on the data features can retain the part of the data features related to the data features of other modalities, improve the information correlation of the pure features, and the fusion features obtained by fusing the pure features with high information correlation have high information accuracy and include multi-modal data information. Based on the accurate fusion features, the intention recognition can effectively improve the accuracy of the intention recognition.
[0060] Other advantages of the present application will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0062] Figure 1 The flow chart of the multi-modal intention recognition method based on a text core provided by an embodiment of the present application;
[0063] Figure 2 The structural schematic diagram of the multi-modal intention recognition device based on a text core provided by an embodiment of the present application;
[0064] Figure 3 The structural schematic diagram of the terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0065] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and
[0066] It will be understood that the term "includes," "including," "has," "having," "comprises," "comprising" and the like when used in the present specification and throughout the claims, specify the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0067] It will be understood that the term "and / or," when used in the present specification and throughout the claims, refers to one or more of the associated listed items, and all possible combinations of one or more of the associated listed items.
[0068] As used in the present specification and throughout the claims, the term "if" can be interpreted as meaning "when" or "once" or "in response to a determination" or "in response to a detection" depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted as meaning "once it is determined" or "in response to the determination" or "once [the described condition or event] is detected" or "in response to the detection [of the described condition or event]" depending on the context.
[0069] In addition, the terms "first," "second," "third," etc. are used herein only to describe different instances of an element, and do not imply relative importance of the elements.
[0070] The terms "one embodiment," "an embodiment," "some embodiments," "other embodiments," "another embodiment," "one implementation," "an implementation," "some implementations," "other implementations," "another implementation," etc. have the same meaning and can be used interchangeably. Each of the expressions "in one embodiment" or "in some embodiments" and the like appearing in various places of the specification are not necessarily referring to the same embodiment or embodiments, but are intended to convey that the feature so described can be present in one or more embodiments.
[0071] In view of the problem that existing intent recognition is not accurate, the embodiment of the present application provides a multi-modal intent recognition method based on a text core. The multi-modal intent recognition method obtains interaction data of multiple modalities when a target object interacts with an environment, and extracts features of each interaction data to obtain data features of each modality. Then, one-way connection weights between each two modalities are obtained, and a non-text modality structural equation and a text modality structural equation are constructed based on all the connection weights. Then, the data features of each modality are eliminated to obtain pure features of each modality according to the non-text modality structural equation and the text modality structural equation. Then, all the pure features are fused to obtain fusion features of the target object. Finally, intent recognition is performed on the target object according to the fusion features to obtain an intent recognition result of the target object. Wherein, the data features are eliminated to retain the part of the data features related to the data features of other modalities, thereby improving the information correlation of the pure features. The fusion features obtained by fusing the pure features with high information correlation have high information accuracy and include multi-modal data information. The intent recognition based on the accurate fusion features can effectively improve the accuracy of intent recognition.
[0072] Next, the multi-modal intent recognition method based on a text core provided by the present application is exemplarily described.
[0073] As shown in Figure 1 , the multi-modal intent recognition method based on a text core provided by the present application includes the following steps:
[0074] Step 11, obtaining interaction data of multiple modalities when a target object interacts with an environment, and extracting features of each interaction data to obtain data features of each modality.
[0075] The multiple modalities include a text modality and multiple non-text modalities. The multiple non-text modalities are a visual modality and an audio modality. The target object is an object that needs to be subjected to intent recognition, such as a mentally ill patient that needs to be subjected to behavior monitoring.
[0076] Exemplarily, for an object that needs to be subjected to behavior monitoring, the environment in which the object is located is a ward or a monitoring environment in which the object is located, the interaction data of the text modality is chat records of the object, the interaction data of the audio modality is telephone and chat audio of the object, and the interaction data of the visual modality is monitoring video of the object.
[0077] In some embodiments of the present application, the interaction data of the target object can be obtained by a monitor, a microphone, a sensor or the like, and the interaction data can be subjected to feature extraction by using a convolutional neural network or the like.
[0078] Step 12, obtain the connection weight of one-way between each two modalities, and construct the non-text modality structural equation and the text modality structural equation based on all connection weights.
[0079] The non-text modality structural equation is used to convert the interaction data of the text modality into the data features of the non-text modality, and the text modality structural equation is used to convert all the interaction data of the non-text modality into the data features of the text modality. The above connection weight is used to describe the correlation degree between two modalities, which is obtained by initialization and parameter training.
[0080] Specifically, the non-text modality structural equation is:
[0081]
[0082]
[0083] wherein, represents the data features of the visual modality converted from the data features of the text modality and the data features of the audio modality, represents the data features of the audio modality converted from the data features of the text modality and the data features of the visual modality, represents the data features of the text modality, represents the data features of the audio modality, represents the data features of the visual modality, represents a strong causal function (usually a function based on a deep multilayer perception machine) from the text modality to the visual modality, represents a weak causal function (usually a function based on a deep multilayer perception machine) from the audio modality to the visual modality, represents a strong causal function from the text modality to the audio modality, represents a weak causal function from the visual modality to the audio modality, represents the inherent noise of the visual modality, represents the inherent noise of the audio modality, represents the variance of the Gaussian distribution, and all represent dynamic gating weights:
[0084]
[0085]
[0086] wherein, represents a Sigmoid function, represents the connection weight from the visual modality to the audio modality, represents the connection weight from the audio modality to the visual modality, represents is an indicator function of condition, output 1 if the condition is met, otherwise 0, denotes the inter-modal local similarity:
[0087] ;
[0088] wherein, , .
[0089] The text modal structure equation is:
[0090] ;
[0091] wherein, denotes the data feature of the text modal converted from the data feature of the visual modal and the data feature of the audio modal, denotes the inherent noise of the text modal, denotes the strong causal function from the audio modal to the text modal, denotes the strong causal function from the visual modal to the text modal.
[0092] It should be noted that in real scenarios, although text is usually the dominant modal of multi-modal interaction, there may be weak associations between non-textual modalities (visual-audio), which may come from physical laws (such as lip synchronization with speech) or cross-modal common features (such as facial expressions and tone of voice in emotional expression). Therefore, allow controlled weak direct connections between non-textual modalities, but their strength should be significantly lower than the text-related connections. The connection weights between non-textual modalities and are allowed to exist, but must satisfy:
[0093] ;
[0094] wherein, controls the maximum relative strength of the weak connection, , . In order to improve the precision of the connection weight, calculate the causal graph regularization loss:
[0095] ;
[0096] If the value of the causal graph regularization loss is less than the loss preset value, it is considered that the precision of the connection weight reaches the expectation, if the value of the causal graph regularization loss is greater than or equal to the loss preset value, adjust the connection weight between the non-textual modalities, and recalculate the causal graph regularization loss until the value of the causal graph regularization loss is less than the loss preset value.
[0097] Step 13, according to the non-textual modal structure equation and the text modal structure equation, false elimination is performed on the data features of each modal to obtain the pure features of each modal.
[0098] The pure feature of each modality is the part of the data features corresponding to that modality that is related to the data features of other modalities. For example, for the visual modality, the interaction data includes images of the target object making gestures and video data of walking. The video data of walking corresponds to the sound made by the target object walking. Therefore, the video data is considered to be related to the audio modality. That is, in the data features of the visual modality, the part corresponding to the video data is related to the data features of the audio data. Therefore, in the data features of the visual modality, the part corresponding to the video data is retained, and the part corresponding to the gesture image is eliminated.
[0099] In some embodiments of this application, the step of performing spurious elimination on the data features of each modality based on the non-text modal structure equation and the text modal structure equation to obtain the pure features of each modality includes:
[0100] The first step is to calculate the spurious association components of the data features for each modality based on the non-text modal structure equation and the text modal structure equation.
[0101] The spurious correlation component of each modality is the part of the data feature corresponding to that modality that is not related to the data feature of other modalities.
[0102] Specifically, through the formula:
[0103] ;
[0104] ;
[0105] ;
[0106] False association components of data features in computational visual modalities spurious correlation components of audio modality data features spurious correlation components of data features in text modality .
[0107] in, This represents the correlation between data features. Indicates the average expected impact:
[0108] ;
[0109] in, , .
[0110] It should be noted that the expression for the average expected impact during the training phase is as follows: , Indicates the first Data features of each training sample , This indicates the number of training samples. In the actual calculations described in this step, the average expected impact value is equal to the value of the data features of the target object.
[0111] The second step is to eliminate spurious correlation components from the data features of each modality to obtain the pure features of the modality.
[0112] Specifically, through the formula:
[0113] ;
[0114] Calculate modes Pure characteristics .
[0115] in, Indicates the intensity of elimination. Representing modes Data characteristics Representing modes spurious correlation components, ,when At that time, mode For text modality, when At that time, mode For visual modality, when At that time, mode It is an audio modal.
[0116] Step 14: Perform feature fusion on all pure features to obtain the fused features of the target object.
[0117] In some embodiments of this application, the step of performing feature fusion on all pure features to obtain the fused features of the target object includes:
[0118] The first step is to calculate the attention weights for the visual modality and the audio modality based on the data features of all modalities.
[0119] Through the formula:
[0120] ;
[0121] Calculate modes Attention weight ;
[0122] in, Representing modes Data characteristics Representing feature dimension, Represents text modality to modality The connection weight, when At that time, mode For the visual modality, when the modality is an audio modality, represents a Softmax activation function, represents a Sigmoid function.
[0123] Secondly, the data features of all modalities are fused according to all attention weights, and the pure features of all modalities are combined to obtain the fusion features of the target object.
[0124] The fusion features are calculated by the formula:
[0125] ;
[0126] The fusion features .
[0127] wherein, represents a batch normalization layer, represents a multi-layer perception.
[0128] Step 15, according to the fusion features, the intention of the target object is recognized, and the intention recognition result of the target object is obtained.
[0129] The intention recognition result can be a label, which is used to describe the behavior, action, etc. of the target object, such as eating, drinking, going out, etc.
[0130] For example, the fusion features can be calculated by using a fully connected layer to obtain the intention recognition result of the target object.
[0131] It should be noted that before the above method is performed, in order to improve the accuracy of the intention recognition of the target object, the parameters in the above formulas can be trained by using sample objects first. The sample object is an object with an actual intention recognition result (which can be obtained by acquiring and analyzing historical behavior intention data, and the time of the interaction data of the sample object is the time of the behavior intention data, such as making a behavior of eating at 8 o'clock, then the acquisition time of the interaction data of the sample object should be before 8 o'clock, such as 7:50), and the intention recognition result of each sample object is obtained through the above process. According to the intention recognition result, a loss function is constructed, and it is judged whether the value of the loss function is less than a preset value. If yes, it is considered that the training is completed, otherwise, the parameters in the above weak causal functions, strong causal functions, formulas for calculating pure features, formulas for calculating attention weights, and formulas for calculating fusion features are adjusted, and the intention recognition result of each sample object is recalculated until the value of the loss function is less than the preset value. The above loss function is:
[0132] ;
[0133] ;
[0134] ;
[0135] in, This represents the value of the loss function. Indicates the intention to identify loss. Indicates causal loss. Indicates the weight of causal discovery loss. Indicates the number of sample objects. , and They represent the first Data features of text modality, video modality, and audio modality corresponding to each sample. , and They represent the first Data features of text modality, video modality, and audio modality corresponding to each sample, obtained through structural equation modeling. Indicates the first Intent recognition results for each sample object Indicates the first The actual intent recognition result of each sample object.
[0136] It is worth mentioning that false feature removal can preserve the parts of data features that are related to data features of other modalities, thereby improving the information relevance of pure features. The information accuracy of the fused features obtained by fusing pure features with high information relevance is high, and it includes multimodal data information. Based on accurate fused features, intent recognition can effectively improve the accuracy of intent recognition.
[0137] Furthermore, the method in this application explicitly models bidirectional causal dependencies by constructing a causal graph with text as the hub and constraining the strength of weak connections between non-textual modalities. It selectively activates modal connections by combining a dynamic gating mechanism that adapts to local similarity, and uses counterfactual intervention to eliminate false associations. Finally, it modulates cross-modal attention weights based on causal strength. While maintaining the complementarity of multimodal information, it effectively solves the problems of causal confusion and noise sensitivity of traditional methods, and provides a more reliable solution for text-dominated multimodal tasks.
[0138] The following is an exemplary description of the text-based multimodal intent recognition device provided in this application.
[0139] like Figure 2 As shown, this application provides a text-based multimodal intent recognition device 200, which includes:
[0140] The feature extraction module 201 is configured to acquire interaction data of multiple modalities when the target object interacts with the environment, and perform feature extraction on each interaction data to obtain data features of each modality; the multiple modalities include a text modality and multiple non-text modalities;
[0141] The construction module 202 is configured to acquire one-way connection weights between each two modalities, and construct a non-text modality structural equation and a text modality structural equation based on all the connection weights; the connection weights are used to describe the correlation between two modalities, the non-text modality structural equation is used to convert the interaction data of the text modality into the data features of the non-text modality, and the text modality structural equation is used to convert all the interaction data of the non-text modalities into the data features of the text modality;
[0142] The false elimination module 203 is configured to perform false elimination on the data features of each modality according to the non-text modality structural equation and the text modality structural equation, to obtain pure features of each modality; the pure features of each modality are the parts of the data features corresponding to the modality that are related to the data features of other modalities;
[0143] The feature fusion module 204 is configured to perform feature fusion on all the pure features to obtain fusion features of the target object.
[0144] The intent recognition module 205 is configured to perform intent recognition on the target object according to the fusion features, to obtain an intent recognition result of the target object.
[0145] It should be noted that the information interaction, execution process and the like between the above apparatus / units are based on the same concept as the method embodiments of the present application, and the specific functions and technical effects brought by them can be referred to the method embodiments part, which will not be repeated here.
[0146] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of the functional units and modules are only for convenient distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can be referred to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0147] AsFigure 3 As shown, the embodiment of the present application provides a terminal device, the terminal device D10 of the embodiment comprises at least one processor D100 (only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps in any of the method embodiments described above when executing the computer program D102. Figure 3
[0148] Specifically, when the processor D100 executes the computer program D102, the processor D100 obtains the interaction data of multiple modalities when the target object interacts with the environment, and extracts the features of each interaction data to obtain the data features of each modality, then obtains the one-way connection weight between each two modalities, and constructs the non-text modality structural equation and the text modality structural equation based on all the connection weights, and then eliminates the false of the data features of each modality according to the non-text modality structural equation and the text modality structural equation to obtain the pure features of each modality, then fuses all the pure features to obtain the fusion features of the target object, and finally performs intent recognition on the target object according to the fusion features to obtain the intent recognition result of the target object. Wherein, the false elimination of the data features can retain the part of the data features related to the data features of other modalities, improve the information correlation of the pure features, the fusion features obtained by fusing the pure features with high information correlation have high information accuracy, and include multi-modal data information, and the intent recognition based on the accurate fusion features can effectively improve the accuracy of intent recognition.
[0149] The processor D100 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0150] The storage D101 can be an internal storage unit of the terminal device D10 in some embodiments, such as a hard disk or a memory of the terminal device D10. The storage D101 can also be an external storage device of the terminal device D10 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device D10. Further, the storage D101 can include both the internal storage unit and the external storage device of the terminal device D10. The storage D101 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of the computer program, etc. The storage D101 can also be used to temporarily store data that has been output or will be output.
[0151] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above-mentioned various method embodiments.
[0152] The embodiments of the present application provide a computer program product. When the computer program product is run on a terminal device, the terminal device is caused to implement the steps in the above-mentioned various method embodiments.
[0153] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the embodiments of the present application can implement all or part of the processes in the above-mentioned method embodiments by a computer program to instruct related hardware to complete, and the computer program can be stored in a computer readable storage medium. The computer program is executed by a processor to implement the steps in the above-mentioned various method embodiments. The computer program includes computer program codes, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium at least includes any entity or device capable of carrying the computer program codes of the text-based multi-modal intent recognition method and device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal and a software distribution medium. For example, a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc.
[0154] In the above embodiments, the description of each embodiment is focused on, and the part not described or recorded in a certain embodiment can be referred to the relevant description of other embodiments.
[0155] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0156] The above is the preferred embodiment of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the principles described in the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.
Claims
1. A multimodal intent recognition method based on text core, characterized in that, include: The system acquires interaction data from multiple modalities when the target object interacts with its environment, and extracts features from each interaction data to obtain the data features of each modality. Multiple modalities include a text modality and multiple non-text modalities; Obtain the unidirectional connection weights between every two modalities, and construct non-text modal structure equations and text modal structure equations based on all connection weights; the connection weights are used to describe the degree of correlation between two modalities, the non-text modal structure equations are used to convert the interaction data of the text modality into non-text modal data features, and the text modal structure equations are used to convert the interaction data of all non-text modalities into text modal data features. Based on the non-text modal structure equation and the text modal structure equation, false positives are eliminated from the data features of each modality to obtain the pure features of each modality; the pure features of each modality are the parts of the data features corresponding to that modality that are related to the data features of other modalities. All pure features are fused to obtain the fused features of the target object; The intent of the target object is identified based on the fusion features to obtain the intent identification result of the target object; The step of performing spurious feature removal on the data features of each modality based on the non-text modal structure equation and the text modal structure equation to obtain the pure features of each modality includes: The spurious correlation component of the data features for each modality is calculated based on the non-text modal structure equation and the text modal structure equation; the spurious correlation component of a modality is the part of the data features corresponding to that modality that is not related to the data features of other modalities. For each of the modalities, spurious correlation components are eliminated from the data features of the modality to obtain the pure features of the modality; the pure features of the modality are the parts of the data features corresponding to the modality that are related to the data features of other modalities.
2. The multimodal intent recognition method according to claim 1, characterized in that, The multiple non-textual modalities are visual modalities and audio modalities; The non-text modal structure equation is: ; ; in, This represents the data features of the visual modality obtained by converting data features from text modality and audio modality. The data features representing the audio modality are obtained by converting the data features of the text modality and the data features of the visual modality. Data features representing text modalities Data features representing audio modalities Data features representing visual modalities This represents a strong causal function from the text modality to the visual modality. This represents a weak causal function representing the transition from the audio modality to the visual modality. This represents a strong causal function representing the transition from text modality to audio modality. This represents a weak causal function from the visual modality to the audio modality. This represents the inherent noise of the visual modality. This represents the inherent noise of the audio mode. This represents the variance of the Gaussian distribution. and All represent dynamic gating weights: ; ; in, This represents the Sigmoid function. This represents the connection weights from the visual modality to the audio modality. This represents the connection weights from the audio modality to the visual modality. Indicated by For conditional indicator functions, This represents the local similarity between modalities.
3. The multimodal intent recognition method according to claim 2, characterized in that, The text modal structure equation is: ; in, Data features of the text modality are obtained by converting data features representing the visual modality and data features representing the audio modality. Indicates the inherent noise of the text modality. This represents a strong causal function representing the transition from the audio modality to the text modality. This represents a strong causal function from the visual modality to the text modality.
4. The multimodal intent recognition method according to claim 1, characterized in that, The step of calculating the spurious association components of the data features for each modality based on the non-text modal structure equation and the text modal structure equation includes: Through the formula: ; ; ; False association components of data features in computational visual modalities spurious correlation components of audio modality data features spurious correlation components of data features in text modality ; in, This represents the correlation between data features. This represents the average expected impact.
5. The multimodal intent recognition method according to claim 4, characterized in that, The step of eliminating spurious correlation components from the data features of the modality to obtain the pure features of the modality includes: Through the formula: ; Calculate the mode Pure characteristics ; in, Indicates the intensity of elimination. Representing modes Data characteristics Representing modes spurious correlation components, ,when At that time, mode For text modality, when At that time, mode For visual modality, when At that time, mode It is an audio modal.
6. The multimodal intent recognition method according to claim 5, characterized in that, The process of fusing all pure features to obtain the fused features of the target object includes: The attention weights for the visual modality and the audio modality are calculated based on the data features of all modalities. The data features of all modalities are fused based on all attention weights, and the pure features of all modalities are combined to obtain the fused features of the target object.
7. The multimodal intent recognition method according to claim 6, characterized in that, The calculation of attention weights for the visual modality and the audio modality based on data features from all modalities includes: Through the formula: ; Calculate the mode Attention weight ; in, Representing modes Data characteristics Representing feature dimension, Represents text modality to modality The connection weight, when At that time, mode For visual modality, when At that time, mode For audio modality, This represents the Softmax activation function. Represents the Sigmoid function; The process of fusing data features from all modalities based on all attention weights and combining the pure features from all modalities to obtain the fused features of the target object includes: Through the formula: ; Computational fusion features ; in, Indicates the batch normalization layer. This represents a multilayer perceptron.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the text-based multimodal intent recognition method as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the text-based multimodal intent recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal intention recognition method and system for uncertain modal deficiency
CN118245846A
Humanoid robot adaptive control method and system, electronic equipment and storage medium
CN119820582A