Text intention recognition method and device, storage medium and electronic device

CN115269774BActive Publication Date: 2026-05-15QINGDAO HAIER TECH +3
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO HAIER TECH
Filing Date
2022-06-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

[0005]本发明实施例提供了一种文本意图的识别方法和装置、存储介质和电子装置,以至少解决现有文本意图的识别方法的识别准确率较低的技术问题

Benefits of technology

[0010]在本发明实施例中,获取待识别的文本信息;对文本信息进行文本特征提取,得到文本特征向量集;对文本特征向量集中的每个文本特征向量分别进行槽填充处理,得到各自对应的语义类别向量,其中,语义类别向量集用于指示文本信息中每个字符各自对应的语义类别;对文本特征向量集和各个语义类别向量进行意图解析,以识别出文本信息所携带的操作意图,从而根据槽填充处理得到的文本信息中的各个字符的语义类别信息,进一步实现准确确定出用于文本信息中的意图类别,进而解决了现有文本意图的识别方法存在的识别准确率较低的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115269774B_ABST
    Figure CN115269774B_ABST
Patent Text Reader

Abstract

The application discloses a text intention recognition method and device, a storage medium and an electronic device, relates to the technical field of smart home, and the text intention recognition method comprises the following steps: obtaining text information to be recognized; text feature extraction is performed on the text information to obtain a text feature vector set; each text feature vector in the text feature vector set is subjected to slot filling processing respectively to obtain a corresponding semantic category vector, wherein the semantic category vector set is used for indicating the respective corresponding semantic category of each character in the text information; and the text feature vector set and the respective semantic category vector are subjected to intention analysis, so as to recognize the operation intention carried by the text information. The application solves the technical problem of low recognition accuracy of the existing text intention recognition method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart home technology, and more specifically, to a method and apparatus for recognizing textual intent, a storage medium, and an electronic device. Background Technology

[0002] In smart home scenarios, users typically control smart home devices via voice commands. However, because user voice commands are sometimes not standardized, devices cannot be controlled directly based on them. It is necessary to accurately recognize the intent behind the user's voice commands before generating precise control instructions. Therefore, accurate intent recognition of user voice commands is crucial for achieving precise device control in smart home scenarios.

[0003] Existing methods typically identify a user's true intent based on their current commands and actions. However, current methods for identifying true intent based on user voice commands are generally quite simplistic, merely using slot-filling networks and intent recognition networks to recognize voice commands separately. They fail to effectively utilize the textual information within the voice commands, resulting in poor intent recognition results. In other words, existing textual intent recognition methods suffer from the technical problem of low accuracy.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] The present invention provides a method and apparatus for identifying text intent, a storage medium and an electronic device, to at least solve the technical problem of low recognition accuracy of existing text intent identification methods.

[0006] According to one aspect of the present invention, a method for identifying text intent is provided, comprising: acquiring text information to be identified; extracting text features from the text information to obtain a text feature vector set; performing slot filling processing on each text feature vector in the text feature vector set to obtain a corresponding semantic category vector, wherein the semantic category vector set is used to indicate the semantic category corresponding to each character in the text information; and performing intent parsing on the text feature vector set and each semantic category vector to identify the operational intent carried by the text information.

[0007] According to another aspect of the present invention, a text intent recognition device is also provided, comprising: an acquisition unit for acquiring text information to be recognized; a feature extraction unit for extracting text features from the text information to obtain a text feature vector set; a semantic parsing unit for performing slot filling processing on each text feature vector in the text feature vector set to obtain a corresponding semantic category vector, wherein the semantic category vector set is used to indicate the semantic category corresponding to each character in the text information; and an intent parsing unit for performing intent parsing on the text feature vector set and each semantic category vector to identify the operational intent carried by the text information.

[0008] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, wherein the computer program is configured to execute the above-described text intent recognition method at runtime.

[0009] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the above-described text intent recognition method through the computer program.

[0010] In this embodiment of the invention, text information to be identified is obtained; text features are extracted from the text information to obtain a text feature vector set; slot filling is performed on each text feature vector in the text feature vector set to obtain its corresponding semantic category vector, wherein the semantic category vector set is used to indicate the semantic category corresponding to each character in the text information; intent parsing is performed on the text feature vector set and each semantic category vector to identify the operational intent carried by the text information, thereby accurately determining the intent category used in the text information based on the semantic category information of each character in the text information obtained by slot filling, thus solving the technical problem of low recognition accuracy in existing text intent recognition methods. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1This is a schematic diagram of the hardware environment for an optional text intent recognition method according to an embodiment of the present invention;

[0014] Figure 2 This is a flowchart of an optional text intent recognition method according to an embodiment of the present invention;

[0015] Figure 3 This is a schematic diagram of another optional method for recognizing text intent according to an embodiment of the present invention;

[0016] Figure 4 This is a schematic diagram of the structure of an optional text intent recognition device according to an embodiment of the present invention;

[0017] Figure 5 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed Implementation

[0018] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0020] According to one aspect of the embodiments of this application, a method for recognizing text intent is provided. This interaction method for IoT devices is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned interaction method for IoT devices can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0021] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0022] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0023] According to one aspect of the embodiments of the present invention, such as Figure 2 As shown, the above method for recognizing text intent includes the following steps:

[0024] S202, Obtain the text information to be recognized;

[0025] It is understandable that the text information to be identified can be the text information converted from the user's voice command to the smart device. By performing intent recognition on the text information, the accurate control intent of the user's voice command can be obtained.

[0026] For example, a user's voice command might be "wash the cherries clean." This command lacks a specific action subject, meaning the device to be controlled cannot be directly derived from it. In other words, the control intent in the voice command is vague, making it impossible to directly obtain control instructions based on it. Therefore, the voice command "wash the cherries clean" is converted into text information for further text intent recognition.

[0027] S204, extract text features from the text information to obtain a set of text feature vectors;

[0028] It is understood that, in this embodiment, the above-mentioned text information can be subjected to preliminary feature extraction to obtain multiple text feature vectors corresponding to each text character, and these multiple text feature vectors corresponding to each text character can be used as the above-mentioned text feature vector set. In one optional approach, the above-mentioned multiple text feature vectors may each correspond to a single text character; in another approach, the above-mentioned text feature vectors may correspond to multiple text feature vectors corresponding to the entire text information.

[0029] S206, each text feature vector in the text feature vector set is subjected to slot filling processing to obtain the corresponding semantic category vector. The semantic category vector set is used to indicate the semantic category corresponding to each character in the text information.

[0030] The results of the slot filling process described above are explained below. In this embodiment, a semantic category vector corresponding to each character can be obtained through slot filling. The semantic category vector is used to indicate the semantic category corresponding to each character. For example, for the five characters in the text information "cherry washed clean", the semantic category corresponding to the character "cherry" is "fruit", the semantic category corresponding to the character "peach" is "fruit", the semantic category corresponding to the character "wash" is "washing action", the semantic category corresponding to the character "dry" is "cleanliness level", and the semantic category corresponding to the character "clean" is "cleanliness level".

[0031] Optionally, a set of category labels can be pre-obtained, which includes multiple semantic category labels. Understandably, the aforementioned semantic category vector is used to indicate the degree of matching between the target character corresponding to the text information and the multiple category labels in the category label set, and the label with the highest matching degree is taken as the semantic category of the character. For example, if the pre-obtained set of category labels includes ("fruit", "washing action", "animal", "pattern"), then for the character "peach", its corresponding semantic category vector could be (90%, 10%, 30%, 20%). This semantic category vector indicates that the probability of the character "peach" being a fruit is 90%, the probability of being a washing action is 10%, the probability of being an animal is 30%, and the probability of being a pattern is 20%. Therefore, the category label "fruit" with the highest probability is determined as the slot-filling result for the character "peach".

[0032] In another alternative approach, the aforementioned semantic category vector can also be used to indicate the semantic category corresponding to each character in the semantic category combination that has the highest degree of matching degree of the overall category label determined based on the overall semantics of the text characters.

[0033] S208 performs intent parsing on the text feature vector set and each semantic category vector to identify the operational intent carried by the text information.

[0034] It should be noted that, in this embodiment, by performing intent parsing on the semantic category vector obtained from slot filling and the text feature vector indicating the overall semantics of the text information, the semantic category vector obtained from slot filling can be introduced into the intent recognition parsing process through an attention mechanism during the semantic parsing process, thereby making full use of the joint information and improving the accuracy of intent recognition.

[0035] Specifically, continuing with the example of "wash cherries clean," the above method is explained as follows: After obtaining the semantic category vectors corresponding to the characters "cherry," "peach," "wash," "dry," and "clean," the intent of "wash cherries clean" is parsed based on the semantic category indicated by each vector. For example, the result of the intent parsing could be "control the washing machine to wash the cherries." This would then trigger subsequent control operations.

[0036] Through the above-described embodiments of this application, text information to be identified is obtained; text features are extracted from the text information to obtain a text feature vector set; slot filling processing is performed on each text feature vector in the text feature vector set to obtain its corresponding semantic category vector, wherein the semantic category vector set is used to indicate the semantic category corresponding to each character in the text information; intent parsing is performed on the text feature vector set and each semantic category vector to identify the operational intent carried by the text information, thereby accurately determining the intent category used in the text information based on the semantic category information of each character in the text information obtained by slot filling processing, thus solving the technical problem of low recognition accuracy in existing text intent recognition methods.

[0037] As an optional implementation, the above-described intent parsing of the text feature vector set and each semantic category vector to identify the operational intent carried by the text information includes:

[0038] S1, obtain the reference feature vector located at the start flag position and each semantic category vector in the text feature vector set, wherein the text feature vector set includes the reference feature vector and the character feature vector corresponding to each character contained in the text information.

[0039] S2, input the reference feature vector and each semantic category vector into the first attention network to obtain the first output vector, wherein the first attention network is used to perform intent parsing on the text feature vector set by utilizing the attention weights matched with the semantic category vectors;

[0040] S3, the first output vector and the reference feature vector are jointly input into the first intent recognition network to obtain the recognized intent category vector, wherein the intent category vector is used to indicate the operation intent of the text information.

[0041] The following combination Figure 3 The above methods will be explained. For example... Figure 3As shown, the multiple feature vectors output by the pre-trained language model in the figure can be an example of the aforementioned text feature vector set. The output vector corresponding to the [CLS] identifier (represented by a box containing two gray dashed circles in the figure) can be an example of the aforementioned reference feature vector. The output vector obtained after slot filling can be an example of the aforementioned semantic category vectors, and an example of the slot-intent attention network and the aforementioned first attention network in the figure. As shown, the output vector corresponding to the [CLS] identifier and the output vector obtained after slot filling are used together as input to the slot-intent attention network to obtain the first output vector, i.e., the black solid circle in the intent recognition module in the figure. The first output vector is concatenated with the reference feature vector and then input into the first intent recognition feedforward network to obtain the final intent recognition category vector, which indicates the operational intent of the text information. It should be noted that the intent recognition category vector can be a 1*N vector, where N is the number of intent categories. The physical meaning of this vector is to represent the probability value of each intent category.

[0042] Through the above-described embodiments of this application, a reference feature vector located at the start flag position in the text feature vector set and each semantic category vector are obtained; the reference feature vector and each semantic category vector are input into a first attention network to obtain a first output vector; the first output vector and the reference feature vector are jointly input into a first intent recognition network to obtain the recognized intent category vector, thereby introducing slot filling result information as attention, and then performing intent recognition on the text, which improves the accuracy of intent recognition and solves the technical problem of inaccurate recognition results in existing text intent recognition methods.

[0043] As an optional implementation, the above-described input of the reference feature vector and each semantic category vector into the first attention network to obtain the first output vector includes:

[0044] S1, obtain the first weight matrix, the second weight matrix, and the third weight matrix that match the first attention network;

[0045] S2, obtain the dot product value calculated by performing dot product on the first feature vector and each reference semantic category vector respectively, wherein the first feature vector is determined by the reference feature vector and the first weight matrix, and the reference semantic category vector is determined by the semantic category vector and the second weight matrix;

[0046] S3, normalize each of the dot product values ​​to obtain multiple first reference values;

[0047] S4, perform a weighted summation of each first reference value and all second feature vectors to obtain the first output vector, wherein the second feature vector is determined by the reference feature vector and the third weight matrix.

[0048] Alternatively, the output of the above steps can be expressed by the following formula:

[0049] i i =β i v0

[0050] Where v0 is the output feature vector and weight matrix W of the pre-trained model. v (i.e., the product of the third weight matrix); β i = softmax(w·v0′), where v0′ is the pre-trained model's output feature vector and weight matrix W. k The product of (i.e., the first weight matrix), where w is the slot vector and W is the label weight matrix. q The product of (i.e., the second weight matrix) and the semantic category vector set composed of the semantic category vectors mentioned above, i i This is the final output of the slot-intent attention module, namely the first output vector mentioned above.

[0051] Through the above-described embodiments of this application, dot product values ​​are obtained by performing dot product calculations on the reference feature vector and each semantic category vector; each dot product value is normalized to obtain multiple first reference values; and each first reference value is weighted and summed with all semantic category vectors to obtain a first output vector.

[0052] As another optional implementation, the above-mentioned slot-filling process for each text feature vector in the text feature vector set to obtain their respective semantic category vectors includes:

[0053] S1, Input the reference feature vector into the second intent recognition network to obtain the intent reference vector;

[0054] S2 uses attention weights matched with the intent reference vector to perform slot filling parsing on the text feature vector set to obtain the semantic category vector.

[0055] Continue to combine Figure 3 The above methods will be explained, such as Figure 3 As shown, the feature vector corresponding to the [CLS] flag is processed by the second intent recognition feedforward network (i.e., the second intent recognition network mentioned above), and the resulting intent reference vector is used as the input of the intent-slot attention network for subsequent slot filling parsing operations.

[0056] It should be noted that the intent information output by the second intent recognition feedforward network here does not contain slot-filled output interaction information, to avoid the problem of interdependence between output results. Therefore, the intent information here is taken from a separate branch.

[0057] Through the above-described embodiments of this application, a reference feature vector is input into a second intent recognition network to obtain an intent reference vector; by using attention weights matched with the intent reference vector, slot filling parsing is performed on the text feature vector set to obtain a semantic category vector, thereby including the preliminary intent recognition result in the slot filling result and improving the accuracy of the slot filling result.

[0058] As an optional implementation, the above-described slot-filling parsing of the text feature vector set, utilizing attention weights matched with the intent reference vector, yields semantic category vectors including:

[0059] S1, input the intent reference vector and the text feature vector set into the second attention network to obtain multiple second output vectors. The second attention network is used to perform slot filling parsing on the text feature vector set using attention weights that match the intent reference vector. The second output vector corresponds one-to-one with each text feature vector in the text feature vector set.

[0060] S2, concatenate each second output vector with its corresponding text feature vector to obtain multiple concatenated vectors;

[0061] S3 uses multiple slot-filling networks connected to the second attention network to parse the multiple concatenated vectors to obtain their respective semantic category vectors.

[0062] Continue to combine Figure 3 The above methods will be explained. For example... Figure 3 As shown, the intent reference vector output by the second intent recognition feedforward network and the text feature vector output by the pre-trained language model are used as inputs to the intent-slot attention network (i.e., the second attention network mentioned above). In this step, the intent information output by the second intent recognition feedforward network is incorporated into each slot-filling feature vector in the form of attention. The fused features are then concatenated with the original features and input into the slot-filling single-layer feedforward network to obtain the final slot-filling output. It should be noted that when receiving the intent reference vector, a dimension transformation layer is required to achieve dimension alignment.

[0063] The method for obtaining the aforementioned second output vectors in this step can be expressed by the following formula:

[0064] S i =α i v

[0065] Where v is the output feature vector and weight matrix W of the pre-trained model. v Product; α i = softmax(w·v′), where v′ is the pre-trained model's output feature vector and weight matrix W. k The product of w and the label weight matrix W, where w is the intention vector and W is the label weight matrix. q The product of s and the aforementioned intention category vectors, i This is the final output of the attention module at the i-th position.

[0066] Furthermore, after obtaining each second output vector, each second output vector (i.e., the solid black circle) is concatenated with its corresponding text feature vector (the hollow dashed circle) to obtain multiple concatenated vectors, and each of these concatenated vectors is then input into the system. Figure 3 The slot-filling forward network shown in the figure is used to obtain the final parsed semantic category vector.

[0067] Through the above-described embodiments of this application, an intent reference vector and a set of text feature vectors are input into a second attention network to obtain multiple second output vectors; each second output vector is concatenated with its corresponding text feature vector to obtain multiple concatenated vectors; multiple slot filling networks connected to the second attention network are used to parse the multiple concatenated vectors to obtain their respective semantic category vectors, thereby including preliminary intent recognition results in the slot filling results and improving the accuracy of the slot filling results.

[0068] As an optional implementation, before obtaining the text information to be recognized, the method further includes:

[0069] S1, obtain multiple training sample text information and their corresponding intent category labels;

[0070] S2, using the training sample text information and the corresponding intent category label, jointly train the initial first attention network, the initial first intent recognition network, the initial second attention network, the initial second intent recognition network, and multiple initial slot filling networks until multiple loss values ​​continuously output by the first intent recognition network are all less than the target threshold. Among them, the initial second intent recognition network is connected to the initial second attention network, the initial second attention network is connected to multiple initial slot filling networks, the multiple initial slot filling networks are connected to the initial first attention network, and the initial first attention network is connected to the initial first intent recognition network.

[0071] The following discusses the aforementioned networks and Figure 3 The correspondence between the multiple networks in the above text is explained. The first attention network mentioned above can be... Figure 3 The slot-intent attention network in the middle, the first intent recognition network can be Figure 3The first intention recognition forward network in, the above second attention network can be Figure 3 The intention-slot attention network in, the above second intention recognition network can be Figure 3 The second intention recognition forward network in, the above multiple slot filling networks can be Figure 3 The slot filling forward network in.

[0072] It should be noted that the structures in the above slot-intention attention network and intention-slot attention network can be the same, but the finally trained network parameters can be different; the structures in the above first intention recognition forward network and second intention recognition forward network can be the same, but the finally trained network parameters can be different.

[0073] Through the above implementation manners of the present application, by jointly training multiple initial networks with multiple samples, the network structure finally used for intention recognition is obtained, thereby improving the recognition accuracy of the intention recognition network structure.

[0074] As an optional manner, the above joint training of the initial first attention network, initial first intention recognition network, initial second attention network, initial second intention recognition network and multiple initial slot filling networks by using the training sample text information and the corresponding intention category labels includes:

[0075] S1. Obtain the classification probability value output by the first intention recognition network during training in the current training round, where the classification probability value is used to indicate whether the intention category recognized based on the current training sample text information is consistent with the category indicated by the intention category label;

[0076] S2. Determine the loss value by using the logarithm value of the classification probability value and the modulation factor determined based on the classification probability value, where the modulation factor is determined by the γ-th power of the difference between the value 1 and the classification probability value, and γ is a constant.

[0077] In this embodiment, the slot filling is trained using the CRF loss function, which can effectively solve the data sparsity problem. In this embodiment, the Focal Loss is also introduced as the intention recognition loss function to effectively control the problem of class imbalance. During the process of training the network in by using the sample set and the intention labels corresponding to each sample Figure 3 the above-mentioned networks are constrained by using the Focal Loss function.

[0078] The formula of the Focal Loss function is as follows:[[]]

[0079] Loss fl =-(1 - p t ) γ log(p t )

[0080] Where, p t Let (1-p) be the classification probability value. t ) γ This is equivalent to learning a modulation factor, p, when the sample is correctly classified. t ≈1, the closer the modulation factor is to 0, the smaller the loss; however, when the sample is not accurately classified, p t Since the modulation factor is approximately 0, the loss remains constant. Therefore, this loss function effectively weakens the influence of correctly classified samples and increases the attention given to misclassified or hard-case samples. When the sample class distribution is imbalanced, the class with fewer samples is very difficult to classify. Therefore, focusing on hard-case samples can help address the technical problem of low accuracy for the class with fewer samples when the sample distribution is imbalanced.

[0081] In the embodiments described above, firstly, the slot filling results are introduced into the intent recognition network layer through an attention mechanism to improve intent recognition performance, thereby indirectly improving slot filling performance. During the training phase, a bidirectional attention information flow forms a loop, fully utilizing joint information through iterative iterations. During the testing phase, the intent recognition network layer can also influence intent recognition through the slot filling results. Furthermore, to address the problem of imbalanced intent recognition samples, a Focal Loss function is introduced into intent recognition to improve the recognition performance of categories with fewer samples.

[0082] The following combination Figure 3 A complete process of the above method will be described.

[0083] S1, input the text information vector corresponding to the control text into the pre-trained language model to obtain the text feature vector set;

[0084] like Figure 3 In the lower right corner, four boxes are shown, each containing three hollow circles. Each box indicates the text information vector corresponding to a character. The first text information vector is the text information vector corresponding to the character identifier "CLS". After processing by the pre-trained language model, the text feature vectors corresponding to each character are obtained. Figure 3 The output vector is represented by four boxes. The first box shows two gray dashed circles, indicating that the output vector is the text feature vector corresponding to the text information vector of the character identifier "CLS". The last three boxes show three dashed circles, indicating the text feature vectors corresponding to each text information vector.

[0085] In this embodiment, the pre-trained language model can be selected as the base encoder according to actual needs, taking into account the language, size, and performance of the pre-trained model. As a preferred approach, BERT and its related variants, such as PhoBERT / Albert / Roberta, can be used as the pre-trained language model.

[0086] The method of obtaining text feature vectors through a pre-trained language model can be expressed by the following formula:

[0087] v i =P model (ω,i)

[0088] Where i is used to identify the i-th text feature vector, and v0 is used for the reference feature vector corresponding to the [CLS] flag.

[0089] S2, input the reference feature vector corresponding to the [CLS] flag into the second intent recognition feedforward network to obtain the result vector, and input the result vector and the text feature vector corresponding to each character into the intent-slot attention network to obtain the second output vector set;

[0090] It should be noted that the result vector here corresponds to the intent reference vector mentioned above, and the intent-slot attention network here corresponds to the second attention network mentioned above. Figure 3 In the diagram, the solid black circles indicate the second output vector corresponding to each character.

[0091] In the above implementation, the second intent recognition feedforward network receives the reference feature vector v0 at the [CLS] position, maps it to an intent result vector, and integrates it into the input feature vector of the slot-filling feedforward network through an attention mechanism. The intent result vector (i.e., the intent reference vector mentioned earlier) is a 1*N vector, where N is the number of intent categories. The physical meaning of this vector is to represent the probability value of each intent category.

[0092] In this step, the intent information is incorporated into each slot-filling feature vector as an attention mechanism. The fused features are then concatenated with the original features and input into a single-layer feedforward network for slot filling to obtain the final slot-filling output. The intent information here does not include slot-filling output interaction information; it is taken from a separate branch. Furthermore, when receiving the intent result vector, a dimension transformation layer is first applied to achieve dimension alignment.

[0093] The method for obtaining the aforementioned second output vectors in this step can be expressed by the following formula:

[0094] s i =α i v

[0095] Where v is the output feature vector and weight matrix W of the pre-trained model. v Product; α i = softmax(w·v′), where v′ is the pre-trained model's output feature vector and weight matrix W. k The product of w and the label weight matrix W, where w is the intention vector and W is the label weight matrix. q The product of s and the aforementioned intention category vectors, i This is the final output of the attention module at the i-th position.

[0096] S3, concatenate the second output vector corresponding to each character with the corresponding text feature vector to obtain a concatenated vector set, and input the above concatenated vector set into the feedforward network with multiple slots to obtain a semantic category vector set;

[0097] like Figure 3 In this process, each second output vector (solid black circle) is first concatenated with the text feature vector (dashed circle), and then fed into multiple slot-filling feedforward networks. This network receives the feature vectors output from the pre-trained model and the output from the intent-slot attention module, mapping the concatenated feature vectors to slot-filling outputs.

[0098] In this model, the slot result vector for each location is a 1*P vector, where P represents the number of slot categories, and its physical meaning is the probability value of each slot category at the current location. All location result vectors are integrated into a Q*P vector, where Q is the length of the input data, and then fused into the intent recognition process after passing through a dimension transformation layer.

[0099] S4. Input the semantic category vector set after dimensional transformation and the reference feature vector corresponding to the [CLS] flag into the slot-intent attention network to obtain the first reference vector set;

[0100] It should be noted that the slot-intent attention network in this step corresponds to the first attention network mentioned earlier.

[0101] In this step, slot-filling information is incorporated into the intent recognition feature vector in the form of attention. The fused features are then concatenated with the original features and input into a single-layer feedforward network for intent recognition to obtain the final intent recognition output. It should be noted that the intent recognition output branch is distinct from the intent information branch of the intent-slot attention mechanism.

[0102] The output of this step can be expressed by the following formula:

[0103] i i =β i v0

[0104] Where v0 is the output feature vector and weight matrix W of the pre-trained model. v The product of; βi = softmax(w·v0′), where v0′ is the pre-trained model's output feature vector and weight matrix W. k The product of w and the label weight matrix W, where w is the slot vector and W is the label weight matrix. q The product of i and the semantic category vector set composed of the above semantic category vectors, i i This is the final output of the attention module, namely the first output vector set mentioned above.

[0105] S5, concatenate the first output vector set and the text feature vector corresponding to the [CLS] flag and input them into the first intent recognition network to obtain the intent result vector;

[0106] It should be noted that the aforementioned intent result vector, also known as the intent category vector, is a 1*N vector, where N is the number of intent categories. The physical meaning of this vector is to represent the probability value of each intent category.

[0107] In this embodiment, slot filling is trained using the CRF loss function, which effectively addresses the data sparsity problem. This embodiment also introduces Focal Loss as the intent recognition loss function to effectively control class imbalance. The system utilizes the sample set and the intent labels corresponding to each sample to... Figure 3 During the training of the networks, the Focal Loss function is used to constrain the aforementioned networks.

[0108] The formula for the Focal Loss function is as follows:

[0109] Loss fl =-(1-p t ) γ log(p t )

[0110] Where, p t Let (1-p) be the classification probability value. t ) γ This is equivalent to learning a modulation factor, p, when the sample is correctly classified. t ≈1, the closer the modulation factor is to 0, the smaller the loss; however, when the sample is not accurately classified, p t Since the modulation factor is approximately 0, the loss remains constant when the modulation factor is close to 1. Therefore, this loss function effectively weakens the influence of correctly classified samples and increases the attention given to misclassified or hard-case samples. When the sample class distribution is imbalanced, the class with fewer samples will be very difficult to classify. Therefore, focusing on hard-case samples can help address the problem of low accuracy for the class with fewer samples when the sample distribution is imbalanced.

[0111] The embodiments described in this application employ a bidirectional attention mechanism. Building upon existing slot-intent attention, slot filling output information is integrated into intent recognition features via this attention mechanism. Slot result information guides intent recognition, improving its effectiveness and indirectly influencing slot filling results. Furthermore, the Focal Loss function addresses the sample imbalance problem in intent recognition, improving the recognition accuracy for categories with fewer samples. Two intent branches are employed: one to provide intent information for slot filling and the other for intent recognition. This avoids interdependence during information interaction, resolving technical issues that prevent model training and use.

[0112] According to another aspect of the present invention, a text intent recognition apparatus for implementing the above-described text intent recognition method is also provided. For example... Figure 4 As shown, the device includes:

[0113] Acquisition unit 402 is used to acquire the text information to be recognized;

[0114] The feature extraction unit 404 is used to extract text features from text information to obtain a set of text feature vectors;

[0115] The semantic parsing unit 406 is used to perform slot filling processing on each text feature vector in the text feature vector set to obtain the corresponding semantic category vector. The semantic category vector set is used to indicate the semantic category corresponding to each character in the text information.

[0116] The intent parsing unit 408 is used to perform intent parsing on the text feature vector set and each semantic category vector in order to identify the operational intent carried by the text information.

[0117] Optionally, in this embodiment, the implementation of each of the above-mentioned unit modules can be referred to the above-mentioned method embodiments, which will not be repeated here.

[0118] According to another aspect of the present invention, an electronic device for implementing the above-described text intent recognition method is also provided, the electronic device being... Figure 5 The terminal device or server shown. This embodiment uses the electronic device as an example of a terminal device for illustration. Figure 5 As shown, the electronic device includes a memory 502 and a processor 504. The memory 502 stores a computer program, and the processor 504 is configured to execute the steps in any of the above method embodiments via the computer program.

[0119] Optionally, in this embodiment, the electronic device may be located in at least one of a plurality of network devices in a computer network.

[0120] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0121] S1, Obtain the text information to be recognized;

[0122] S2 is used to extract text features from text information to obtain a set of text feature vectors;

[0123] S3, each text feature vector in the text feature vector set is subjected to slot filling to obtain the corresponding semantic category vector. The semantic category vector set is used to indicate the semantic category corresponding to each character in the text information.

[0124] S4 performs intent parsing on the text feature vector set and each semantic category vector to identify the operational intent carried by the text information.

[0125] Alternatively, as those skilled in the art will understand, Figure 5 The structure shown is for illustrative purposes only. The electronic device can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. Figure 5 This does not limit the structure of the aforementioned electronic device. For example, the electronic device may also include components that are more... Figure 5 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 5 The different configurations shown.

[0126] The memory 502 can be used to store software programs and modules, such as the program instructions / modules corresponding to the text intent recognition method and apparatus in this embodiment of the invention. The processor 504 executes various functional applications and data processing by running the software programs and modules stored in the memory 502, thereby realizing the aforementioned text intent recognition method. The memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 502 may further include memory remotely located relative to the processor 504, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 502 may be used, but is not limited to, storing device control information and other information. As an example, such as... Figure 5As shown, the memory 502 may include, but is not limited to, the acquisition unit 402, feature extraction unit 404, semantic parsing unit 406, and intent parsing unit 408 from the text intent recognition device. Furthermore, it may include, but is not limited to, other module units from the text intent recognition device, which will not be elaborated upon in this example.

[0127] Optionally, the transmission device 506 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 506 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 506 is a radio frequency (RF) module, used for wireless communication with the Internet.

[0128] In addition, the above-mentioned electronic device also includes: a display 508 for displaying the interface device control operation interface; and a connection bus 510 for connecting the various module components in the above-mentioned electronic device.

[0129] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.

[0130] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.

[0131] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0132] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of a computer device reads computer instructions from the computer-readable storage medium, and executes the computer instructions to cause the computer device to perform the aforementioned device control method.

[0133] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0134] S1, Obtain the text information to be recognized;

[0135] S2 is used to extract text features from text information to obtain a set of text feature vectors;

[0136] S3, each text feature vector in the text feature vector set is subjected to slot filling to obtain the corresponding semantic category vector. The semantic category vector set is used to indicate the semantic category corresponding to each character in the text information.

[0137] S4 performs intent parsing on the text feature vector set and each semantic category vector to identify the operational intent carried by the text information.

[0138] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0139] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0140] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0142] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0143] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0144] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for recognizing textual intent, characterized in that, include: Obtain the text information to be recognized; Text features are extracted from the text information to obtain a set of text feature vectors; The reference feature vector located at the start flag in the text feature vector set is input into the second intent recognition network to obtain the intent reference vector; Using attention weights that match the intent reference vector, slot filling is performed on each text feature vector in the text feature vector set to obtain their respective semantic category vectors. The semantic category vector set is used to indicate the semantic category corresponding to each character in the text information. Obtain the reference feature vector and each of the semantic category vectors, wherein the text feature vector set includes the reference feature vector and the character feature vectors corresponding to each character contained in the text information; The process of inputting the reference feature vector and each of the semantic category vectors into a first attention network to obtain a first output vector includes: obtaining a first weight matrix, a second weight matrix, and a third weight matrix that match the first attention network; obtaining the dot product values ​​of the first feature vector and each of the reference semantic category vectors, wherein the first feature vector is determined by the reference feature vector and the first weight matrix, and the reference semantic category vector is determined by the semantic category vector and the second weight matrix; normalizing each of the dot product values ​​to obtain multiple first reference values; and performing a weighted summation of each first reference value with all the second feature vectors to obtain the first output vector, wherein the second feature vector is determined by the reference feature vector and the third weight matrix. The first attention network is used to perform intent parsing on the text feature vector set using attention weights that match the semantic category vectors. The first output vector and the reference feature vector are jointly input into the first intent recognition network to obtain the recognized intent category vector, wherein the intent category vector is used to indicate the operation intent of the text information.

2. The method according to claim 1, characterized in that, The process of using attention weights matched with the intent reference vector to perform slot filling on each text feature vector in the text feature vector set to obtain their respective semantic category vectors includes: The intent reference vector and the text feature vector set are input into the second attention network to obtain multiple second output vectors. The second attention network is used to perform slot filling parsing on the text feature vector set using attention weights that match the intent reference vector. The second output vector corresponds one-to-one with each text feature vector in the text feature vector set. Each of the second output vectors is concatenated with its corresponding text feature vector to obtain multiple concatenated vectors; The multiple spliced ​​vectors are parsed using multiple slot-filling networks connected to the second attention network to obtain their respective semantic category vectors.

3. The method according to any one of claims 1 to 2, characterized in that, Before obtaining the text information to be recognized, the process also includes: Obtain multiple training sample texts and their corresponding intent category labels; The initial first attention network, the initial first intent recognition network, the initial second attention network, the initial second intent recognition network, and multiple initial slot-filling networks are jointly trained using the training sample text information and the corresponding intent category labels until multiple loss values ​​continuously output by the first intent recognition network are all less than the target threshold. The initial second intent recognition network is connected to the initial second attention network, the initial second attention network is connected to the multiple initial slot-filling networks, the multiple initial slot-filling networks are connected to the initial first attention network, and the initial first attention network is connected to the initial first intent recognition network.

4. The method according to claim 3, characterized in that, The step of jointly training the initial first attention network, the initial first intent recognition network, the initial second attention network, the initial second intent recognition network, and multiple initial slot-filling networks using the training sample text information and corresponding intent category labels includes: Obtain the classification probability value output by the first intent recognition network in the current training round, wherein the classification probability value is used to indicate whether the intent category identified based on the text information of the current training sample is consistent with the category indicated by the intent category label; The loss value is determined using the logarithm of the classification probability value and the modulation factor determined based on the classification probability value, wherein the modulation factor is determined by the power of γ of the difference between the value 1 and the classification probability value, where γ is a constant.

5. A text intent recognition device, characterized in that, include: The acquisition unit is used to acquire the text information to be recognized; The feature extraction unit is used to extract text features from the text information to obtain a set of text feature vectors; A semantic parsing unit is used to input the reference feature vector located at the start flag in the text feature vector set into the second intent recognition network to obtain an intent reference vector; using attention weights matched with the intent reference vector, slot filling is performed on each text feature vector in the text feature vector set to obtain their respective semantic category vectors, wherein the semantic category vector set is used to indicate the semantic category corresponding to each character in the text information. The apparatus is further configured to acquire the reference feature vector and each of the semantic category vectors, wherein the text feature vector set includes the reference feature vector and character feature vectors corresponding to each character contained in the text information; inputting the reference feature vector and each of the semantic category vectors into a first attention network to obtain a first output vector, including: acquiring a first weight matrix, a second weight matrix, and a third weight matrix matched with the first attention network; acquiring the dot product values ​​of the first feature vector and each of the reference semantic category vectors, wherein the first feature vector is determined by the reference feature vector and the first weight matrix, and the reference semantic category vectors are determined by the first weight matrix and the second weight matrix, respectively. The semantic category vector and the second weight matrix are determined; each of the dot product values ​​is normalized to obtain multiple first reference values; each first reference value is weighted and summed with all the second feature vectors to obtain the first output vector, the second feature vector is determined by the reference feature vector and the third weight matrix; the first attention network is used to perform intent parsing on the text feature vector set using attention weights matched with the semantic category vector; the first output vector and the reference feature vector are jointly input into the first intent recognition network to obtain the recognized intent category vector, wherein the intent category vector is used to indicate the operation intent of the text information.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 4.

7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 4 through the computer program.