Multi-intention recognition method, knowledge graph-based question and answer method, device and equipment

Through the multi-intent recognition method of RoBERTa pre-training model and dynamic weight fusion, bidirectional gated recurrent unit and local attention mechanism, the problem of multi-intent parsing in converter steelmaking is solved, the accuracy and efficiency of the question-answering model are improved, and intelligent control is supported.

CN120654708APending Publication Date: 2025-09-16NANCHANG YANNUO TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510778875.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional knowledge graph question answering methods based on semantic parsing cannot effectively handle multi-intent parsing, especially in the field of converter steelmaking. It is difficult to meet the multi-intent parsing requirements of complex questions, and it is difficult to achieve precise control by relying on manual experience, which affects production efficiency.

Method used

A multi-intent recognition method that combines the RoBERTa pre-training model with a dynamic weight fusion mechanism, a bidirectional gated recurrent unit, and a local attention mechanism is used. Through feature extraction, encoding, and decoding modules, multi-intent recognition of user questions and parallel recognition of complex relationship attributes are achieved.

Benefits of technology

It improves the ability to represent professional terms and complex semantics in the steelmaking field, reduces misjudgments, improves the response efficiency and accuracy of the question-answering model, and supports intelligent and adaptive control of converter steelmaking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654708A_ABST
    Figure CN120654708A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-intention recognition method, a knowledge graph-based question and answer method, device and equipment, and the method achieves the parallel recognition of composite relation attributes in a user question through feature space decoupling and a local attention mechanism, and effectively constructs a dynamic mapping system of natural language expression and a knowledge graph relation mode. Through dynamic weight fusion, two-way time sequence modeling and a local attention mechanism, the expression capability of professional terms and composite semantics in the steelmaking field is effectively improved; through the multi-intention analysis capability of the multi-intention identification model, the composite relation attribute in the complex question can be identified and processed more accurately, misjudgment is reduced, and the response efficiency of the question and answer model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a multi-intention recognition method, a question-answering method based on a knowledge graph, a device and an apparatus. Background Art

[0002] Steel is widely used in construction, transportation, machinery manufacturing, and other fields, and is an indispensable multifunctional material in the modern industrial system. BOF steelmaking, due to its high efficiency, energy conservation, and low emissions, has gradually become the mainstream steelmaking process. This process has undergone significant technological innovations, from air-oxidized iron smelting to scrap smelting and then to a combination of top-blowing and bottom-blowing furnaces.

[0003] However, the converter steelmaking process is complex, involving multiple operations such as molten iron requirements, scrap steel ratios, and converter gun position control. Relying on manual experience makes precise control difficult and fails to meet the efficiency and precision demands of modern industry. Furthermore, traditional methods are struggling to effectively utilize massive amounts of real-time data and address the unknowns presented by the development of new steel grades. Steel grade manuals are complex and inefficient to search, further impacting production efficiency.

[0004] Most current knowledge graph question-answering methods based on semantic parsing are typically limited to single-intent parsing, which has certain limitations in practical applications. For example, when asking "What are the process flow, molten iron requirements, and included elements for CCA031A and CCA008H steel grades?", multi-intent parsing of the question can provide a deeper understanding of the connection between CCA031A and CCA008H steel grades in terms of process flow, molten iron requirements, and included element properties, providing users with more accurate and rich answers. In addition, while knowledge graph question-answering based on semantic parsing can accurately answer questions in existing knowledge bases, it cannot handle questions outside the knowledge graph, such as predicting the properties of new steel grades.

[0005] To address these challenges, multi-intent analysis plays a vital role in knowledge management. As a structured knowledge management tool, knowledge graphs can integrate dispersed data and empirical knowledge from converter steelmaking. Using question-answering models, they enable multi-intent analysis. This not only uncovers similarities between steel grades but also, combined with the reasoning power of large language models, provides scientific advice for predicting the operational properties of new steel grades. This fusion approach expands the application scope of question-answering models, enhances their intelligence and adaptability, and provides a new technical path for optimized control of converter steelmaking.

[0006] After performing named entity recognition, entity linking, and semantic mapping, the relationship attributes involved in the question need to be mapped and matched with the relationship attributes in the knowledge graph, which is called intent resolution. Intent resolution is one of the most important steps in completing the question-answering process. Currently, most scholars are committed to studying relationship recognition in fields such as medicine, planting, and culture. They complete the relationship recognition operation by converting the task of understanding the user's question intent into a text classification task. However, they ignore the similarities and differences in the relationship attributes between domain entities and only focus on a specific relationship attribute of a certain domain entity. This limits the inherent correlation that may exist between domain entities or between domain relationships during the domain question-answering process, and at the same time curbs the flexibility of user questions and answers. Summary of the Invention

[0007] Based on this, it is necessary to provide a multi-intent recognition method, a knowledge graph-based question-answering method, device and equipment to address the above technical problems.

[0008] A multi-intent recognition method, the method comprising: Preprocess the obtained user question text to obtain a token sequence.

[0009] The token sequence is input into the feature extraction module to obtain a fusion vector. The feature extraction module is used to process the token sequence using the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder, and a dynamic weight fusion mechanism is used to fuse the multi-level semantic representation to obtain a fusion vector.

[0010] The fusion vector is input into the encoding module to obtain a hidden state sequence; the encoding module is used to rationalize the fusion vector using a context-aware network constructed with a bidirectional gated recurrent unit to obtain a hidden state sequence.

[0011] The hidden state sequence is input into the decoding module to obtain the multi-intent recognition result; the decoding module is used to process the hidden state sequence using the local attention mechanism to generate an attention-weighted context vector, and the context vector is processed through the GRU layer, the fully connected layer and the Sigmoid function to achieve multi-label parallel prediction.

[0012] In one embodiment, the feature extraction module includes: a RoBERTa pre-trained model and a dynamic weight fusion mechanism; the RoBERTa pre-trained model includes a 12-layer Transformer encoder.

[0013] Input the token sequence into the feature extraction module to obtain the fusion vector, including: The token sequence is input into the RoBERTa pre-trained model to obtain the multi-level semantic representation output by each layer of Transformer encoder.

[0014] The multi-level semantic representation output by each layer of Transformer encoder is input into the dynamic weight fusion mechanism, weighted summed by learnable weight parameters, and then projected through the fully connected layer to obtain the fusion vector; the learnable weight parameters are calculated and generated by the fully connected layer; i The weight parameters are:

[0015] in, For the i weight parameters, is the Sigmoid activation function, and are two trainable parameters, R is the set of real numbers, For the i The multi-level semantic representation output by the layer Transformer encoder.

[0016] In one embodiment, the decoding module includes: a local attention mechanism, a GRU layer, a fully connected layer, and a Sigmoid function.

[0017] The hidden state sequence is input into the decoding module to obtain the multi-intent recognition results, including: The hidden state sequence is input into the local attention mechanism to obtain the attention-weighted context vector.

[0018] The attention-weighted context vector is concatenated with the hidden state sequence and processed through the GRU layer, fully connected layer, and Sigmoid function to obtain the multi-intent recognition result.

[0019] In one embodiment, in the local attention mechanism: Predict the center position of the attention window based on the hidden state sequence; the center position of the attention window is:

[0020] in, is the center position of the attention window, T is the encoder output sequence length, σ is the Sigmoid function, Wp∈ is the location prediction parameter matrix, is the current state of the encoder, is a learnable parameter.

[0021] According to the hidden state sequence, the energy score of each position is calculated using additive attention in a fixed width window:

[0022] in, is the energy fraction, W q 、W k ∈ are two linear transformation matrices, v∈ is the learnable weight vector, is a learnable parameter.

[0023] The energy scores are normalized using Softmax to generate local attention weights.

[0024] The local attention weight is weighted and summed with the hidden state sequence to obtain the attention-weighted context vector.

[0025] In one embodiment, the loss function of the multi-intent recognition model composed of a feature extraction module, an encoding module, and a decoding module during the training process is:

[0026] in, is the loss of the multi-intent recognition model, σ(.) is the Sigmoid function, ŷ The model predicts logits vector, y∈{0,1} C is the true label.

[0027] A multi-intent recognition device, comprising: The data preprocessing unit is used to preprocess the obtained user question text to obtain a token sequence.

[0028] The feature extraction unit is used to input the token sequence into the feature extraction module to obtain a fusion vector. The feature extraction module is used to process the token sequence using the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder, and to fuse the multi-level semantic representation using a dynamic weight fusion mechanism to obtain a fusion vector.

[0029] The encoding unit is used to input the fusion vector into the encoding module to obtain a hidden state sequence; the encoding module is used to rationally perform the fusion vector on the context-aware network constructed by using a bidirectional gated recurrent unit to obtain a hidden state sequence.

[0030] The decoding unit is used to input the hidden state sequence into the decoding module to obtain the multi-intent recognition result; the decoding module is used to process the hidden state sequence using the local attention mechanism, generate an attention-weighted context vector, and process the context vector through the GRU layer, the fully connected layer and the Sigmoid function to achieve multi-label parallel prediction.

[0031] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any of the above-mentioned multi-intent recognition methods when executing the computer program.

[0032] A question-answering method based on a converter steelmaking knowledge graph, the method comprising: Get the user question text.

[0033] According to the user's question text, use any of the above multi-intent recognition methods to perform intent recognition and obtain a multi-intent recognition result.

[0034] The multi-intent recognition results are input into the preset graph query statement generation model to obtain the graph query statement.

[0035] Based on the graph query statement, a query is performed in the converter steelmaking knowledge graph to generate the answer text.

[0036] A question-answering device based on a converter steelmaking knowledge graph, the device comprising: The user question text obtaining unit is used to obtain the user question text.

[0037] The multi-intent recognition unit is used to perform intent recognition based on the user question text using any of the above multi-intent recognition methods to obtain a multi-intent recognition result.

[0038] The graph query statement determination unit is used to input the multi-intent recognition results into the preset graph query statement generation model to obtain a graph query statement.

[0039] The question-answering unit is used to query the converter steelmaking knowledge graph based on the graph query statement and generate the answer text.

[0040] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned question-answering method based on the converter steelmaking knowledge graph are implemented.

[0041] The above-mentioned multi-intention recognition method, knowledge graph-based question-answering method, device and equipment, the method realizes the parallel recognition of complex relational attributes in user questions through feature space decoupling and local attention mechanism, and effectively constructs a dynamic mapping system between natural language expression and knowledge graph relational pattern; through dynamic weight fusion, bidirectional time series modeling and local attention mechanism, it effectively improves the representation ability of professional terminology and complex semantics in the steelmaking field; through the multi-intention parsing capability of the multi-intention recognition model, it can more accurately identify and process complex relational attributes in complex questions, reduce misjudgment, and improve the response efficiency of the question-answering model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 1 is a flow chart of a multi-intent recognition method according to an embodiment; Figure 2 A framework diagram of a multi-intent recognition model in another embodiment; Figure 3 A schematic flow chart of a question-answering method based on a converter steelmaking knowledge graph in another embodiment; Figure 4 is a diagram of the internal structure of a computer device in one embodiment; Figure 5 A flow chart of a method for constructing converter steelmaking multi-intention recognition data in another embodiment; Figure 6 A multi-intention recognition dataset for converter steelmaking in another embodiment is shown; Figure 7 Graphs showing the loss function and various indicators of a multi-intent recognition model in another embodiment, where (a) is a graph showing the loss function of the multi-intent recognition model, and (b) is a graph showing various indicators of the multi-intent recognition model. Figure 8 1 is a schematic diagram of the F1 value of the multi-intention recognition model in different intention categories in another embodiment, wherein (a) is a schematic diagram of the F1 value of the multi-intention recognition model in categories related to steel grade properties, (b) is a schematic diagram of the F1 value of the multi-intention recognition model in categories related to process parameters, (c) is a schematic diagram of the F1 value of the multi-intention recognition model in categories related to converter smelting operations, (d) is a schematic diagram of the F1 value of the multi-intention recognition model in categories related to steel grade standards, (e) is a schematic diagram of the F1 value of the multi-intention recognition model in categories related to deoxidation alloying operations, and (f) is a schematic diagram of the F1 value of the multi-intention recognition model in categories related to conditions for molten steel supply for refining. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0044] In one embodiment, Figure 1 As shown, a multi-intent recognition method is provided, which includes the following steps: Step 100: Preprocess the acquired user question text to obtain a token sequence.

[0045] Step 102: Input the token sequence into the feature extraction module to obtain a fusion vector; the feature extraction module is used to process the token sequence using the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder, and use a dynamic weight fusion mechanism to fuse the multi-level semantic representation to obtain a fusion vector.

[0046] Specifically, in the feature extraction module, the RoBERTa pre-training model with dynamic weight fusion mechanism is innovatively introduced. To address the problem of limited representation ability of traditional single-layer Transformer, this module uses learnable weight parameters to The outputs of RoBERTa's 12-layer Transformer encoder are dynamically weighted and fused. Specifically, the 768-dimensional semantic representations generated by each Transformer layer are dynamically weighted, and the weighted summation forms a fused vector with multi-level semantic information. This design effectively captures the semantic features of text at different levels of abstraction, significantly improving the ability to represent specialized terminology and complex intent in the steelmaking field.

[0047] Step 104: Input the fused vector into the encoding module to obtain a hidden state sequence; the encoding module is used to rationalize the fused vector using a context-aware network constructed using a bidirectional gated recurrent unit to obtain a hidden state sequence.

[0048] Specifically, the encoding module uses a bidirectional gated recurrent unit (BiGRU) to construct a context-aware network. Through bidirectional temporal modeling, this module converts the fused vector output by feature extraction into a hidden state sequence that embodies bidirectional semantic dependencies. The BiGRU layer adaptively adjusts the propagation strength of forward and backward information through reset and update gate mechanisms, accurately modeling the semantic associations between text in the converter steelmaking domain while preserving long-range dependencies. The encoding module inputs the hidden states of all time steps and the last time step into the decoding module. This ensures that the encoding module incorporates both global semantic features and local contextual information, providing a high-density semantic representation for subsequent decoding.

[0049] In the study of recurrent neural networks (RNNs), the vanishing gradient problem has long hampered the model's ability to model long-term dependencies. Traditional RNNs require continuous multiplication of gradients along time steps during backpropagation, resulting in a near-stagnant weight update at early time steps. To address this issue, the Gated Recurrent Unit (GRU) introduces a gating mechanism to dynamically control the flow of information. GRU The core structure of includes reset gate and update gate, and the corresponding calculation formula is as follows:

[0050]

[0051] The reset gate r t The Sigmoid activation function generates a scalar between 0 and 1 to control the hidden state of the previous moment h t-1 For the current candidate status The contribution of r t When it approaches 0, the model will ignore historical information and focus on the current input x t , this feature makes it good at capturing short-term dependencies in the sequence. The calculation of the candidate state is as follows:

[0052] The update gate determines the proportion of the current state that inherits the historical memory. The final hidden state is generated by weighted fusion of the historical state and the candidate state. The calculation formula is as follows:

[0053] When z t When it approaches 1, the model tends to retain the current input information, effectively modeling long-term dependencies across multiple time steps. To further enhance context-awareness, this application adopts a bidirectional gated recurrent unit (BiGRU) architecture. The BiGRU consists of two independent GRU layers, one in the forward direction and the other in the reverse direction, processing the input sequence X = {x1, x2, ..., xt} in time and in reverse time, respectively. The hidden state calculation process is as follows:

[0054] in and denote the hidden states from left to right and from right to left, respectively. x t Indicates the first t elements, [;] represents the vector splicing operation, which splices the hidden states in two directions together to obtain a bidirectional hidden state .

[0055] Through bidirectional information fusion, BiGRU not only strengthens the contextual relevance of local features, but also builds a global semantic network across time steps, providing a robust sequence representation basis for the downstream local attention mechanism and GRU decoding module.

[0056] Step 106: Input the hidden state sequence into the decoding module to obtain the multi-intent recognition result; the decoding module is used to process the hidden state sequence using the local attention mechanism to generate an attention-weighted context vector, and process the context vector through the GRU layer, the fully connected layer and the Sigmoid function to achieve multi-label parallel prediction.

[0057] Specifically, the decoding module innovatively constructs a hierarchical decoding architecture guided by local attention. The local attention mechanism dynamically focuses on the encoder hidden state, uses a sliding window to constrain the attention range, and strengthens the association constraints between adjacent operation parameter labels. This mechanism generates an attention weight matrix with spatial locality through interactive calculations between a trainable query vector (Query) and the encoding output. The attention-weighted context vector is then multimodally fused with the GRU hidden state, and multi-label parallel prediction is achieved through a fully connected layer and a Sigmoid activation function. This hierarchical decoding strategy effectively solves the label coupling problem in the steelmaking field in scenarios where multiple intents co-occur while ensuring computational efficiency. This model breaks through the limitations of traditional single-intent parsing frameworks through the coordinated optimization of dynamic feature fusion, bidirectional context encoding, and local attention decoding.

[0058] By utilizing the dynamic fusion of model features and the local attention mechanism, we can deeply explore the potential correlation between the operating characteristics of different steel grades and provide data support and scientific suggestions for the development of new steel grades. The multi-intention recognition model used in the multi-intention recognition method includes three core modules: feature extraction module, encoding module and decoding module. The construction of the multi-intention recognition model is as follows: Figure 2 shown.

[0059] In the above-mentioned multi-intent recognition method, the method realizes the parallel recognition of complex relational attributes in user questions through feature space decoupling and local attention mechanism, and effectively constructs a dynamic mapping system between natural language expression and knowledge graph relational pattern; through dynamic weight fusion, bidirectional time series modeling and local attention mechanism, it effectively improves the representation ability of professional terminology and complex semantics in the steelmaking field; through the multi-intent parsing capability of the multi-intent recognition model, it can more accurately identify and process the complex relational attributes in complex questions, reduce misjudgment, and improve the response efficiency of the question-answering model.

[0060] In one embodiment, the feature extraction module includes: a RoBERTa pre-trained model and a dynamic weight fusion mechanism; the RoBERTa pre-trained model includes a 12-layer Transformer encoder; step 102 includes: inputting the token sequence into the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder; inputting the multi-level semantic representation output by each layer of the Transformer encoder into the dynamic weight fusion mechanism, performing weighted summation through learnable weight parameters, and then projecting through a fully connected layer to obtain a fusion vector; wherein the learnable weight parameters are calculated and generated by the fully connected layer; step 103 includes: inputting the token sequence into the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder; inputting the multi-level semantic representation output by each layer of the Transformer encoder into the dynamic weight fusion mechanism, performing weighted summation through learnable weight parameters, and then projecting through a fully connected layer to obtain a fusion vector; wherein the learnable weight parameters are calculated and generated by the fully connected layer; step 104 includes: i The expression of the weight parameter is:

[0061] in, For the i weight parameters, is the Sigmoid activation function, and are two trainable parameters, R is the set of real numbers, For the i The multi-level semantic representation output by the layer Transformer encoder.

[0062] Specifically, each layer of Transformer encoder in BERT or RoBERTa has a different understanding of text. The dynamic weight fusion mechanism proposed in this application aims to effectively integrate the multi-level semantic representations output by the 12-layer Transformer encoder of the RoBERTa model, allowing the model to dynamically select the semantic representation that is more important for the classification task. The mechanism first extracts the output vector of each Transformer encoder layer. respresent i ∈R 768 (i=1,2,...,12), and then through the learnable weight parameters Perform weighted fusion. Specifically, each weight parameter It is generated by the fully connected layer, as mentioned above i The expressions of the weight parameters are shown in Figure 2.

[0063] Through this process, the feature extraction module can adaptively assign different contribution weights to different layers. For example, the shallow layer may capture local grammatical features, while the deep layer models global logical associations. After the weight calculation is completed, the representations of each layer are fused through weighted summation to obtain the fusion vector ; The expression of the fusion vector is:

[0064] in, is the fusion vector.

[0065] This fusion vector Fused_Vector∈R768 It is further projected into the 768-dimensional semantic space through the fully connected layer to obtain the final vector output after dynamic weight fusion. The final expression of the vector output after dynamic weight fusion is:

[0066] Where Wd∈R 768×768 , bd∈R 768 is a trainable parameter. This design not only alleviates the overfitting problem that high-dimensional features can cause, but also enhances the model's ability to focus on domain-specific semantics. Compared to traditional static average fusion methods, the dynamic weight fusion mechanism achieves autonomous allocation of layer contributions through end-to-end training.

[0067] In one embodiment, the decoding module includes: a local attention mechanism, a GRU layer, a fully connected layer, and a Sigmoid function; step 106 includes: inputting the hidden state sequence into the local attention mechanism to obtain an attention-weighted context vector; concatenating the attention-weighted context vector with the hidden state sequence and processing them through the GRU layer, the fully connected layer, and the Sigmoid function to obtain a multi-intent recognition result.

[0068] In one embodiment, step 106 is used in a local attention mechanism to predict the center position of the attention window based on the hidden state sequence; the expression of the center position of the attention window is:

[0069] in, is the center position of the attention window, T is the encoder output sequence length, σ is the Sigmoid function, Wp∈ is the location prediction parameter matrix, is the current state of the encoder, is a learnable parameter.

[0070] The energy score of each position is calculated using additive attention in a fixed-width window according to the hidden state sequence; the expression of the energy score is:

[0071] in, is the energy fraction, W q 、W k ∈ is the linear transformation matrix, v∈ is the learnable weight vector, , is a learnable parameter.

[0072] The energy score is normalized using Softmax to generate local attention weights; the local attention weights are weighted and summed with the hidden state sequence to obtain the attention-weighted context vector.

[0073] Specifically, the local attention mechanism adopted in this application effectively solves the computational redundancy and semantic sparsity problems of traditional global attention in long sequence tasks by dynamically constraining the attention range. This mechanism consists of three core links. On the one hand, the center position of the attention window is predicted based on the current hidden state of the decoder, and a dynamic positioning signal is generated through a learnable position prediction network, so that the window can move adaptively with the semantic focus; on the other hand, the additive attention score is calculated within a fixed-width window, which significantly reduces the computational complexity; in addition, the local attention weight is generated through Softmax normalization, and the contextual information within the window is aggregated. When a given encoder output sequence Henc={h1enc,..., hTenc} and the current state of the encoder htenc , the window center position is calculated using the above expression of the window center position. This design enables the window center to move adaptively with the semantic focus. In the encoder, the energy score of each position is calculated using additive attention. For the encoder hidden state Henc, its energy score The definition is as shown in the expression for the energy fraction above. Determined by hyperparameter tuning. The energy score is normalized by Softmax to generate the local attention weight α t,i , and weighted sum to get the attention-weighted context vector c t ∈ , the calculation formula is:

[0074]

[0075] in, α t,i is the local attention weight, c t The context vector weighted for attention.

[0076] This process is only for the window The calculation is performed at each position, which reduces the complexity from the global attention down to Finally, the context vector Embedded with the current input After splicing, input the GRU unit and the update formula is as follows:

[0077] in, The current state of the encoder is the final output.

[0078] In one embodiment, the loss function of the multi-intent recognition model composed of a feature extraction module, an encoding module, and a decoding module during the training process is expressed as:

[0079] in, is the loss of the multi-intent recognition model, σ(.) is the Sigmoid function, which maps logits to the interval [0,1], indicating the independent probability of each label. ŷ The model predicts logits vector, y∈{0,1} C is the true label.

[0080] Specifically, this method uses the binary cross entropy loss function as the optimization target of the multi-intent recognition model. By combining the Sigmoid activation function with the cross entropy loss calculation, it effectively adapts to the probability output characteristics of the multi-label classification task. x , the model predicts logits vector (Because at this time is the unnormalized raw output, so the output should be a real number vector without constraints) and the true label , where C is the total number of categories. The loss function is calculated using the above loss function expression. This loss function can optimize the prediction results of multiple intent labels simultaneously by calculating the loss independently for each label and averaging them.

[0081] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0082] In one embodiment, a multi-intent recognition device is provided, comprising: a data preprocessing unit, a feature extraction unit, an encoding unit, and a decoding unit, wherein: The data preprocessing unit is used to preprocess the obtained user question text to obtain a token sequence.

[0083] The feature extraction unit is used to input the token sequence into the feature extraction module to obtain a fusion vector. The feature extraction module is used to process the token sequence using the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder, and to fuse the multi-level semantic representation using a dynamic weight fusion mechanism to obtain a fusion vector.

[0084] The encoding unit is used to input the fusion vector into the encoding module to obtain a hidden state sequence; the encoding module is used to rationally perform the fusion vector on the context-aware network constructed by using a bidirectional gated recurrent unit to obtain a hidden state sequence.

[0085] The decoding unit is used to input the hidden state sequence into the decoding module to obtain the multi-intent recognition result; the decoding module is used to process the hidden state sequence using the local attention mechanism, generate an attention-weighted context vector, and process the context vector through the GRU layer, the fully connected layer and the Sigmoid function to achieve multi-label parallel prediction.

[0086] In one embodiment, the feature extraction module includes: a RoBERTa pre-trained model and a dynamic weight fusion mechanism; the RoBERTa pre-trained model includes a 12-layer Transformer encoder; the encoding unit is also used to input the token sequence into the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder; the multi-level semantic representation output by each layer of the Transformer encoder is input into the dynamic weight fusion mechanism, weighted summed by learnable weight parameters, and then projected through a fully connected layer to obtain a fusion vector; wherein the learnable weight parameters are calculated and generated by the fully connected layer; the i The weight parameters are:

[0087] in, For the i weight parameters, is the Sigmoid activation function, and are two trainable parameters, R is the set of real numbers, For the iThe multi-level semantic representation output by the layer Transformer encoder.

[0088] In one embodiment, the decoding module includes: a local attention mechanism, a GRU layer, a fully connected layer, and a Sigmoid function; the decoding unit is also used to input the hidden state sequence into the local attention mechanism to obtain an attention-weighted context vector; the attention-weighted context vector is concatenated with the hidden state sequence and processed through the GRU layer, the fully connected layer, and the Sigmoid function to obtain a multi-intent recognition result.

[0089] In one embodiment, the decoding unit is further configured to, in a local attention mechanism, predict the center position of the attention window based on the hidden state sequence; the center position of the attention window is:

[0090] in, is the center position of the attention window, T is the encoder output sequence length, σ is the Sigmoid function, Wp∈ is the location prediction parameter matrix, is the current state of the encoder, is a learnable parameter.

[0091] According to the hidden state sequence, the energy score of each position is calculated using additive attention in a fixed width window:

[0092] in, is the energy fraction, W q 、W k ∈ are two linear transformation matrices, v∈ is the learnable weight vector, is a learnable parameter.

[0093] The energy score is normalized using Softmax to generate local attention weights; the local attention weights are weighted and summed with the hidden state sequence to obtain the attention-weighted context vector.

[0094] In one embodiment, the loss function of the multi-intent recognition model composed of a feature extraction module, an encoding module, and a decoding module during the training process is:

[0095] in, is the loss of the multi-intent recognition model, σ(.) is the Sigmoid function, ŷ The model predicts logits vector, y∈{0,1} C is the true label.

[0096] For the specific definition of the multi-intention recognition device, please refer to the definition of the multi-intention recognition method above, which will not be repeated here. The various modules in the above-mentioned multi-intention recognition device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0097] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a multi-intention recognition method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0098] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0099] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned multi-intent recognition method embodiment when executing the computer program.

[0100] In one embodiment, in order to solve the problems of information dispersion, reliance on manual experience and low data utilization efficiency in traditional processes, such as Figure 4A question-answering method based on a converter steelmaking knowledge graph is provided, the method comprising: Step 200: Obtain user question text.

[0101] Step 202: Perform intent recognition based on the user question text using any of the above multi-intent recognition methods to obtain a multi-intent recognition result.

[0102] Specifically, by constructing a high-quality converter steelmaking knowledge graph, integrating multi-source information, and adopting a multi-intent parsing model, accurate parsing of complex questions and multi-intent recognition are achieved.

[0103] The multi-intent parsing capability of the multi-intent recognition model can more accurately identify and process the complex relationship attributes in complex questions, reduce misjudgments, and improve the response efficiency of the question-answering model.

[0104] Step 204: Input the multi-intent recognition results into the preset graph query statement generation model to obtain a graph query statement.

[0105] Step 206: Based on the graph query statement, query the converter steelmaking knowledge graph and generate an answer text.

[0106] This method provides a knowledge graph and question-answering method based on the converter steelmaking domain. It also provides technical support for intelligent knowledge management and optimized control in complex industrial scenarios, and also provides an important reference for future cross-domain applications. Compared with existing methods, this method significantly enhances the associative reasoning ability across entity relationship networks while maintaining domain characteristics, providing a new technical path for multidimensional knowledge query in complex industrial scenarios.

[0107] In a verification example, since there is currently no high-quality dataset for the classification of converter steelmaking problems, this method is constructed by setting a relationship slot correspondence table and applying a prescribed template, such as Figure 5 Each relationship slot can be expressed in over 30 different ways. Fixed templates include single-category, double-category, and triple-category templates, each of which specifies 30 different expressions. Based on the ontological relationships, entity relationships, and entity attribute categories in the converter steelmaking domain discussed in Chapter 2, this study categorizes question intent into 30 categories.

[0108] The specific information of the question intent category is shown in Table 1, which includes 25 relationships defined in the converter steelmaking knowledge graph and 5 ontology relationships. In order to facilitate the subsequent relationship mapping, most of the intent labels are the same as the entity relationship and ontology relationship labels. In addition, in order to ensure the balance of the sample size of the question category, this embodiment shuffles all samples and then divides them into training set, validation set and test set in a ratio of 7:2:1. The final training set is 27015, the test set is 7718, and the validation set is 3863. The constructed dataset is as follows Figure 6 shown.

[0109] Table 1 Question intention categories

[0110] (1) Experimental environment and parameter settings This example uses the Pytorch framework to build the experimental model, using the same experimental environment as the named entity recognition experiment. For experimental parameter settings, the RoBERTa pre-trained language model was used, consisting of a 12-layer Transformer structure and 768 hidden dimensions. Table 2 shows the hyperparameter settings used in the experiment, with a minimum number of samples per training run set to 64, a maximum sequence length of 72, a learning rate of 3e-5, and 30 iterations to ensure optimal model performance.

[0111] Table 2 Hyperparameter settings for intent recognition experiments

[0112] (2) Experimental results and analysis In the multi-intent recognition task, this embodiment uses precision, recall, and F1 score to evaluate the final quality of the model.

[0113] The training and testing results of the multi-intent recognition model are as follows: Figure 7 As shown, Figure 7 (a) in the figure is the loss function curve of the multi-intent recognition model. Figure 7(b) shows the performance of the multi-intent recognition model. After 30 epochs of iterative optimization, the training set loss converged to 0.0063, and the test set loss stabilized at 0.0087, demonstrating the model's strong generalization ability on unseen data. Regarding evaluation metrics, the precision, recall, and F1-score on the training set reached 95.68%, 96.39%, and 96.92%, respectively, while the corresponding metrics on the test set were 94.12%, 95.33%, and 95.63%, respectively. This performance degradation was less than 2%, demonstrating the model's robustness to the multi-intent parsing task in converter steelmaking. Notably, the training curve exhibited a slow convergence trend in the first 10 epochs. This is primarily due to the 30-class multi-label classification task (where a single sample contains at most three labels). The model must simultaneously learn both independence and potential correlations between labels, making it difficult to capture complex label co-occurrence patterns in the initial stages. Secondly, during the prediction phase, a fixed threshold (0.7) was used to binarize logits. This means that a label is considered present if the predicted probability is greater than 0.7. This hard threshold strategy exacerbated optimization difficulty during the initial training phase, causing the model's initial output to tend toward 0. However, as training progressed, the model gradually adjusted its parameters to accommodate the threshold constraint, resulting in a rapid improvement in prediction performance over the last 20 epochs. Furthermore, in multi-label classification tasks, predictions are considered correct only when all relevant labels are correct. Misclassification of a single label can also lead to sample-level errors. This strict evaluation mechanism amplifies errors introduced by the randomness of the model's initial parameters. Figure 8 More specifically, the F1 performance of the multi-intent recognition model in various categories is shown. The overall trend shows that the convergence is slow in the early stage, but it improves rapidly in the last 20 epochs. Figure 8 (a) is a schematic diagram of the F1 value of the multi-intention recognition model on the steel attribute related categories. Figure 8 (b) is a schematic diagram of the F1 value of the multi-intention recognition model on the process parameter related categories. Figure 8 (c) is a schematic diagram of the F1 value of the multi-intention recognition model on categories related to converter smelting operations. Figure 8 (d) is a schematic diagram of the F1 value of the multi-intention recognition model on the steel grade standard related categories. Figure 8 (e) is a schematic diagram of the F1 value of the multi-intention recognition model on the deoxidation alloy operation related categories. Figure 8 Figure (f) shows the F1 score of the multi-intent recognition model for categories related to molten steel supply conditions. This model effectively improves its representation of specialized terminology and complex semantics in the steelmaking field through dynamic weight fusion, bidirectional temporal modeling, and a local attention mechanism. Experimental results demonstrate that it outperforms existing methods in precision, recall, and F1 score, demonstrating its robustness and generalization capabilities in multi-intent resolution tasks.

[0114] To verify the effectiveness of the model, this example selected the RoBERTa and RoBERTa-TextCNN models as comparison methods and conducted experiments on self-developed converter steelmaking multi-intent recognition training data. The results show that the proposed model (multi-intent recognition model) outperforms the other two models in all metrics. Compared with the RoBERTa-TextCNN model, the precision, recall, and F1 score increased by 2.78%, 3.65%, and 3.89%, respectively. The experimental results are shown in Table 3. To demonstrate the effectiveness of the various modules of the multi-intent recognition model in this application, ablation comparison experiments were conducted on the local attention mechanism and the dynamic fusion mechanism, as shown in Table 4. The experimental results show that after removing the local attention mechanism, the F1 score of the multi-intent recognition model dropped from 95.63% to 94.82%, a decrease of 0.81%. This indicates that the local attention mechanism can enhance the ability to capture contextual dependencies by focusing on key information in the question. Removing this module will reduce the model's sensitivity to local semantics. When the dynamic fusion mechanism is removed, the F1 score drops significantly to 93.29%, demonstrating its central role in multi-level feature integration. Dynamic fusion integrates Transformer features from each layer through an adaptive weight allocation strategy, allowing the model to automatically select features that are more important for intent recognition based on actual semantic information, thereby improving recognition efficiency.

[0115] Table 3 Experimental results of converter steelmaking intention recognition

[0116] Table 4 Ablation experiment results

[0117] Ablation experiments have verified that the local attention and dynamic fusion mechanisms effectively improve the performance of the model in multi-label classification tasks and enhance the model's application capabilities in complex industrial scenarios.

[0118] In one embodiment, a question-answering device based on a converter steelmaking knowledge graph is provided, comprising: a user question text acquisition unit, a multi-intent recognition unit, a graph query statement determination unit, and a question-answering unit, wherein: The user question text obtaining unit is used to obtain the user question text.

[0119] The multi-intent recognition unit is used to perform intent recognition based on the user question text using any of the above multi-intent recognition methods to obtain a multi-intent recognition result.

[0120] The graph query statement determination unit is used to input the multi-intent recognition results into the preset graph query statement generation model to obtain a graph query statement.

[0121] The question-answering unit is used to query the converter steelmaking knowledge graph based on the graph query statement and generate the answer text.

[0122] Regarding the specific limitations of the question-answering device based on the converter steelmaking knowledge graph, please refer to the limitations of the question-answering method based on the converter steelmaking knowledge graph above, which will not be repeated here. The various modules in the above-mentioned question-answering device based on the converter steelmaking knowledge graph can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0123] In one embodiment, an electronic device is provided, which may be a terminal. The electronic device includes a processor, memory, a network interface, a display, and an input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a question-and-answer method based on a converter steelmaking knowledge graph. The display of the electronic device may be a liquid crystal display or an electronic ink display. The input device of the electronic device may be a touch layer covering the display, or may be buttons, a trackball, or a touchpad provided on the electronic device housing, or may be an external keyboard, touchpad, or mouse.

[0124] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps in the embodiment of the question-answering method based on the converter steelmaking knowledge graph are implemented.

[0125] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0126] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A multi-intent recognition method, characterized in that: The method comprises: Preprocess the obtained user question text to obtain a token sequence; Inputting the token sequence into a feature extraction module to obtain a fusion vector; the feature extraction module is used to process the token sequence using a RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder, and fusing the multi-level semantic representation using a dynamic weight fusion mechanism to obtain a fusion vector; The fusion vector is input into an encoding module to obtain a hidden state sequence; the encoding module is used to rationally process the fusion vector using a context-aware network constructed by a bidirectional gated recurrent unit to obtain a hidden state sequence; The hidden state sequence is input into the decoding module to obtain a multi-intent recognition result; the decoding module is used to process the hidden state sequence using a local attention mechanism to generate an attention-weighted context vector, and the context vector is processed through a GRU layer, a fully connected layer and a Sigmoid function to achieve multi-label parallel prediction.

2. The multi-intention recognition method according to claim 1, characterized in that: The feature extraction module includes: a RoBERTa pre-trained model and a dynamic weight fusion mechanism; the RoBERTa pre-trained model includes a 12-layer Transformer encoder; The token sequence is input into the feature extraction module to obtain a fusion vector, including: Input the token sequence into the RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder; The multi-level semantic representation output by each layer of Transformer encoder is input into the dynamic weight fusion mechanism, weighted summed by learnable weight parameters, and then projected through the fully connected layer to obtain the fusion vector; the learnable weight parameters are calculated and generated by the fully connected layer; i The weight parameters are: in, For the i weight parameters, is the Sigmoid activation function, and are two trainable parameters, R is the set of real numbers, For the i The multi-level semantic representation output by the layer Transformer encoder.

3. The multi-intention recognition method according to claim 1, characterized in that: The decoding module includes: a local attention mechanism, a GRU layer, a fully connected layer, and a Sigmoid function; The hidden state sequence is input into the decoding module to obtain the multi-intent recognition result, including: Inputting the hidden state sequence into the local attention mechanism to obtain an attention-weighted context vector; The attention-weighted context vector is concatenated with the hidden state sequence and processed through the GRU layer, the fully connected layer, and the Sigmoid function to obtain a multi-intent recognition result.

4. The multi-intention recognition method according to claim 3, characterized in that: In the local attention mechanism: The center position of the attention window is predicted according to the hidden state sequence; the center position of the attention window is: in, is the center position of the attention window, T is the encoder output sequence length, σ is the Sigmoid function, Wp∈ is the location prediction parameter matrix, is the current state of the encoder, is a learnable parameter; According to the hidden state sequence, the energy score of each position is calculated using additive attention in a fixed width window: in, is the energy fraction, W q ,W k ∈ are two linear transformation matrices, v∈ is the learnable weight vector, is a learnable parameter; The energy scores are normalized using Softmax to generate local attention weights; The local attention weight is weighted and summed with the hidden state sequence to obtain an attention-weighted context vector.

5. The multi-intention recognition method according to claim 1, characterized in that: The loss function of the multi-intent recognition model composed of feature extraction module, encoding module and decoding module during the training process is: in, is the loss of the multi-intent recognition model, σ(.) is the Sigmoid function, The model predicts logits vector, ∈{0,1} C is the true label.

6. A multi-intention recognition device, characterized in that: The device comprises: The data preprocessing unit is used to preprocess the obtained user question text to obtain a token sequence; A feature extraction unit is configured to input the token sequence into a feature extraction module to obtain a fusion vector; the feature extraction module is configured to process the token sequence using a RoBERTa pre-trained model to obtain a multi-level semantic representation output by each layer of the Transformer encoder, and to fuse the multi-level semantic representation using a dynamic weight fusion mechanism to obtain a fusion vector; The encoding unit is used to input the fusion vector into the encoding module to obtain a hidden state sequence; the encoding module is used to rationally process the fusion vector using a context-aware network constructed by a bidirectional gated recurrent unit to obtain a hidden state sequence; A decoding unit is used to input the hidden state sequence into a decoding module to obtain a multi-intent recognition result; the decoding module is used to process the hidden state sequence using a local attention mechanism to generate an attention-weighted context vector, and process the context vector through a GRU layer, a fully connected layer, and a Sigmoid function to achieve multi-label parallel prediction.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multi-intent recognition method according to any one of claims 1 to 5 are implemented.

8. A question-answering method based on converter steelmaking knowledge graph, characterized in that: The method comprises: Get the user question text; Performing intent recognition using the multi-intent recognition method according to any one of claims 1 to 5 according to the user question text to obtain a multi-intent recognition result; Inputting the multi-intent recognition result into a preset graph query statement generation model to obtain a graph query statement; According to the graph query statement, a query is performed in the converter steelmaking knowledge graph to generate an answer text.

9. A question-answering device based on converter steelmaking knowledge graph, characterized in that: The device comprises: A user question text acquisition unit, used to acquire the user question text; a multi-intent recognition unit, configured to perform intent recognition according to the user question text using the multi-intent recognition method according to any one of claims 1 to 5 to obtain a multi-intent recognition result; A graph query statement determination unit, configured to input the multi-intent recognition result into a preset graph query statement generation model to obtain a graph query statement; The question-answering unit is used to query the converter steelmaking knowledge graph according to the graph query statement and generate an answer text.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the question-answering method based on the converter steelmaking knowledge graph as described in claim 8 are implemented.

Citation Information

Cited By

  • Intention recognition response method and system based on forest farmer question and answer data

    CN121542442A