Intelligent dialogue method, system and device based on business scene and medium

By obtaining user input text information in the intelligent dialogue system, dynamically selecting processing models, identifying intentions and performing business operations, the problem of low efficiency and accuracy of traditional question-and-answer systems in complex business processes is solved, and efficient and flexible business process processing is achieved.

CN119938822APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411814378.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When facing complex and multi-step business processes, traditional question-and-answer systems have low answer efficiency and accuracy, high cost, and large-scale technology resources are consumed and uncertain, so model calls are not flexible enough.

Method used

By obtaining the target text information entered by the user, selecting an appropriate target processing model based on the preset model call information, performing intention identification, determining business processes and business operations, and executing business operations to obtain processing results.

Benefits of technology

It realizes the rapid and accurate identification of user intentions by the intelligent dialogue system in complex business processes, and flexibly arrange the process, improving service response speed and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938822A_ABST
    Figure CN119938822A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent dialogue method, system and device based on a business scene and a medium. The method comprises the steps that target text information input by a user is acquired; selecting a target processing model corresponding to the target text information according to preset model calling information; performing intention recognition on the target text information according to the target processing model, and generating an intention prediction result of the target text information; determining a business process and a business operation corresponding to the target text information according to the intention prediction result; according to the method, the business operation is executed, the business processing result of the target text information is obtained, different processing models are flexibly called through the intelligent dialogue to quickly and accurately recognize the intention of the user, then the intention of the user is mapped to the corresponding business process, the corresponding business operation is triggered, the flexible process arrangement service is achieved, and the personalized requirements of the user are met. And the response speed of the service and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent dialogue technology, and in particular to an intelligent dialogue method based on a business scenario, an intelligent dialogue system based on a business scenario, an electronic device, and a computer-readable storage medium. Background Art

[0002] In today's wave of digital transformation, enterprises have an increasing demand for intelligence and automation. As an important part of enterprise intelligent services, traditional question-answering systems are facing new challenges. Traditional question-answering systems are usually based on rules or simple machine learning models and can handle some basic question-answering tasks, but their limitations gradually become apparent when faced with complex, multi-step business processes.

[0003] Existing technologies usually use large model technology to fine-tune models to adapt to different business needs. However, although adjustments based on large models can improve accuracy, large model technology requires a large amount of knowledge base and computing resource support, resulting in high resource consumption and uncertainty. In addition, model calling technology is not flexible enough, and it is difficult to dynamically select and call the most appropriate model according to specific business scenarios, affecting the adaptability and scalability of the system. Summary of the invention

[0004] The embodiments of the present invention provide an intelligent dialogue method, system, device and medium based on business scenarios to solve or partially solve the problems that the existing traditional question-and-answer system has limitations when facing complex and multi-step business processes, resulting in low answering efficiency and answering accuracy and high cost.

[0005] The embodiment of the present invention discloses an intelligent dialogue method based on a business scenario, the method comprising:

[0006] Get the target text information entered by the user;

[0007] Selecting a target processing model corresponding to the target text information according to preset model calling information;

[0008] Performing intent recognition on the target text information according to the target processing model to generate an intent prediction result of the target text information;

[0009] Determine the business process and business operation corresponding to the target text information according to the intention prediction result;

[0010] Execute the business operation to obtain the business processing result of the target text information.

[0011] In some feasible implementations, the step of obtaining target text information input by a user includes:

[0012] Get the initial text information entered by the user;

[0013] Recognize the initial text information through a natural language understanding model to obtain key information of the initial text;

[0014] generating guidance information according to the key information;

[0015] Obtaining the supplementary information input by the user according to the guidance information;

[0016] According to the key information and the supplementary information, target text information corresponding to the initial text information is generated.

[0017] In some feasible implementations, the obtaining of initial text information input by the user includes:

[0018] Obtaining first text information input by a user;

[0019] Performing semantic analysis on the first text information to generate a semantic analysis result;

[0020] Determining whether the first text information has a semantic error according to the semantic analysis result;

[0021] If there is a semantic error, generating feedback information for the first text information;

[0022] According to the feedback information, second text information input by the user is obtained, and the second text information is used as the initial text information.

[0023] In some feasible implementations, the target processing model includes a target context perception model and a target intent classification model, and the selecting the target processing model corresponding to the target text information according to the preset model calling information includes:

[0024] Analyze the target text information according to the model calling information to obtain the processing requirements of the target text information;

[0025] According to the processing requirements, a context-aware large model or a context-aware small model is selected from a preset model pool as a target context-aware model, and a large intent classification model or a small intent classification model is selected as a target intent classification model.

[0026] In some feasible implementations, the target processing model includes a target context perception model and a target intent classification model, and the performing intent recognition on the target text information according to the target processing model to generate an intent prediction result of the text information includes:

[0027] Performing word segmentation processing on the target text information to obtain a word vector corresponding to the target text information;

[0028] Processing the word vector by the target context-aware model to obtain context information of the target text information;

[0029] The target intention classification model is used to perform intention recognition on the target text information according to the context information to obtain an intention prediction result of the target text information.

[0030] In some feasible implementations, the method further includes:

[0031] Automatically search the user information database to obtain the user's job information;

[0032] The step of performing intent recognition on the target text information according to the context information by using the target intent classification model to obtain an intent prediction result of the target text information includes:

[0033] The target intention classification model is used to identify the intention of the target text information according to the context information and the position information to obtain the intention prediction result of the target text information.

[0034] In some feasible implementations, before executing the business operation to obtain the business processing result of the target text information, the method further includes:

[0035] Determining whether the user can call the business process according to the position information;

[0036] If the user cannot call the business process, the business operation will not be executed and a prompt message will be issued.

[0037] In some feasible implementations, the performing the business operation to obtain the business processing result of the target text information includes:

[0038] Generate a business operation request according to the business operation, and send the business operation request to the multi-task learning model;

[0039] The business operation requests are processed in parallel through the multi-task learning model to obtain the business processing result.

[0040] In some feasible implementations, the multi-task learning model further includes an interrupt interface, a cancel interface, a switch interface, and a resubmit interface, and the method further includes:

[0041] Performing an interrupt operation on the business operation request through the interrupt interface, and suspending execution of the business operation corresponding to the business operation request;

[0042] Perform a cancel operation on the business operation request through the cancel interface to cancel the business operation corresponding to the business operation request;

[0043] Performing a switching operation on the business operation request through the switching interface, switching and executing the business operation corresponding to the business operation request;

[0044] The suspended or cancelled business operation is resubmitted through the re-interface to generate a business operation request for the suspended or cancelled business operation.

[0045] In some feasible implementations, the method further includes:

[0046] Analyze the target text information through a sentiment analysis model to generate a sentiment score for the target text information;

[0047] Obtaining historical conversation information of the user;

[0048] Generate comfort information based on the emotion score and the historical conversation information.

[0049] The embodiment of the present invention further discloses an intelligent dialogue system based on a business scenario, the system comprising:

[0050] An information acquisition module is used to acquire target text information input by a user;

[0051] A model calling module, used for selecting a target processing model corresponding to the target text information according to preset model calling information;

[0052] An intention recognition module, used to perform intention recognition on the target text information according to the target processing model, and generate an intention prediction result of the target text information;

[0053] A business process determination module, used to determine the business process and business operation corresponding to the target text information according to the intention prediction result;

[0054] The business processing module is used to execute the business operation and obtain the business processing result of the target text information.

[0055] In some feasible implementations, the information acquisition module includes:

[0056] The acquisition submodule is used to obtain the initial text information input by the user;

[0057] A recognition submodule, used to recognize the initial text information through a natural language understanding model to obtain key information of the initial text;

[0058] A guidance submodule, used to generate guidance information according to the key information;

[0059] A supplementary submodule, used to obtain the supplementary information input by the user according to the guidance information;

[0060] The improvement submodule is used to generate target text information corresponding to the initial text information according to the key information and the supplementary information.

[0061] In some feasible implementations, the information acquisition module includes:

[0062] An acquisition submodule, used to acquire first text information input by a user;

[0063] A semantic analysis submodule, used to perform semantic analysis on the first text information and generate a semantic analysis result;

[0064] A judgment submodule, used for judging whether the first text information has a semantic error according to the semantic analysis result;

[0065] The repair submodule generates feedback information for the first text information if there is a semantic error; obtains the second text information input by the user according to the feedback information, and uses the second text information as the initial text information.

[0066] In some feasible implementations, the target processing model includes a target context perception model and a target intent classification model, and the model calling module includes:

[0067] An analysis submodule, used for analyzing the target text information according to the model calling information to obtain the processing requirements of the target text information;

[0068] The model selection submodule is used to select a context-aware large model or a context-aware small model from a preset model pool as a target context-aware model according to the processing requirements, and to select an intent classification large model or an intent classification small model as a target intent classification model.

[0069] In some feasible implementations, the target processing model includes a target context perception model and a target intent classification model, and the intent recognition module includes:

[0070] A word segmentation submodule, used to perform word segmentation processing on the target text information to obtain a word vector corresponding to the target text information;

[0071] A context information extraction submodule, used to process the word vector through the target context perception model to obtain the context information of the target text information;

[0072] The intent classification submodule is used to perform intent recognition on the target text information according to the context information through the target intent classification model to obtain the intent prediction result of the target text information.

[0073] In some feasible implementations, the system further includes an automatic retrieval module, which is used to automatically search the user information database to obtain the user's position information;

[0074] The intention classification submodule is also used to perform intent recognition on the target text information according to the context information and the position information through the target intention classification model to obtain the intention prediction result of the target text information.

[0075] In some feasible implementations, the system further includes a permission determination module, and the permission determination module is used to:

[0076] Determining whether the user can call the business process according to the position information;

[0077] If the user cannot call the business process, the business operation will not be executed and a prompt message will be issued.

[0078] In some feasible implementations, the business processing module includes:

[0079] A request generation submodule, used to generate a business operation request according to the business operation, and send the business operation request to the multi-task learning model;

[0080] The operation execution submodule is used to process the business operation requests in parallel through the multi-task learning model to obtain the business processing results.

[0081] In some feasible implementations, the multi-task learning model further includes an interrupt interface, a cancel interface, a switch interface, and a resubmit interface, and the operation execution submodule is further used to:

[0082] Performing an interrupt operation on the business operation request through the interrupt interface, and suspending execution of the business operation corresponding to the business operation request;

[0083] Perform a cancel operation on the business operation request through the cancel interface to cancel the business operation corresponding to the business operation request;

[0084] Performing a switching operation on the business operation request through the switching interface, switching and executing the business operation corresponding to the business operation request;

[0085] The suspended or cancelled business operation is resubmitted through the re-interface to generate a business operation request for the suspended or cancelled business operation.

[0086] The embodiment of the present invention further discloses an electronic device, comprising: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0087] The memory is used to store computer programs;

[0088] The processor is used to implement the method described in the embodiment of the present invention when executing the program stored in the memory.

[0089] The embodiment of the present invention further discloses a computer-readable storage medium on which instructions are stored. When executed by one or more processors, the processors are enabled to execute the method described in the embodiment of the present invention.

[0090] The embodiments of the present invention include the following advantages:

[0091] In an embodiment of the present invention, target text information input by a user is obtained; a target processing model corresponding to the target text information is selected according to preset model calling information; intent recognition is performed on the target text information according to the target processing model to generate an intent prediction result of the target text information; the business process and business operation corresponding to the target text information are determined according to the intent prediction result; the business operation is executed to obtain the business processing result of the target text information, and different processing models are flexibly called through intelligent dialogue to quickly and accurately identify the user's intent, and then the user's intent is mapped to the corresponding business process, and the corresponding business operation is triggered, thereby realizing flexible process orchestration services, meeting the personalized needs of users, and improving the response speed of services and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 is a flowchart of a business scenario-based intelligent dialogue method provided in an embodiment of the present invention;

[0093] Figure 2 It is a schematic diagram of a scenario of calling different business processes based on user input provided in an embodiment of the present invention;

[0094] Figure 3 It is a solution architecture diagram for calling different business processes based on user input provided in an embodiment of the present invention;

[0095] Figure 4 is a structural block diagram of an intelligent dialogue system based on a business scenario provided in an embodiment of the present invention;

[0096] Figure 5 is a block diagram of an electronic device provided in an embodiment of the present invention;

[0097] Figure 6 is a schematic diagram of a computer-readable storage medium provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0098] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0099] As an example, existing technologies usually use big model technology to fine-tune models to adapt to different business needs. However, although adjustments based on big models can improve accuracy, big model technology requires a large amount of knowledge base and computing resource support, resulting in high resource consumption and uncertainty. In addition, the model calling technology is not flexible enough, and it is difficult to dynamically select and call the most appropriate model according to specific business scenarios, which affects the adaptability and scalability of the system.

[0100] In this regard, in the present invention, target text information input by the user is obtained; a target processing model corresponding to the target text information is selected according to preset model calling information; intent recognition is performed on the target text information according to the target processing model to generate an intent prediction result of the target text information; the business process and business operation corresponding to the target text information are determined according to the intent prediction result; the business operation is executed to obtain the business processing result of the target text information, and different processing models are flexibly called through intelligent dialogue to quickly and accurately identify the user's intent, and then the user's intent is mapped to the corresponding business process, and the corresponding business operation is triggered, thereby realizing flexible process orchestration services, meeting the personalized needs of users, and improving the response speed of services and user experience.

[0101] Reference Figure 1 , shows a flowchart of a business scenario-based intelligent dialogue method provided in an embodiment of the present invention, which may specifically include the following steps:

[0102] Step 101: Obtain target text information input by a user;

[0103] In the embodiment of the present invention, the system obtains the target text information input by the user in real time as the basis for intelligent dialogue, and the target text information represents the user's needs and intentions. Specifically, the user communicates by inputting text information, such as questions, requests, and instructions, and then the system obtains the text information input by the user in real time as the basic data for subsequent intention recognition and business process calls, realizing an instant interactive experience, thereby realizing automated, efficient and personalized services.

[0104] In some embodiments, the method of obtaining target text information input by a user includes: obtaining initial text information input by a user; identifying the initial text information through a natural language understanding model to obtain key information of the initial text; generating guidance information based on the key information; obtaining supplementary information input by the user based on the guidance information; and generating target text information corresponding to the initial text information based on the key information and the supplementary information.

[0105] In an embodiment of the present invention, after the user inputs the initial text information, the natural language understanding model will be called to identify the text information input by the user and extract the key information therein. The key information includes important information such as the user's intention and entity, and then judge whether the necessary information is lacking based on the key information. If it is lacking, guidance information is generated. The guidance information is prompt information or questions, which are used to guide the user to further explain the initial text information or supplement detailed information, complete the input initial text information, and obtain the target text information that fully expresses the user's intention. The embodiment of the present invention generates guidance information to guide the user to provide supplementary information, ensure the integrity and accuracy of the target text information, and obtain complete information through multiple rounds of interaction with the user, so as to provide more complete data for subsequent intent recognition and improve the accuracy of intent recognition.

[0106] As an example, in the specific implementation process, the key information in the text information input by the user is first identified through the natural language understanding model, and then the dialogue state tracker used to record and update the dialogue state is called. The dialogue state tracker dynamically fills the key information into the predefined slots, and then generates guidance information based on the key information in the slots. The guidance information guides the user to provide necessary information to complete the slots, so that the target text information finally obtained can fully represent the user's needs and intentions. For example, the user inputs "Please book a room at Hotel A from 12.1 to 12.3", the key information is identified and filled into the corresponding slots, and the intention is obtained: book a hotel, location: Hotel A, check-in date: 12.1, check-out date: 12.3, and the key information of the room type slot is missing, and then the guidance information "What room type do you want to book?" is generated to guide the user to output supplementary information "Book a double room", and the target text information that fully expresses the user's needs and intentions is generated based on the key information and supplementary information.

[0107] In some embodiments, the obtaining of initial text information input by a user includes: obtaining first text information input by a user; performing semantic analysis on the first text information to generate a semantic analysis result; judging whether the first text information has a semantic error based on the semantic analysis result; if a semantic error exists, generating feedback information for the first text information; and obtaining second text information input by the user based on the feedback information, and using the second text information as the initial text information.

[0108] In the embodiment of the present invention, the global dialogue will also be monitored. After the user inputs text information, the input text will be semantically analyzed to determine whether there is an error in the user input. If there is an error, feedback information will be generated to guide the user to restate and re-enter accurate text information. For example, if the user inputs "I want to buy a high-speed rail ticket from Xiamen to Guangzhou on November 1st" and there is a grammatical error, feedback information "Do you want to order a high-speed rail ticket from Xiamen to Guangzhou on November 1st, or a plane ticket" will be issued to guide the user to output the accurate text information "I want to buy a plane ticket from Xiamen to Guangzhou on November 1st".

[0109] In some embodiments, the method further includes: analyzing the target text information through a sentiment analysis model to generate a sentiment score for the target text information; obtaining historical conversation information of the user; and generating comforting information based on the sentiment score and the historical conversation information.

[0110] In the embodiment of the present invention, the user's emotional understanding and comfort will also be provided. Specifically, after receiving the text information input by the user, the sentiment analysis model will be called to analyze the text information to obtain the sentiment score of the user's input text information. The sentiment score represents the user's emotional tendency, and the comfort information is further generated in combination with the user's historical conversation data. By analyzing the historical conversation information, the user's emotional state can be better understood, and more appropriate comfort information can be generated, so as to provide personalized emotional comfort and guidance to the user through the comfort information, thereby improving the user experience and satisfaction.

[0111] As an example, the text information input by the user is input into the sentiment analysis model, and the sentiment analysis model is calculated according to formula (1) to obtain the sentiment score of the user input.

[0112]

[0113] Among them, S(u) represents the sentiment score of the user input text information u, w i represents the weight of the i-th sentiment feature, s i represents the score of the i-th sentiment feature.

[0114] Step 102: selecting a target processing model corresponding to the target text information according to preset model calling information;

[0115] In an embodiment of the present invention, after obtaining the target text information input by the user, the target processing model corresponding to the target text information will be selected according to the model calling information. The model calling information is a preset model calling strategy, which is used to dynamically select the most suitable processing model according to the user input. Specifically, a model pool is pre-set to store large models and small models corresponding to various processing models. After receiving the user input, the large model or small model is dynamically called for processing according to the processing requirements of the target text information, such as task complexity and computing resource requirements. The model calling strategy can also be further adjusted in combination with the system load. For example, when the system load is too high, the small model processing task is called first to reduce resource consumption, so as to realize the dynamic calling of the large model or small model for processing according to the processing requirements of the target text information, and optimize the allocation and management of computing resources.

[0116] In some embodiments, the target processing model includes a target context-aware model and a target intent classification model, and the target processing model corresponding to the target text information is selected according to preset model calling information, including: analyzing the target text information according to the model calling information to obtain the processing requirements of the target text information; selecting a context-aware large model or a context-aware small model from a preset model pool as the target context-aware model according to the processing requirements, and selecting an intent classification large model or an intent classification small model as the target intent classification model.

[0117] In an embodiment of the present invention, identifying the intent of the target text information mainly involves a context-aware model and an intent classification model. The context-aware model is used to understand the context information input by the user, and the intent classification model is used to identify the intent of the user input. The target text information is analyzed according to the model call information to identify the processing requirements. The processing requirements may include task complexity, context requirements, intent recognition requirements, etc., and then according to the processing requirements, a suitable context-aware model is selected from a preset model pool as the target context-aware model, and a suitable intent classification model is selected as the target intent classification model. The model pool includes at least a context-aware large model, a context-aware small model, an intent classification large model and an intent classification small model. The context-aware large model is used to process complex context information, and the context-aware small model is used to process simple context information. The intent classification large model is used to process complex intent recognition, and the intent classification small model is used to process simple intent recognition, so as to realize dynamic selection of suitable models according to processing requirements, improve processing efficiency, and reduce unnecessary resource consumption.

[0118] As an example, the model call information can be the task complexity of the user input text information, and the task complexity of the target text information is calculated according to formula (2):

[0119] C = a·L + β·K (2)

[0120] Where C represents the task complexity, L represents the length of the input text information, K represents the knowledge requirement, and a and β are weight parameters;

[0121] Furthermore, the task complexity is compared with a preset threshold T. When C < T, a small model is called to process the task; when C ≥ T, a large model is called to process the task. Among them, the preset threshold can be determined through big data training and analysis.

[0122] As another example, the model call information can also include the system load. The resource usage of the system is monitored in real time, such as the CPU (Central Processing Unit), memory, and bandwidth, and the system load is further calculated according to formula (3):

[0123] L s = γ·CPU + δ·MEM + ε·BW (3)

[0124] Where, L s represents the system load, CPU, MEM, and BW respectively represent the usage of CPU, memory, and bandwidth, and γ, δ, and ε are weight parameters.

[0125] Step 103, perform intent recognition on the target text information according to the target processing model, and generate an intent prediction result of the target text information;

[0126] In the embodiment of the present invention, the target text information is subjected to intent recognition through the target processing model to generate an intent prediction result for characterizing the user's needs and goals. Further, according to the intent prediction result, the corresponding business process and specific business operations are determined, and accurate services are provided according to the user's intent to meet the user's personalized needs.

[0127] In the specific implementation process, first, the text information input by the user is obtained and text cleaning is performed to remove irrelevant characters. Then, the text information is subjected to word segmentation processing to obtain words, and the words are converted into vector representations according to the word embedding model to obtain word vectors to capture the semantic relationships between words. The specific word embedding model is processed according to formula (4):

[0128]

[0129] Where, P(w i |context(w i)) means predicting word w given the context i probability.

[0130] Secondly, the context vector representation of the input text information is generated according to the context-aware model. Specifically, the BERT model is used to capture the context information according to the following formula based on the bidirectional Transformer architecture (Formula 5) and further combined with the self-attention mechanism (Formula 6) and the dynamic context weight adjustment mechanism (Formula 7):

[0131]

[0132] W t =a t ·W t-1 +(1-a t )·C t (7) Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key vector, W t is the context weight at the current moment, W t-1 is the context weight of the previous moment, C t is the context information at the current moment, a t is the dynamic adjustment coefficient.

[0133] Furthermore, the LSTM (Long Short-Term Memory) model is trained using the following formula to obtain the intent classification model:

[0134] i t =σ(W i ·[h t-1 ,x t ]+b i ) (8)

[0135] f t =σ(W f ·[h t-1 ,x t ]+b f ) (9)

[0136] o t =σ(W o ·[h t-1 ,x t ]+b o ) (10)

[0137] ζ t =tanh(W c ·[h t-1 ,x t ]+b c) (11)

[0138] C t =f t *C t-1 +i t *ζ t (12)

[0139] n t =o t *tanh(ι t ) (13)

[0140]

[0141] Among them, i t is the input gate, f t It is the forget gate, t is the output gate, t is the tanh function used by the input gate, C t is the memory unit, n t is the tanh function used by the output gate, L is the cross entropy loss function, y i is the true label, is the predicted probability;

[0142] Finally, the trained intent recognition model is used to identify the user input, generate prediction results for one or more intents, further calculate the confidence of each intent, and select the intent with the highest confidence as the final intent prediction result.

[0143] In some embodiments, global conversation monitoring involves not only monitoring the text information input by the user but also the intention prediction results output by the system. After the intention prediction results are output, the intention prediction results will be error detected according to predefined rules. When an error is detected in the intention prediction result, the user will be redirected to express himself by generating prompt information or questions so as to re-determine the user's intention.

[0144] As an example, the probability of false detection is calculated as follows:

[0145]

[0146] Among them, P(e|u) represents the probability of error e occurring when the user inputs u, P(u|e) represents the probability of the user inputting u when error e occurs, P(e) represents the prior probability of error e occurring, and P(u) represents the prior probability of the user inputting u.

[0147] In some embodiments, the method also includes: automatically searching a user information database to obtain the user's job information; identifying the intent of the target text information according to the context information through the target intention classification model to obtain the intention prediction result of the target text information, including: identifying the intent of the target text information according to the context information and the job information through the target intention classification model to obtain the intention prediction result of the target text information.

[0148] In an embodiment of the present invention, the user information database used to store user information will also be automatically retrieved and the user's position information will be extracted. The position information can be used to provide additional contextual information, such as the user's position, department, and responsibilities to better help understand the user's needs and intentions. The user's input text information is further identified by combining the contextual information and position information to obtain more accurate intention prediction results, so that the target intention classification model can understand user input from multiple dimensions and improve the accuracy of intention recognition. Optionally, RAG (Retrieval-Augmented Generation) can be used to retrieve relevant documents or information libraries, and RAG also supports the configuration of its own or external knowledge base, which can integrate multi-source information and provide more comprehensive and accurate data.

[0149] Step 104: determining the business process and business operation corresponding to the target text information according to the intention prediction result;

[0150] In the embodiment of the present invention, the business process and specific business operation corresponding to the user input are further determined based on the intention prediction result. The business process refers to a series of steps to complete a specific task, and the business operation refers to a specific execution action. For example, the user asks "What is the status of my order?" It can be mapped to a business process that includes order query, status update and result return, and the corresponding business operations include "query order database", "get order status", and "generate query results". The embodiment of the present invention realizes flexible process orchestration services for users by mapping the identified user intentions to predefined business processes, providing customized responses and services, and thus improving user experience.

[0151] Step 105: Execute the business operation to obtain the business processing result of the target text information.

[0152] In an embodiment of the present invention, after determining the business process corresponding to the user input, the corresponding business operation will be executed, and the business processing result corresponding to the user input will be obtained and returned to the user, thereby realizing the ability to dynamically call different business operations based on the text information and processing requirements input by the user, so as to flexibly respond to different business needs and scenarios, provide customized responses and services, and enhance user experience.

[0153] In some embodiments, position information can also be used as permission information for calling business processes. After confirming that the user has entered the corresponding business process, it will also be determined based on the user's position information whether the user has the authority to call the business process. If so, the business operation corresponding to the business process will be executed. If not, a prompt message will be returned to inform the user that he or she does not have the authority to call the business process and execute the corresponding business operation. Through permission control, sensitive business processes and data can be protected to prevent unauthorized access and operations.

[0154] In some embodiments, executing the business operation and obtaining the business processing result of the target text information includes: generating a business operation request based on the business operation, and sending the business operation request to a multi-task learning model; and performing parallel processing on the business operation requests through the multi-task learning model to obtain the business processing result.

[0155] In an embodiment of the present invention, when determining the business process and business operation corresponding to the user input, a business operation request corresponding to the business operation will be generated and sent to the multi-task learning model, and then the business operation requests will be processed in parallel through the multi-task learning model to improve processing efficiency. The multi-task learning model can also dynamically adjust the priority of tasks, i.e., business operation requests, to ensure that key tasks are processed first.

[0156] As an example, the total loss function of the multi-task learning model is calculated as follows:

[0157]

[0158] Among them, L total represents the total loss function, λ j represents the weight of the jth task, L j represents the loss function of the jth task.

[0159] Furthermore, the multi-task learning model also provides multiple interfaces such as interrupt interface, cancel interface, switch interface and resubmit interface to monitor the execution status of business operations in real time. It can further perform interrupt operation on business operation request through the interrupt interface to suspend the current business operation; perform cancel operation on business operation request through the cancel interface to cancel the current business operation; perform switch operation on business operation request through the switch interface to switch the current business operation; resubmit the suspended or canceled business operation through the resubmit interface to generate a business operation request for the suspended or canceled business operation, so as to actively intervene and manage business operations and flexibly control the execution of business operations through the interrupt, cancel, switch and resubmit interfaces.

[0160] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following examples are used for exemplary description:

[0161] Reference Figure 2 , showing a scenario schematic diagram of calling different business processes based on user input provided in an embodiment of the present invention. First, the user inputs a question, and the user's job data is automatically extracted to determine the user's job information, which is used as auxiliary information for intent recognition and permission information for business process permission judgment. The user input is intended to be recognized, and the business process corresponding to the user input is determined. Before calling the corresponding business process, a permission judgment is made based on the user's job information to determine whether the user has the permission to call the business process. If so, the business process is called and the corresponding operation is performed, so that users can automatically call and execute business processes through intelligent dialogue, provide customized business processes, and meet the needs of different users.

[0162] Reference Figure 3 , showing a scheme architecture diagram for calling different business processes based on user input provided in an embodiment of the present invention. In a business scenario, the user inputs questions used to characterize the user's needs and intentions. By automatically retrieving and extracting knowledge content files, the user's job information is obtained as auxiliary information for intent recognition, and the job information is segmented and vectorized for subsequent model processing. The user input questions are further identified through the model mixing mode of "small model + large model", the intention prediction results of the user input are determined, and the corresponding business processes are called according to the intention recognition results to realize flexible process orchestration services. In this process, the global dialogue monitoring can also be used to perform error detection on the user input text information and the intention prediction results output by the system to realize dialogue repair. When an error in the intention prediction result is detected, the user is redirected to express, and then the user's intention is re-determined, and the third-party system or service is quickly called through the plug-in center to realize the call of the business process, and the business operations can also be processed in parallel through the multi-task learning model.

[0163] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0164] It should be noted that the embodiments of the present invention include but are not limited to the above examples. It is understandable that those skilled in the art can also make settings according to actual needs under the guidance of the ideas of the embodiments of the present invention, and the present invention is not limited to this.

[0165] In an embodiment of the present invention, target text information input by a user is obtained; a target processing model corresponding to the target text information is selected according to preset model calling information; intent recognition is performed on the target text information according to the target processing model to generate an intent prediction result of the target text information; the business process and business operation corresponding to the target text information are determined according to the intent prediction result; the business operation is executed to obtain the business processing result of the target text information, and different processing models are flexibly called through intelligent dialogue to quickly and accurately identify the user's intent, and then the user's intent is mapped to the corresponding business process, and the corresponding business operation is triggered, thereby realizing flexible process orchestration services, meeting the personalized needs of users, and improving the response speed of services and the user experience.

[0166] Reference Figure 4 , shows a structural block diagram of an intelligent dialogue system based on a business scenario provided in an embodiment of the present invention, and the intelligent dialogue system based on a business scenario includes:

[0167] An information acquisition module is used to acquire target text information input by a user;

[0168] A model calling module, used for selecting a target processing model corresponding to the target text information according to preset model calling information;

[0169] An intention recognition module, used to perform intention recognition on the target text information according to the target processing model, and generate an intention prediction result of the target text information;

[0170] A business process determination module, used to determine the business process and business operation corresponding to the target text information according to the intention prediction result;

[0171] The business processing module is used to execute the business operation and obtain the business processing result of the target text information.

[0172] In some feasible implementations, the information acquisition module includes:

[0173] The acquisition submodule is used to obtain the initial text information input by the user;

[0174] A recognition submodule, used to recognize the initial text information through a natural language understanding model to obtain key information of the initial text;

[0175] A guidance submodule, used to generate guidance information according to the key information;

[0176] A supplementary submodule, used to obtain the supplementary information input by the user according to the guidance information;

[0177] The improvement submodule is used to generate target text information corresponding to the initial text information according to the key information and the supplementary information.

[0178] In some feasible implementations, the information acquisition module includes:

[0179] An acquisition submodule, used to acquire first text information input by a user;

[0180] A semantic analysis submodule, used to perform semantic analysis on the first text information and generate a semantic analysis result;

[0181] A judgment submodule, used for judging whether the first text information has a semantic error according to the semantic analysis result;

[0182] The repair submodule generates feedback information for the first text information if there is a semantic error; obtains the second text information input by the user according to the feedback information, and uses the second text information as the initial text information.

[0183] In some feasible implementations, the target processing model includes a target context perception model and a target intent classification model, and the model calling module includes:

[0184] An analysis submodule, used for analyzing the target text information according to the model calling information to obtain the processing requirements of the target text information;

[0185] The model selection submodule is used to select a context-aware large model or a context-aware small model from a preset model pool as a target context-aware model according to the processing requirements, and to select an intent classification large model or an intent classification small model as a target intent classification model.

[0186] In some feasible implementations, the target processing model includes a target context perception model and a target intent classification model, and the intent recognition module includes:

[0187] A word segmentation submodule, used to perform word segmentation processing on the target text information to obtain a word vector corresponding to the target text information;

[0188] A context information extraction submodule, used to process the word vector through the target context perception model to obtain the context information of the target text information;

[0189] The intent classification submodule is used to perform intent recognition on the target text information according to the context information through the target intent classification model to obtain the intent prediction result of the target text information.

[0190] In some feasible implementations, the system further includes an automatic retrieval module, which is used to automatically search the user information database to obtain the user's position information;

[0191] The intention classification submodule is also used to perform intent recognition on the target text information according to the context information and the position information through the target intention classification model to obtain the intention prediction result of the target text information.

[0192] In some feasible implementations, the system further includes a permission determination module, and the permission determination module is used to:

[0193] Determining whether the user can call the business process according to the position information;

[0194] If the user cannot call the business process, the business operation will not be executed and a prompt message will be issued.

[0195] In some feasible implementations, the business processing module includes:

[0196] A request generation submodule, used to generate a business operation request according to the business operation, and send the business operation request to the multi-task learning model;

[0197] The operation execution submodule is used to process the business operation requests in parallel through the multi-task learning model to obtain the business processing results.

[0198] In some feasible implementations, the multi-task learning model further includes an interrupt interface, a cancel interface, a switch interface, and a resubmit interface, and the operation execution submodule is further used to:

[0199] Performing an interrupt operation on the business operation request through the interrupt interface, and suspending execution of the business operation corresponding to the business operation request;

[0200] Perform a cancel operation on the business operation request through the cancel interface to cancel the business operation corresponding to the business operation request;

[0201] Performing a switching operation on the business operation request through the switching interface, switching and executing the business operation corresponding to the business operation request;

[0202] The suspended or cancelled business operation is resubmitted through the re-interface to generate a business operation request for the suspended or cancelled business operation.

[0203] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the system embodiment.

[0204] In addition, an embodiment of the present invention further provides an electronic device, such as Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503 and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0205] Memory 503, used for storing computer programs;

[0206] The processor 501 is used to execute the program stored in the memory 503 to implement the various processes of the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be described here.

[0207] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0208] The communication interface is used for communication between the above terminal and other devices.

[0209] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0210] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0211] like Figure 6As shown, in another embodiment provided by the present invention, a computer-readable storage medium 601 is also provided, in which instructions are stored. When executed by one or more processors, the processors execute the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, they are not repeated here.

[0212] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0213] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0214] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

[0215] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0216] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0217] In the embodiments provided by the present invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0218] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0219] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0220] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical disks.

[0221] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. An intelligent dialogue method based on business scenarios, characterized in that: The method comprises: Get the target text information entered by the user; Selecting a target processing model corresponding to the target text information according to preset model calling information; Performing intent recognition on the target text information according to the target processing model to generate an intent prediction result of the target text information; Determine the business process and business operation corresponding to the target text information according to the intention prediction result; Execute the business operation to obtain the business processing result of the target text information.

2. The method according to claim 1, characterized in that: The step of obtaining target text information input by the user includes: Get the initial text information entered by the user; Recognize the initial text information through a natural language understanding model to obtain key information of the initial text; generating guidance information according to the key information; Obtaining the supplementary information input by the user according to the guidance information; According to the key information and the supplementary information, target text information corresponding to the initial text information is generated.

3. The method according to claim 2, characterized in that: The obtaining of initial text information input by the user includes: Obtaining first text information input by a user; Performing semantic analysis on the first text information to generate a semantic analysis result; Determining whether the first text information has a semantic error according to the semantic analysis result; If there is a semantic error, generating feedback information for the first text information; According to the feedback information, second text information input by the user is obtained, and the second text information is used as the initial text information.

4. The method according to claim 1, characterized in that: The target processing model includes a target context perception model and a target intent classification model, and the target processing model corresponding to the target text information is selected according to the preset model calling information, including: Analyze the target text information according to the model calling information to obtain the processing requirements of the target text information; According to the processing requirements, a context-aware large model or a context-aware small model is selected from a preset model pool as a target context-aware model, and a large intent classification model or a small intent classification model is selected as a target intent classification model.

5. The method according to claim 1, characterized in that: The target processing model includes a target context perception model and a target intent classification model, and the performing intent recognition on the target text information according to the target processing model to generate an intent prediction result of the text information includes: Performing word segmentation processing on the target text information to obtain a word vector corresponding to the target text information; Processing the word vector by the target context-aware model to obtain context information of the target text information; The target intention classification model is used to perform intention recognition on the target text information according to the context information to obtain an intention prediction result of the target text information.

6. The method according to claim 5, characterized in that: The method further comprises: Automatically search the user information database to obtain the user's job information; The step of performing intent recognition on the target text information according to the context information by using the target intent classification model to obtain an intent prediction result of the target text information includes: The target intention classification model is used to identify the intention of the target text information according to the context information and the position information to obtain the intention prediction result of the target text information.

7. The method according to claim 6, characterized in that: Before executing the business operation and obtaining the business processing result of the target text information, the method further includes: Determining whether the user can call the business process according to the position information; If the user cannot call the business process, the business operation will not be executed and a prompt message will be issued.

8. The method according to claim 1, characterized in that: The performing of the business operation to obtain the business processing result of the target text information includes: Generate a business operation request according to the business operation, and send the business operation request to the multi-task learning model; The business operation requests are processed in parallel through the multi-task learning model to obtain the business processing result.

9. The method according to claim 8, characterized in that: The multi-task learning model further includes an interrupt interface, a cancel interface, a switch interface, and a resubmit interface, and the method further includes: Performing an interrupt operation on the business operation request through the interrupt interface, and suspending execution of the business operation corresponding to the business operation request; Perform a cancel operation on the business operation request through the cancel interface to cancel the business operation corresponding to the business operation request; Performing a switching operation on the business operation request through the switching interface, switching and executing the business operation corresponding to the business operation request; The suspended or cancelled business operation is resubmitted through the re-interface to generate a business operation request for the suspended or cancelled business operation.

10. The method according to claim 1, characterized in that: The method further comprises: Analyze the target text information through a sentiment analysis model to generate a sentiment score for the target text information; Obtaining historical conversation information of the user; Generate comfort information based on the emotion score and the historical conversation information.

11. An intelligent dialogue system based on business scenarios, characterized in that: The system comprises: An information acquisition module is used to acquire target text information input by a user; A model calling module, used for selecting a target processing model corresponding to the target text information according to preset model calling information; An intention recognition module, used to perform intention recognition on the target text information according to the target processing model, and generate an intention prediction result of the target text information; A business process determination module, used to determine the business process and business operation corresponding to the target text information according to the intention prediction result; The business processing module is used to execute the business operation and obtain the business processing result of the target text information.

12. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 10 when executing the program stored in the memory.

13. A computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Service providing method and device and service request processing method

    CN120633609A

  • Knowledge extraction method and system, electronic equipment and storage medium

    CN120822593A