Dialogue intention information determination method and device, equipment and storage medium

By assessing the complexity of user dialogue information and selecting appropriate inference modules for intent analysis, the problem of imbalance between recognition accuracy and resource utilization in existing technologies is solved, achieving high-precision and low-cost user intent recognition.

CN122088698APending Publication Date: 2026-05-26YI REN HENG YE TECH DEV (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YI REN HENG YE TECH DEV (BEIJING) CO LTD
Filing Date
2026-02-12
Publication Date
2026-05-26

Smart Images

  • Figure CN122088698A_ABST
    Figure CN122088698A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue intention information determination method and device, equipment and a storage medium. The method comprises the following steps: acquiring target speaking information of a target user in a communication dialogue; and performing complexity analysis processing according to the target speaking information, and determining a comprehensive complexity value corresponding to the target speaking information. And screening and determining a target reasoning module for intention analysis from a pre-constructed reasoning module set according to a pre-constructed complexity level threshold value and the comprehensive complexity value. And performing intention recognition processing on the target speaking information according to the target reasoning module, and determining dialogue intention information of the target user. According to the technical scheme, different resources are called according to the complexity of the user dialogue information, so that stable recognition of the intention category and intention intensity of the user is achieved, and the computing power and the cost are controlled while the recognition effect is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dialogue semantic analysis technology, and in particular to a method, apparatus, device and storage medium for determining dialogue intent information. Background Technology

[0002] With the widespread application of intelligent customer service and chatbots in industries such as finance, telecommunications, and the internet, the system needs to accurately determine the user's current intent during real-time dialogue, such as complaints, inquiries, applications, rejections, and hesitations, in order to provide refined services and business decisions and provide decision-making support for subsequent business processes.

[0003] Currently, existing technologies typically analyze user dialogues to determine user intent through keyword matching, single neural networks, and large-scale language models. However, while rule-based keyword matching is simple to deploy, it struggles to adapt to the flexible and varied expressions of natural language, resulting in limited recall. Single neural network-based models, while improving generalization ability, are unstable in complex or contradictory contexts and incur significant computational costs. Solutions relying directly on large-scale language models, while offering enhanced understanding capabilities, suffer from high costs and response delays due to a lack of invocation control, posing challenges for practical implementation.

[0004] Therefore, existing technologies cannot achieve an effective balance between recognition accuracy and resource allocation based on the scenario. The clarity of user intent in the dialogue lacks a quantitative evaluation mechanism, which fails to meet the needs of practical applications. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and storage medium for determining dialogue intent information. By calling different resources based on the complexity of user dialogue information, it achieves stable identification of user intent categories and intent strength, while controlling computing power and costs while ensuring identification effectiveness.

[0006] According to one aspect of the present invention, a method for determining dialogue intent information is provided. The method includes: Obtain target speech information of a target user in a dialogue, wherein the target speech information includes current speech information, historical speech information, and dialogue parameter information; Complexity analysis is performed on the target speech information to determine the overall complexity value corresponding to the target speech information; Based on the pre-built complexity level threshold and the comprehensive complexity value, target reasoning modules for intent analysis are selected from the pre-built set of reasoning modules. The set of reasoning modules includes at least rule-layer reasoning modules, lightweight model reasoning modules, and large model-layer reasoning modules. The target reasoning module performs intent recognition processing on the target speech information to determine the dialogue intent information of the target user, wherein the dialogue intent information includes the target intent result and the intent confidence value.

[0007] According to another aspect of the present invention, a device for determining dialogue intent information is provided. The device includes: The speech information acquisition module is used to acquire the target speech information of the target user in the communication dialogue, wherein the target speech information includes current speech information, historical speech information and dialogue parameter information; The complexity value determination module is used to perform complexity analysis processing based on the target speech information to determine the comprehensive complexity value corresponding to the target speech information; The reasoning module filtering module is used to filter and determine the target reasoning module for intent analysis from a pre-built set of reasoning modules based on a pre-built complexity level threshold and the comprehensive complexity value. The set of reasoning modules includes at least a rule-layer reasoning module, a lightweight model reasoning module, and a large model-layer reasoning module. The intent information determination module is used to perform intent recognition processing on the target speech information based on the target reasoning module to determine the dialogue intent information of the target user, wherein the dialogue intent information includes the target intent result and the intent confidence value.

[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the dialogue intent information determination method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the dialogue intent information determination method according to any embodiment of the present invention.

[0010] The technical solution of this invention involves acquiring target speech information from a target user during a conversation. Complexity analysis is performed on the target speech information to determine its overall complexity value. Based on a pre-built complexity level threshold and the overall complexity value, a target inference module for intent analysis is selected from a pre-built set of inference modules. The target inference module performs intent recognition processing on the target speech information to determine the target user's dialogue intent information, which includes a target intent result and an intent confidence value. This invention, through complexity assessment and complexity level thresholds, calls the target inference module corresponding to the complexity level only when needed, significantly reducing unnecessary calls and resource usage. While ensuring recognition effectiveness, it controls computational power and costs, guarantees high-precision recognition and semantic understanding capabilities in key scenarios, and further improves overall performance and robustness.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a method for determining dialogue intent information according to an embodiment of the present invention; Figure 2 This is a flowchart of a method for determining dialogue intent information according to an embodiment of the present invention; Figure 3 This is a structural diagram of a dialogue intent information determination device provided according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the dialogue intent information determination method of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Figure 1 This is a flowchart illustrating a method for determining dialogue intent information according to an embodiment of the present invention. This embodiment is applicable to analyzing and determining a user's communication dialogue intent. The method can be executed by a dialogue intent information determining device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S101. Obtain the target user's speech information in the communication dialogue.

[0017] In this context, the target user can refer to the user whose intent is to be analyzed during communication. The target speech information can refer to the speech information of the target user during communication. For example, the target speech information includes current speech information, historical speech information, and dialogue parameter information.

[0018] The current speech information can refer to the target user's most recent speech text. Historical speech information can refer to the historical speech information between the target user and customer service during the current call. Preferably, the historical speech information in this invention can refer to a preset number of historical speech texts, such as the speech text information between the user and customer service 3-5 times before the current speech information. Dialogue parameter information can refer to the target user's basic information in the communication dialogue, including the current dialogue round, call duration (in seconds), number of user confirmations (such as the number of times expressions like "okay," "suffice," "um," etc. appear), number of negations (such as the number of times expressions like "no," "don't need," "never mind," etc. appear), number of proactive questions, and the silence duration before the most recent speech.

[0019] Specifically, the method of obtaining target user's speech information in the context of a dialogue can be chosen according to the actual situation, and this invention does not impose specific limitations or explanations on it.

[0020] S102. Perform complexity analysis processing based on the target speech information to determine the comprehensive complexity value corresponding to the target speech information.

[0021] Among them, overall complexity can refer to the complexity information of the target speech information.

[0022] Specifically, the present invention performs complexity analysis on the target speech information, evaluates the overall complexity of the current dialogue based on the speech features in the target speech information, determines whether there is significant ambiguity, and determines the target reasoning module required for subsequent reasoning based on the overall complexity.

[0023] S103. Based on the pre-built complexity level threshold and the comprehensive complexity value, select and determine the target reasoning module for intent analysis from the pre-built reasoning module set.

[0024] The reasoning module is primarily used to analyze and reason about the target's speech information, outputting corresponding intent prediction results and confidence levels. For example, the inference module set includes at least a rule-layer inference module, a lightweight model inference module, and a large model-layer inference module.

[0025] The rule-based reasoning module uses a configurable library of keywords or regular expressions to quickly match text, suitable for simple and clearly defined intent categories. Different priority rules can be set for this module; for example, complaints and insults can be set to high priority, ordinary rejections and applications to medium priority, and general inquiries to low priority. The module outputs the inferred intent category and its corresponding confidence level, and marks whether a strong rule has been hit. When the complexity is low and a high-priority rule is hit, the final result is output directly.

[0026] The lightweight model inference module can use FastText, linear models, or small deep learning models to perform multi-class predictions on target speech information, outputting the predicted intent category, confidence level, and selectable Top-K candidate intents. When the model confidence level is high and the text ambiguity is low, the results from the rule layer and the lightweight model layer can be weighted and fused and directly output; when the confidence level is insufficient or the ambiguity is high, further analysis by the large model layer is triggered.

[0027] The large model layer inference module is triggered when the text complexity is high, the lightweight model's confidence level is low, or the prediction results from the rule layer and the lightweight model layer are inconsistent or conflict significantly. Within the large model layer inference module, prompt templates can be constructed to input information such as the recent dialogue texts, rule layer results, and lightweight model prediction results into the large model, and a small number of typical examples can be used for prompting. The large model outputs structured results, including intent category, key evidence or reasons, and model confidence estimates.

[0028] Specifically, the complexity level threshold and the overall complexity value are matched to determine the complexity level of the target speech information, and then a matching target reasoning module is selected from the pre-built set of reasoning modules based on the complexity level.

[0029] For example, the step of selecting and determining target inference modules for intent analysis from a pre-built set of inference modules based on a pre-built complexity level threshold and the overall complexity value includes: If the overall complexity value belongs to the first complexity level threshold, the rule layer reasoning module is determined as the target reasoning module for analyzing the target speech information; If the overall complexity value belongs to the second complexity level threshold, both the rule layer reasoning module and the lightweight model reasoning module are determined as the target reasoning modules for analyzing the target speech information; If the overall complexity value belongs to the third complexity level threshold, the rule layer reasoning module, the lightweight model reasoning module, and the large model layer reasoning module are all identified as target reasoning modules for analyzing the target speech information.

[0030] In this invention, the complexity level thresholds include at least a first complexity level threshold, a second complexity level threshold, and a third complexity level threshold. The first complexity level threshold can be less than 0.3, the second complexity level threshold can be greater than or equal to 0.3 and less than 0.6, and the third complexity level threshold can be greater than or equal to 0.6.

[0031] Specifically, when the overall complexity value is at the first complexity level threshold, only the rule layer reasoning module is used as the target reasoning module; when the overall complexity value is at the second complexity level threshold, both the rule layer reasoning module and the lightweight model reasoning module are used as target reasoning modules; when the overall complexity value is at the third complexity level threshold, the rule layer reasoning module, the lightweight model reasoning module, and the large model layer reasoning module are used as target reasoning modules, and the large model fallback is triggered when necessary.

[0032] S104. Based on the target reasoning module, the target speech information is processed for intent recognition to determine the dialogue intent information of the target user.

[0033] The dialogue intent information includes the target intent result and the intent confidence value.

[0034] Specifically, based on the selected target inference module, the target speech information is processed for intent analysis. Based on the output of the target inference module, the target user's target intent result and intent confidence value can be determined.

[0035] For example, the step of performing intent recognition processing on the target speech information according to the target inference module to determine the target user's target intent result includes: for each target inference module, performing intent recognition processing on the target speech information according to the target inference module to obtain the primary intent result and primary confidence level output by the target inference module; weighting and fusing all primary intent results according to the intent prediction weights corresponding to each target inference module to obtain the target intent result; and weighting and fusing all primary confidence levels according to the confidence level weights corresponding to each target inference module to obtain the intent confidence level value.

[0036] The initial intent result can be the intent result predicted by each target inference module, and the initial confidence level can refer to the confidence level information predicted by each target inference module.

[0037] Specifically, the target speech information is processed by the target inference module to identify intent, thereby obtaining the initial intent result and initial confidence level output by the target inference module. When the target inference module is a rule-layer inference module, the initial intent result is directly determined as the target intent result, and the initial confidence level is directly determined as the intent confidence level value.

[0038] When there are multiple target inference modules, the initial intent results of each target inference module can be fused in a weighted manner according to the intent prediction weights corresponding to each target inference module, so as to obtain the final target intent result.

[0039] When there are multiple target inference modules, the initial confidence scores are weighted and fused according to the confidence scores of each target inference module to obtain the intent confidence score. For example: Let the confidence level of the rule layer be conf_rule; let the confidence level of the lightweight model layer be conf_model; and let the confidence level of the large model layer be conf_llm.

[0040] The final intent confidence value is obtained by weighting and summing the confidence coefficients α, β, and γ. ; The confidence weight can be dynamically adjusted according to the stage of the dialogue. For example, in the early rounds, the rule-based reasoning module and the lightweight model reasoning module can be given higher confidence weights, while in the later stages of the dialogue, the confidence weight of the large model reasoning module can be appropriately increased.

[0041] For example, the method further includes: performing intent result analysis on the dialogue intent information obtained within a preset round, and adjusting the filtering rules of the target inference module and the intent confidence value according to the intent analysis results.

[0042] In other words, in the technical solution of this invention, a lightweight consistency check can be performed only on the prediction results of the most recent few rounds of dialogue. Without introducing complex feature engineering, the filtering rules of the target inference module and the intent confidence value are adjusted based solely on the intent prediction results of the most recent few rounds and simple behavioral statistics. Specifically, this includes: when there are few dialogue rounds and insufficient information, appropriately lowering the current confidence value or increasing the probability of triggering the large model inference module; when the intent category of the recent prediction results is stable and the user behavior features support the intent, appropriately increasing the current confidence value; when the intent category of recent predictions changes frequently or there is a significant inconsistency between the previous and subsequent expressions, lowering the current confidence value and re-triggering the large model layer for review. Through the above mechanism, the stability and reliability of the overall prediction results can be improved without increasing the dependence on complex features.

[0043] The technical solution of this invention involves acquiring target speech information from a target user during a conversation. Complexity analysis is performed on the target speech information to determine its overall complexity value. Based on a pre-built complexity level threshold and the overall complexity value, a target inference module for intent analysis is selected from a pre-built set of inference modules. The target inference module performs intent recognition processing on the target speech information to determine the target user's dialogue intent information, which includes a target intent result and an intent confidence value. This invention, through complexity assessment and complexity level thresholds, calls the target inference module corresponding to the complexity level only when needed, significantly reducing unnecessary calls and resource usage. While ensuring recognition effectiveness, it controls computational power and costs, guarantees high-precision recognition and semantic understanding capabilities in key scenarios, and further improves overall performance and robustness.

[0044] Figure 2 This is a flowchart illustrating a method for determining dialogue intent information according to an embodiment of the present invention. Based on the above embodiments, this embodiment further refines the comprehensive complexity value for determining the target speech information. For example... Figure 2 As shown, the method includes: S201. Obtain the target user's speech information in the communication dialogue.

[0045] S202. Perform feature extraction processing on the target speech information to obtain speech feature information.

[0046] The speech feature information includes, but is not limited to, text normalization features, lexical diversity features, round normalization features, and behavior duration features.

[0047] Specifically, this invention can perform feature extraction processing on target speech information through a feature extraction module. It extracts multi-dimensional numerical features from the current text, recent context, and basic statistical information to obtain speech feature information including, but not limited to, text length and normalized features, lexical diversity features, turn number and normalized features, and behavior and duration features and their normalized values. All speech feature information is obtained from the original dialogue data through simple statistical methods, without relying on complex feature engineering, making it easy to implement and migrate across various business systems.

[0048] S203. Perform ambiguity detection processing on the target speech information to determine the speech ambiguity score of the target speech information.

[0049] Among them, ambiguity detection processing is used to measure whether the current input content is easy to judge or contains significant ambiguity.

[0050] Specifically, ambiguous words in the target speech information can be analyzed and detected, and a speech ambiguity score of the target speech information can be calculated based on the ambiguous words in the target speech information.

[0051] For example, the step of performing ambiguity detection processing on the target speech information and determining the speech ambiguity score of the target speech information includes: determining the hesitant expression information contained in the current speech information in the target speech information; determining the negative expression information contained in the historical speech information in the target speech information; and calculating the speech ambiguity score of the target speech information based on the hesitant expression information and the negative expression information.

[0052] Specifically, the system checks whether the current speech contains typical hesitant or uncertain expressions, such as "let me think about it," "maybe," "it seems," "uncertain," "I can't say for sure," and "we'll see." It also checks whether obvious negative expressions, such as "no," "not needed," "never mind," and "let's leave it at that for now," appear repeatedly in recent rounds of historical speech. Based on pre-assigned ambiguity scores and weighting coefficients for hesitant and negative expressions, the system calculates the ambiguity score of the target speech and normalizes it to the 0–1 range.

[0053] S204. Perform complexity analysis processing based on the speech feature information and the speech ambiguity score to determine the comprehensive complexity value corresponding to the target speech information.

[0054] The overall complexity value needs to take into account speech features such as text normalization features, lexical diversity features, and round normalization features, as well as speech ambiguity score features. The complexity analysis can be performed by averaging or weighted averaging the speech features and speech ambiguity scores to obtain the overall complexity value.

[0055] S205. Based on the pre-built complexity level threshold and the comprehensive complexity value, select and determine the target reasoning module for intent analysis from the pre-built reasoning module set.

[0056] S206. Based on the target reasoning module, the target speech information is processed for intent recognition to determine the dialogue intent information of the target user.

[0057] For example, the dialogue intent information further includes an intent intensity level; determining the intent intensity level of the target user includes: performing behavioral analysis processing based on the target speech information to determine the user behavior characteristics of the target user; performing behavioral intensity calculation based on the user behavior characteristics to determine the behavioral intensity score corresponding to the user behavior characteristics; and determining the intent intensity level of the target user based on the behavioral intensity score and a preset level threshold.

[0058] For example, in this invention, the dialogue intent information may also include an intent intensity level. Determining the intent intensity level of a target user is primarily based on an assessment of the target user's behavioral characteristics during the dialogue. A higher number of confirmations, questions, rounds, and durations in the user's behavioral characteristics generally indicates a higher level of attention the target user pays to the target behavior or product, and a stronger intent intensity; conversely, a higher number of denials generally indicates a lower intensity of the current target intent.

[0059] Specifically, behavioral analysis is performed on the target user's speech information to determine the user's behavioral characteristics. For example, user behavioral characteristics may include the number of confirmations and normalized values, the number of denials and normalized values, the number of proactive questions and normalized values, the number of dialogue turns and normalized values, and the call duration and normalized values. By weighted summing and normalizing the user behavioral characteristics, a behavioral intensity score within the 0–1 range can be obtained, and then the intent intensity level is classified according to a preset level threshold.

[0060] For example: when the behavior intensity score is ≥0.7, the intention intensity level is determined to be high intensity; when 0.4≤behavior intensity score<0.7, the intention intensity level is determined to be medium intensity; when the behavior intensity score<0.4, the intention intensity level is determined to be low intensity.

[0061] The technical solutions of this invention can be applied to various scenarios, such as single-round explicit complaint scenarios: 1. User inputs: "Your service is terrible, I want to complain." 2. The feature extraction module calculates basic features such as text length and rounds, while the complexity evaluation module determines that the current complexity is low; 3. The rule layer hits high-priority rules related to "complaints", indicating a high level of confidence at the rule layer. 4. The system can output the following without triggering subsequent models: the intent category is complaint, the intent strength is high, and the confidence level is relatively high.

[0062] For example, a scenario where the application intent gradually strengthens over multiple rounds: 1. In the initial stage, users mainly consult and raise product-related questions; 2. As the conversation progressed, users gradually began to explicitly ask questions such as, "Can you help me submit the application?" 3. As the number of dialogue rounds, call duration, number of confirmations, and number of questions gradually increase, the calculated intent intensity rises from medium to high. 4. Final system output: The intent category is application, the intent strength is high, and the confidence level is relatively high, which can provide a basis for subsequent business conversion.

[0063] For example, consider a scenario where there is hesitation regarding the target behavior: 1. User input: "I don't know whether I want it or not right now. I'll think about what you said." 2. The ambiguity detection module identified hesitant expressions such as "let me think about it" and "I don't know whether to do it or not," and scored highly on the ambiguity detection. 3. The text complexity assessment results are too high, while the confidence scores for both the rule layer and the lightweight model layer are too low; 4. The system triggers large model layer analysis and, combined with the outputs of the rule layer and lightweight model layer, gives an explanatory conclusion of "temporarily uncertain / low-intensity demand"; 5. Final output: The intent category is low demand or pending, the intent strength is low, and the overall confidence level is medium.

[0064] Figure 3 This is a schematic diagram of a dialogue intent information determination device provided in an embodiment of the present invention. Figure 3 As shown, the device includes: The speech information acquisition module 301 is used to acquire the target speech information of the target user in the communication dialogue, wherein the target speech information includes current speech information, historical speech information and dialogue parameter information; The complexity value determination module 302 is used to perform complexity analysis processing based on the target speech information to determine the comprehensive complexity value corresponding to the target speech information; The reasoning module filtering module 303 is used to filter and determine the target reasoning module for intent analysis from a pre-built set of reasoning modules based on a pre-built complexity level threshold and the comprehensive complexity value. The set of reasoning modules includes at least a rule-layer reasoning module, a lightweight model reasoning module, and a large model-layer reasoning module. The intent information determination module 304 is used to perform intent recognition processing on the target speech information based on the target reasoning module to determine the dialogue intent information of the target user, wherein the dialogue intent information includes the target intent result and the intent confidence value.

[0065] The technical solution of this invention involves acquiring target speech information from a target user during a conversation. Complexity analysis is performed on the target speech information to determine its overall complexity value. Based on a pre-built complexity level threshold and the overall complexity value, a target inference module for intent analysis is selected from a pre-built set of inference modules. The target inference module performs intent recognition processing on the target speech information to determine the target user's dialogue intent information, which includes a target intent result and an intent confidence value. This invention, through complexity assessment and complexity level thresholds, calls the target inference module corresponding to the complexity level only when needed, significantly reducing unnecessary calls and resource usage. While ensuring recognition effectiveness, it controls computational power and costs, guarantees high-precision recognition and semantic understanding capabilities in key scenarios, and further improves overall performance and robustness.

[0066] Optionally, the complexity value determination module 302 includes: The speech feature determination unit is used to perform feature extraction processing on the target speech information to obtain speech feature information, wherein the speech feature information includes, but is not limited to, text normalization features, lexical diversity features, round normalization features, and behavior duration features. The ambiguity score determination unit is used to perform ambiguity detection processing on the target speech information and determine the speech ambiguity score of the target speech information. The complexity value determination unit is used to perform complexity analysis processing based on the speech feature information and the speech ambiguity score to determine the comprehensive complexity value corresponding to the target speech information.

[0067] Optional, ambiguous fraction determination unit, specifically used for: Determine the hesitation expression information contained in the current speech information within the target speech information; Determine the negative expression information contained in the historical speech information within the target speech information; Based on the hesitant expression information and the negative expression information, calculate the speech ambiguity score of the target speech information.

[0068] Optionally, the complexity level thresholds include at least a first complexity level threshold, a second complexity level threshold, and a third complexity level threshold; The reasoning module filtering module 303 is specifically used for: If the overall complexity value belongs to the first complexity level threshold, the rule layer reasoning module is determined as the target reasoning module for analyzing the target speech information; If the overall complexity value belongs to the second complexity level threshold, both the rule layer reasoning module and the lightweight model reasoning module are determined as the target reasoning modules for analyzing the target speech information; If the overall complexity value belongs to the third complexity level threshold, the rule layer reasoning module, the lightweight model reasoning module, and the large model layer reasoning module are all identified as target reasoning modules for analyzing the target speech information.

[0069] Optionally, the intent information determination module 304 is specifically used for: For each target inference module, the target speech information is processed by the target inference module to obtain the initial intent result and initial confidence level output by the target inference module; Based on the intent prediction weights corresponding to each target inference module, all primary intent results are weighted and fused to obtain the target intent result. Based on the confidence weights corresponding to each target inference module, all primary confidence values ​​are weighted and fused to obtain the intent confidence value.

[0070] Optionally, the dialogue intent information further includes an intent intensity level; correspondingly, the intent information determination module 304 is also specifically used for: Based on the target's speech information, behavioral analysis and processing are performed to determine the target user's user behavior characteristics; Based on the user behavior characteristics, the behavior intensity is calculated to determine the behavior intensity score corresponding to the user behavior characteristics; The intent intensity level of the target user is determined based on the behavior intensity score and a preset level threshold.

[0071] Optionally, the device further includes a confidence calibration module. The confidence calibration module is used for: The intent information obtained within a preset number of rounds is analyzed for intent results, and the filtering rules of the target inference module and the intent confidence value are adjusted based on the intent analysis results.

[0072] The dialogue intent information determination device provided in the embodiments of the present invention can execute the dialogue intent information determination method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0073] Figure 4A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0074] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0075] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0076] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for determining dialogue intent information.

[0077] In some embodiments, the dialogue intent information determination method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the dialogue intent information determination method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the dialogue intent information determination method by any other suitable means (e.g., by means of firmware).

[0078] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0079] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0080] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0081] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0082] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0083] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0084] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0085] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for determining dialogue intent information, characterized in that, include: Obtain target speech information of a target user in a dialogue, wherein the target speech information includes current speech information, historical speech information, and dialogue parameter information; Complexity analysis is performed on the target speech information to determine the overall complexity value corresponding to the target speech information; Based on the pre-built complexity level threshold and the comprehensive complexity value, target reasoning modules for intent analysis are selected from the pre-built set of reasoning modules. The set of reasoning modules includes at least rule-layer reasoning modules, lightweight model reasoning modules, and large model-layer reasoning modules. The target reasoning module performs intent recognition processing on the target speech information to determine the dialogue intent information of the target user, wherein the dialogue intent information includes the target intent result and the intent confidence value.

2. The method according to claim 1, characterized in that, The step of performing complexity analysis processing based on the target speech information to determine the comprehensive complexity value corresponding to the target speech information includes: The target speech information is subjected to feature extraction processing to obtain speech feature information, wherein the speech feature information includes, but is not limited to, text normalization features, lexical diversity features, round normalization features, and behavior duration features; The target speech information is subjected to ambiguity detection processing to determine the speech ambiguity score of the target speech information; Complexity analysis is performed based on the speech feature information and the speech ambiguity score to determine the comprehensive complexity value corresponding to the target speech information.

3. The method according to claim 2, characterized in that, The step of performing ambiguity detection processing on the target speech information and determining the speech ambiguity score of the target speech information includes: Determine the hesitation expression information contained in the current speech information within the target speech information; Determine the negative expression information contained in the historical speech information within the target speech information; Based on the hesitant expression information and the negative expression information, calculate the speech ambiguity score of the target speech information.

4. The method according to claim 1, characterized in that, The complexity level thresholds include at least a first complexity level threshold, a second complexity level threshold, and a third complexity level threshold; The step of selecting and determining target inference modules for intent analysis from a pre-built set of inference modules based on a pre-built complexity level threshold and the overall complexity value includes: If the overall complexity value belongs to the first complexity level threshold, the rule layer reasoning module is determined as the target reasoning module for analyzing the target speech information; If the overall complexity value belongs to the second complexity level threshold, both the rule layer reasoning module and the lightweight model reasoning module are determined as the target reasoning modules for analyzing the target speech information; If the overall complexity value belongs to the third complexity level threshold, the rule layer reasoning module, the lightweight model reasoning module, and the large model layer reasoning module are all identified as target reasoning modules for analyzing the target speech information.

5. The method according to claim 1, characterized in that, The step of performing intent recognition processing on the target speech information based on the target reasoning module to determine the target user's target intent result includes: For each target inference module, the target speech information is processed by the target inference module to obtain the initial intent result and initial confidence level output by the target inference module; Based on the intent prediction weights corresponding to each target inference module, all primary intent results are weighted and fused to obtain the target intent result. Based on the confidence weights corresponding to each target inference module, all primary confidence values ​​are weighted and fused to obtain the intent confidence value.

6. The method according to claim 1, characterized in that, The dialogue intent information also includes the intent strength level; Determining the intent strength level of the target user includes: Based on the target's speech information, behavioral analysis and processing are performed to determine the target user's user behavior characteristics; Based on the user behavior characteristics, the behavior intensity is calculated to determine the behavior intensity score corresponding to the user behavior characteristics; The intent intensity level of the target user is determined based on the behavior intensity score and a preset level threshold.

7. The method according to claim 1, characterized in that, The method further includes: The intent information obtained within a preset number of rounds is analyzed for intent results, and the filtering rules of the target inference module and the intent confidence value are adjusted based on the intent analysis results.

8. A device for determining dialogue intent information, characterized in that, include: The speech information acquisition module is used to acquire the target speech information of the target user in the communication dialogue, wherein the target speech information includes current speech information, historical speech information and dialogue parameter information; The complexity value determination module is used to perform complexity analysis processing based on the target speech information to determine the comprehensive complexity value corresponding to the target speech information; The reasoning module filtering module is used to filter and determine the target reasoning module for intent analysis from a pre-built set of reasoning modules based on a pre-built complexity level threshold and the comprehensive complexity value. The set of reasoning modules includes at least a rule-layer reasoning module, a lightweight model reasoning module, and a large model-layer reasoning module. The intent information determination module is used to perform intent recognition processing on the target speech information based on the target reasoning module to determine the dialogue intent information of the target user, wherein the dialogue intent information includes the target intent result and the intent confidence value.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, which is executed by the at least one processor to enable the at least one processor to perform the dialogue intent information determination method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the dialogue intent information determination method according to any one of claims 1-7.