Intent recognition model training method, device, equipment, storage medium and product

By mapping the category labels of the intent sample data to the hyperspherical space and controlling the distribution of category features, combining calibration direction and scale correction of the initial intent recognition model, the reliability problem of the intent recognition model when the probability of category prediction is high is solved, and the accuracy of the identification results is improved.

CN114565036BActive Publication Date: 2025-09-02BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210179911.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-09-02
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The existing intention identification model cannot accurately measure the reliability of the identification results when the category prediction probability is high, resulting in overconfidence in the identification results and affecting user experience and accuracy.

Method used

Map the category labels of the intent sample data to the hyperspherical space, control the distance between multiple category features, make them evenly distributed, and correct the initial intent identification model based on the calibration direction and calibration scale to obtain the target intent identification model.

Benefits of technology

The accuracy of the intent recognition results of the intent recognition model is improved, and the reliability and accuracy of the identification results are characterized by more refined category characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565036B_ABST
    Figure CN114565036B_ABST
Patent Text Reader

Abstract

The present application provides a training method, device, equipment, storage medium and product for an intent recognition model, which belongs to the field of Internet technology. The method includes: obtaining multiple intent sample data and an intent category label for each intent sample data; mapping the intent category label for each intent sample data to a hypersphere space to obtain multiple first category features, and the distance between each first category feature and the other first category features is not greater than a preset distance; based on the multiple intent sample data and the multiple first category features, determining the calibration direction and calibration scale of the initial intent recognition model; based on the calibration direction and calibration scale, correcting the initial intent recognition model to obtain a target intent recognition model. This method can improve the accuracy of the intent recognition result of the target intent recognition model obtained after correcting the initial intent recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a training method, apparatus, device, storage medium, and product for an intent recognition model. Background Art

[0002] Intent recognition models are widely used in tasks such as dialogue tasks, search tasks, and other tasks to identify user intentions. During the training process of the intent recognition model, the user's historical intent data is collected as intent sample data, and then the intent recognition model is trained based on the intent sample data and the intent category labels corresponding to the intent sample data. However, the intent recognition model trained in the prior art has a high category prediction probability for the intent data, and the category prediction probability reflects the possibility that the recognition result of the intent recognition model is correct. In this way, relying on a high category prediction probability cannot accurately measure the reliability of the recognition result of the intent recognition model, that is, the category prediction probability output by the intent recognition model is generally too high, making the recognition result overly confident, which in turn affects the user experience and the accuracy of the intent recognition result obtained by the intent recognition model. Summary of the Invention

[0003] The present invention provides a method, apparatus, device, storage medium, and product for training an intent recognition model, which can improve the accuracy of the intent recognition results of the intent recognition model. The technical solution is as follows:

[0004] In one aspect, a method for training an intent recognition model is provided, the method comprising:

[0005] Obtain multiple intent sample data and the intent category label of each intent sample data;

[0006] Mapping the intent category label of each intent sample data to the hypersphere space to obtain multiple first category features, where the distance between each first category feature and any other first category feature is no greater than a preset distance;

[0007] Determining a calibration direction and a calibration scale of an initial intent recognition model based on the plurality of intent sample data and the plurality of first category features;

[0008] Based on the calibration direction and calibration scale, the initial intent recognition model is corrected to obtain the target intent recognition model.

[0009] In one implementation, based on the calibration direction and the calibration scale, the initial intent recognition model is corrected to obtain the target intent recognition model, including:

[0010] Determine category prediction results of multiple intent sample data based on the calibration direction and calibration scale;

[0011] Based on the first loss value between the category prediction result of each intent sample data and the preset category result of the intent sample data, the initial intent recognition model is corrected to obtain the target intent recognition model.

[0012] In one implementation, determining category prediction results for a plurality of intent sample data based on the calibration direction and the calibration scale includes:

[0013] Determine a first probability feature based on the calibration direction and the calibration scale, where the first probability feature includes a predicted probability of each intent sample data corresponding to each first category feature;

[0014] For each intention sample data, determine the maximum probability among multiple prediction probabilities corresponding to the intention sample data, and use the first category feature corresponding to the maximum probability as the category prediction result of the intention sample data.

[0015] In one implementation, based on a first loss value between a category prediction result of each intent sample data and a preset category result of the intent sample data, the initial intent recognition model is corrected to obtain a target intent recognition model, including:

[0016] Obtaining a second loss value, where the second loss value is obtained based on first parameters and second parameters of the plurality of intent sample data, where the first parameter and second parameter of each intent sample data are respectively used to indicate whether a category prediction result corresponding to the intent sample data is accurate and whether it is certain;

[0017] A weighted sum of the average of the multiple first loss values ​​corresponding to the multiple intent sample data and the second loss value is obtained to obtain a third loss value;

[0018] Based on the third loss value, the initial intent recognition model is corrected to obtain the target intent recognition model.

[0019] In one implementation, obtaining the second loss value includes:

[0020] Obtaining first and second parameters of multiple intent sample data;

[0021] Classifying the plurality of intent sample data based on first parameters and second parameters of the plurality of intent sample data to obtain a plurality of sample sets;

[0022] A second loss value is determined based on the amount of intent sample data in each sample set.

[0023] In one implementation, the multiple sample sets include a first set, a second set, a third set, and a fourth set, the number of intent sample data corresponding to the first set, the second set, the third set, and the fourth set are a first number, a second number, a third number, and a fourth number, respectively. Determining the second loss value based on the number of intent sample data in each sample set includes:

[0024] determining a first ratio between the second quantity and a first sum, and determining a second ratio between the fourth quantity and the second sum, the first sum being the sum of the first quantity and the second quantity, and the second sum being the sum of the third quantity and the fourth quantity;

[0025] determining a logarithm of the sum of the first ratio, the second ratio, and a preset constant to obtain a second loss value;

[0026] Among them, the first parameter of each intention sample data in the first set indicates that its category prediction result is accurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the second set indicates that its category prediction result is accurate, and the second parameter is greater than the first threshold; the first parameter of each intention sample data in the third set indicates that its category prediction result is inaccurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the fourth set indicates that its category prediction result is inaccurate, and the second parameter is greater than the first threshold.

[0027] In one implementation, the process of obtaining the first threshold includes:

[0028] An average value of the second parameter of the plurality of intention sample data is determined to obtain a first threshold.

[0029] In one implementation, the process of obtaining the second parameter of a plurality of intent sample data includes:

[0030] For each intention sample data, obtain the predicted probability of the intention sample data corresponding to each first category feature to form a second probability feature;

[0031] A second parameter is determined based on a product of the second probability feature of the intent sample data and a logarithmic value of the second probability feature.

[0032] In one implementation, the intent category label of each intent sample data is mapped to a hypersphere space to obtain multiple first category features, including:

[0033] Encode multiple intent category labels separately to obtain multiple second category features;

[0034] For each second category feature, determining the distance between the second category feature and other second category features;

[0035] Based on the multiple distances corresponding to the second category feature, the second category feature is corrected, and when the multiple distances are not greater than the preset distance, the first category feature corresponding to the second category feature is obtained.

[0036] In one implementation, based on multiple distances corresponding to the second category feature, correcting the second category feature, and obtaining the first category feature corresponding to the second category feature when none of the multiple distances is greater than a preset distance, includes:

[0037] determining an average between a maximum value among the plurality of distances and a maximum value corresponding to each second category feature;

[0038] Correcting the second category features based on the maximum and mean values;

[0039] When the difference between the maximum value and the average value corresponding to the corrected second category feature is not greater than the second threshold, it is determined that the multiple distances are not greater than the preset distance, and the corrected second category feature is determined as the first category feature.

[0040] In one implementation, determining a calibration direction and a calibration scale of an initial intent recognition model based on a plurality of intent sample data and a plurality of first category features includes:

[0041] Determine the intent sample features of each intent sample data;

[0042] Perform dot product on multiple intent sample features and multiple first category features to obtain a calibration direction;

[0043] The sum of the moduli of the plurality of first category features is determined to obtain a calibration scale.

[0044] In another aspect, a training apparatus for an intent recognition model is provided, the apparatus comprising:

[0045] An acquisition module is used to obtain multiple intent sample data and the intent category label of each intent sample data;

[0046] a mapping module, configured to map the intent category label of each intent sample data to a hypersphere space to obtain a plurality of first category features, wherein the distance between each first category feature and each other first category feature is no greater than a preset distance;

[0047] A first determination module is configured to determine a calibration direction and a calibration scale of an initial intent recognition model based on the plurality of intent sample data and the plurality of first category features;

[0048] A correction module is used to correct the initial intention recognition model based on the calibration direction and the calibration scale to obtain a target intention recognition model.

[0049] In one implementation, the correction module includes:

[0050] a first determining unit, configured to determine category prediction results of the plurality of intent sample data based on the calibration direction and the calibration scale;

[0051] The first correction unit is used to correct the initial intent recognition model based on a first loss value between a category prediction result of each intent sample data and a preset category result of the intent sample data to obtain the target intent recognition model.

[0052] In one implementation, the first determining unit is configured to:

[0053] Determining a first probability feature based on the calibration direction and the calibration scale, wherein the first probability feature includes a predicted probability of each intent sample data corresponding to each first category feature;

[0054] For each intention sample data, determine the maximum probability among multiple prediction probabilities corresponding to the intention sample data, and use the first category feature corresponding to the maximum probability as the category prediction result of the intention sample data.

[0055] In one implementation, the first correction unit includes:

[0056] an acquisition subunit, configured to obtain a second loss value, where the second loss value is obtained based on first parameters and second parameters of the plurality of intent sample data, where the first parameter and second parameter of each intent sample data are respectively used to indicate whether a category prediction result corresponding to the intent sample data is accurate and whether it is certain;

[0057] a determination subunit, configured to perform a weighted summation of an average of a plurality of first loss values ​​corresponding to the plurality of intent sample data and the second loss value to obtain a third loss value;

[0058] A correction subunit is used to correct the initial intent recognition model based on the third loss value to obtain the target intent recognition model.

[0059] In one implementation, the acquisition subunit is configured to:

[0060] Acquire first parameters and second parameters of the plurality of intent sample data;

[0061] Classifying the plurality of intent sample data based on the first parameter and the second parameter of the plurality of intent sample data to obtain a plurality of sample sets;

[0062] The second loss value is determined based on the amount of intent sample data in each sample set.

[0063] In one implementation, the multiple sample sets include a first set, a second set, a third set, and a fourth set, and the quantities of intent sample data corresponding to the first set, the second set, the third set, and the fourth set are respectively a first quantity, a second quantity, a third quantity, and a fourth quantity, and the acquisition subunit is used to:

[0064] determining a first ratio between the second amount and a first sum value, which is the sum of the first amount and the second amount, and determining a second ratio between the fourth amount and a second sum value, wherein the first sum value is the sum of the first amount and the second amount, and the second sum value is the sum of the third amount and the fourth amount;

[0065] determining a logarithm of the sum of the first ratio, the second ratio, and a preset constant to obtain the second loss value;

[0066] Among them, the first parameter of each intention sample data in the first set indicates that its category prediction result is accurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the second set indicates that its category prediction result is accurate, and the second parameter is greater than the first threshold; the first parameter of each intention sample data in the third set indicates that its category prediction result is inaccurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the fourth set indicates that its category prediction result is inaccurate, and the second parameter is greater than the first threshold.

[0067] In one implementation, the apparatus further includes:

[0068] The second determination module is used to determine the average value of the second parameter of the multiple intention sample data to obtain the first threshold.

[0069] In one implementation, the acquisition subunit is configured to acquire, for each piece of intent sample data, a predicted probability of the intent sample data corresponding to each first category feature, to form a second probability feature;

[0070] The second parameter is determined based on the product of the second probability feature of the intention sample data and the logarithm of the second probability feature.

[0071] In one implementation, the mapping module includes:

[0072] an encoding unit, configured to encode the plurality of intent category labels respectively to obtain a plurality of second category features;

[0073] A second determining unit is configured to determine, for each second category feature, a distance between the second category feature and each of the other second category features;

[0074] The second correction unit is configured to correct the second category feature based on a plurality of distances corresponding to the second category feature, and obtain the first category feature corresponding to the second category feature when none of the plurality of distances is greater than the preset distance.

[0075] In one implementation, the second correction unit is configured to:

[0076] determining an average value between a maximum value among the plurality of distances and a maximum value corresponding to each second category feature;

[0077] Correcting the second category feature based on the maximum value and the average value;

[0078] When the difference between the maximum value and the average value corresponding to the corrected second category feature is not greater than the second threshold, it is determined that the multiple distances are not greater than the preset distance, and the corrected second category feature is determined as the first category feature.

[0079] In one implementation, the first determining module is configured to:

[0080] Determine the intent sample features of each intent sample data;

[0081] Performing a dot product on the plurality of intent sample features and the plurality of first category features to obtain the calibration direction;

[0082] The sum of the moduli of the plurality of first category features is determined to obtain the calibration scale.

[0083] On the other hand, a computer device is provided, which includes one or more processors and one or more memories, wherein at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the training method of the intent recognition model described in any of the above implementation methods.

[0084] On the other hand, a computer-readable storage medium is provided, in which at least one program code is stored. The at least one program code is loaded and executed by a processor to implement the training method of the intent recognition model described in any of the above implementation methods.

[0085] On the other hand, a computer program product is provided, which includes a computer program code, the computer program code is stored in a computer-readable storage medium, a processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the training method of the intent recognition model described in any of the above implementations.

[0086] The present application provides a training method for an intent recognition model. Since the method maps intent category labels to a hypersphere space and controls the distances between multiple category features corresponding to multiple intent category labels, the multiple category features are distributed as evenly as possible in the hypersphere space, thereby obtaining category features that represent the intent recognition labels more finely, so that the category features can more accurately represent the intent category labels; and then, based on such category features and intent sample data, the calibration direction and calibration scale for correcting the initial intent recognition model are determined, thereby improving the accuracy of the intent recognition results of the target intent recognition model obtained after correcting the initial intent recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0088] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0089] Figure 2 This is a flowchart of a method for training an intent recognition model provided in an embodiment of the present application;

[0090] Figure 3 This is a flowchart of a method for training an intent recognition model provided in an embodiment of the present application;

[0091] Figure 4 This is a schematic diagram of encoding an intent category label provided in an embodiment of the present application;

[0092] Figure 5 This is a flowchart of a method for training an intent recognition model provided in an embodiment of the present application;

[0093] Figure 6 This is a flowchart of a method for training an intent recognition model provided in an embodiment of the present application;

[0094] Figure 7 This is a block diagram of a training device for an intent recognition model provided in an embodiment of the present application;

[0095] Figure 8 This is a block diagram of a terminal provided in an embodiment of the present application;

[0096] Figure 9 This is a block diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0097] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0098] It should be noted that the user information involved in this application (including but not limited to user device information, user personal information, etc.) is information authorized by the user or fully authorized by all parties.

[0099] The terms "first," "second," "third," and "fourth," etc. in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, not to describe a specific order. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0100] An embodiment of the present application provides an implementation environment for a training method of an intent recognition model, which includes a computer device 10 and an intent recognition device 20. The intent recognition device 20 is any device that can perform intent recognition on the user's intent data. The computer device 10 is used to obtain intent sample data, perform model training based on the intent sample data, obtain an intent recognition model, and send it to the intent recognition device 20, so that the intent recognition device 20 can realize intent recognition based on the trained intent recognition model. In some embodiments, the intent recognition device 20 can be an intelligent robot, an intelligent customer service device, or a human-computer dialogue device, etc., but is not limited thereto. Among them, the computer device 10 can be provided as a terminal or a server; in the embodiment of the present application, the computer device 10 is provided as a server as an example for explanation.

[0101] In some embodiments, the intent recognition device 20 is an intelligent robot, which is used to provide services in exhibition halls, airports, hospitals, libraries, and other places. The intelligent robot stores an intent recognition model. After obtaining the user's intent data, the intelligent robot uses the intent recognition model to determine the intent recognition result corresponding to the intent data, and then provides services to the user based on the intent recognition result.

[0102] In some embodiments, the intent recognition device 20 is an intelligent customer service device, which is used in scenarios such as flight ticket inquiries, communication package inquiries, weather inquiries, and location inquiries. The intelligent customer service device stores an intent recognition model. After obtaining user intent data, the intelligent customer service device uses the intent recognition model to determine the intent recognition result corresponding to the intent data, and then provides query services to the user based on the intent recognition result.

[0103] In some embodiments, the intent recognition device 20 is a human-machine interactive device used in scenarios such as shopping and ticket purchase. The human-machine interactive device stores an intent recognition model. After obtaining user intent data, the human-machine interactive device uses the intent recognition model to determine the intent recognition result corresponding to the intent data, and then provides the user with purchase services based on the intent recognition result.

[0104] The present application embodiment provides a method for training an intent recognition model, the execution subject is a computer device, see Figure 2 , methods include:

[0105] Step 201: Obtain multiple intent sample data and the intent category label of each intent sample data.

[0106] Step 202: Map the intent category label of each intent sample data to the hypersphere space to obtain a plurality of first category features, wherein the distance between each first category feature and any other first category feature is no greater than a preset distance.

[0107] Step 203: Determine a calibration direction and calibration scale of the initial intent recognition model based on the plurality of intent sample data and the plurality of first category features.

[0108] Step 204: Based on the calibration direction and calibration scale, calibrate the initial intent recognition model to obtain a target intent recognition model.

[0109] In one implementation, based on the calibration direction and the calibration scale, the initial intent recognition model is corrected to obtain the target intent recognition model, including:

[0110] Determine category prediction results of multiple intent sample data based on the calibration direction and calibration scale;

[0111] Based on the first loss value between the category prediction result of each intent sample data and the preset category result of the intent sample data, the initial intent recognition model is corrected to obtain the target intent recognition model.

[0112] In one implementation, determining category prediction results for a plurality of intent sample data based on the calibration direction and the calibration scale includes:

[0113] Determine a first probability feature based on the calibration direction and the calibration scale, where the first probability feature includes a predicted probability of each intent sample data corresponding to each first category feature;

[0114] For each intention sample data, determine the maximum probability among multiple prediction probabilities corresponding to the intention sample data, and use the first category feature corresponding to the maximum probability as the category prediction result of the intention sample data.

[0115] In one implementation, based on a first loss value between a category prediction result of each intent sample data and a preset category result of the intent sample data, the initial intent recognition model is corrected to obtain a target intent recognition model, including:

[0116] Obtaining a second loss value, where the second loss value is obtained based on first parameters and second parameters of the plurality of intent sample data, where the first parameter and second parameter of each intent sample data are respectively used to indicate whether a category prediction result corresponding to the intent sample data is accurate and whether it is certain;

[0117] A weighted sum of the average of the multiple first loss values ​​corresponding to the multiple intent sample data and the second loss value is obtained to obtain a third loss value;

[0118] Based on the third loss value, the initial intent recognition model is corrected to obtain the target intent recognition model.

[0119] In one implementation, obtaining the second loss value includes:

[0120] Obtaining first and second parameters of multiple intent sample data;

[0121] Classifying the plurality of intent sample data based on first parameters and second parameters of the plurality of intent sample data to obtain a plurality of sample sets;

[0122] A second loss value is determined based on the amount of intent sample data in each sample set.

[0123] In one implementation, the multiple sample sets include a first set, a second set, a third set, and a fourth set, the number of intent sample data corresponding to the first set, the second set, the third set, and the fourth set are a first number, a second number, a third number, and a fourth number, respectively. Determining the second loss value based on the number of intent sample data in each sample set includes:

[0124] determining a first ratio between the second quantity and a first sum, and determining a second ratio between the fourth quantity and the second sum, the first sum being the sum of the first quantity and the second quantity, and the second sum being the sum of the third quantity and the fourth quantity;

[0125] determining a logarithm of the sum of the first ratio, the second ratio, and a preset constant to obtain a second loss value;

[0126] Among them, the first parameter of each intention sample data in the first set indicates that its category prediction result is accurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the second set indicates that its category prediction result is accurate, and the second parameter is greater than the first threshold; the first parameter of each intention sample data in the third set indicates that its category prediction result is inaccurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the fourth set indicates that its category prediction result is inaccurate, and the second parameter is greater than the first threshold.

[0127] In one implementation, the process of obtaining the first threshold includes:

[0128] An average value of the second parameter of the plurality of intention sample data is determined to obtain a first threshold.

[0129] In one implementation, the process of obtaining the second parameter of a plurality of intent sample data includes:

[0130] For each intention sample data, obtain the predicted probability of the intention sample data corresponding to each first category feature to form a second probability feature;

[0131] A second parameter is determined based on a product of the second probability feature of the intent sample data and a logarithmic value of the second probability feature.

[0132] In one implementation, the intent category label of each intent sample data is mapped to a hypersphere space to obtain multiple first category features, including:

[0133] Encode multiple intent category labels separately to obtain multiple second category features;

[0134] For each second category feature, determining the distance between the second category feature and other second category features;

[0135] Based on the multiple distances corresponding to the second category feature, the second category feature is corrected, and when the multiple distances are not greater than the preset distance, the first category feature corresponding to the second category feature is obtained.

[0136] In one implementation, based on multiple distances corresponding to the second category feature, correcting the second category feature, and obtaining the first category feature corresponding to the second category feature when none of the multiple distances is greater than a preset distance, includes:

[0137] determining an average between a maximum value among the plurality of distances and a maximum value corresponding to each second category feature;

[0138] Correcting the second category features based on the maximum and mean values;

[0139] When the difference between the maximum value and the average value corresponding to the corrected second category feature is not greater than the second threshold, it is determined that the multiple distances are not greater than the preset distance, and the corrected second category feature is determined as the first category feature.

[0140] In one implementation, determining a calibration direction and a calibration scale of an initial intent recognition model based on a plurality of intent sample data and a plurality of first category features includes:

[0141] Determine the intent sample features of each intent sample data;

[0142] Perform dot product on multiple intent sample features and multiple first category features to obtain a calibration direction;

[0143] The sum of the moduli of the plurality of first category features is determined to obtain a calibration scale.

[0144] The present application provides a training method for an intent recognition model. Since the method maps intent category labels to a hypersphere space and controls the distances between multiple category features corresponding to multiple intent category labels, the multiple category features are distributed as evenly as possible in the hypersphere space, thereby obtaining category features that represent the intent recognition labels more finely, so that the category features can more accurately represent the intent category labels; and then, based on such category features and intent sample data, the calibration direction and calibration scale for correcting the initial intent recognition model are determined, thereby improving the accuracy of the intent recognition results of the target intent recognition model obtained after correcting the initial intent recognition model.

[0145] This application embodiment provides a method for training an intent recognition model. Figure 3 , methods include:

[0146] Step 301: A computer device obtains a plurality of intent sample data and an intent category label for each intent sample data.

[0147] Optionally, multiple intent sample data are user historical query data, which are original intent data output by the user, and the intent sample data can be sample data corresponding to any field; for example, the intent sample data is "How much is cabbage?" in the shopping field, or "Buy me a plane ticket" in the intelligent customer service field, etc., which are not specifically limited in the embodiments of the present application. Among them, the intent category label of each intent sample data is used to represent the intent corresponding to the intent sample data, such as the intent category label corresponding to the intent sample data "How much is cabbage?" is "Inquiry", and the intent category label corresponding to "Buy me a plane ticket" is "Buy a ticket". Furthermore, the corresponding intent category label can also be more detailed "Ask about the price of cabbage", "Buy a plane ticket", etc. In the embodiments of the present application, the intent category label can be set and changed as needed, and there is no specific limitation on this. Optionally, the computer device performs model training on the intent sample data in different fields respectively to obtain intent recognition models in different fields.

[0148] Optionally, the number of the plurality of intent sample data is N, and the plurality of intent sample data are respectively represented as {Q1, ..., Q i ,…,Q N}, Q1, Q i and Q N etc. represent an intent sample data respectively; the corresponding intent category labels are {T1,…,T i ,…,T N}, T1, T i and T N etc. represent an intent category label, where T i ∈ C. C = {1, ..., K} represents the set of K intent category labels.

[0149] Step 302: The computer device maps the intent category label of each intent sample data to the hypersphere space to obtain multiple first category features.

[0150] Wherein, the distance between each first category feature and each other first category feature is not greater than a preset distance. In some embodiments, step 302 includes the following steps (1)-(2):

[0151] (1) The computer device encodes the multiple intent category labels respectively to obtain multiple second category features; for each second category feature, the computer device determines the distance between the second category feature and other second category features.

[0152] Optionally, the computer device encodes multiple category labels respectively to obtain multiple second category features, which are all encoding vectors. The first category features and the second category features correspond to the first category vector and the second category vector respectively. The distance between the second category features determined by the computer device and other second category features is the cosine distance between the vectors.

[0153] It should be noted that before encoding multiple intent category labels, the computer device encodes the H-dimensional output space S composed of multiple intent category labels. H The data is uniformly divided into K subspaces, so that the number of these subspaces is the same as the number of labels in the set of intent category labels, and then encoded in the hypersphere space to obtain multiple second category vectors {h1, ..., h i ,…,h K}, h1, h i and h K Represents a second category vector respectively. Optionally, the norm of each second category vector satisfies ||h i ||=1, that is, the norm is 1; optionally, the computer device encodes the intent category label through a hypersphere encoder.

[0154] (2) The computer device corrects the second category feature based on the multiple distances corresponding to the second category feature, and obtains the first category feature corresponding to the second category feature when the multiple distances are not greater than the preset distance.

[0155] In one implementation, a computer device determines an average value between a maximum value among a plurality of distances and a maximum value corresponding to each second category feature; the computer device corrects the second category feature based on the maximum value and the average value; and when the difference between the maximum value and the average value corresponding to the corrected second category feature is not greater than a second threshold, the computer device determines that the plurality of distances are not greater than a preset distance, and determines the corrected second category feature as a first category feature.

[0156] It should be noted that for each second-category vector in the hypersphere space, there is a cosine distance of K-1 between itself and all K-1 second-category vectors. The maximum cosine distance between multiple cosine distances is defined as D i =max(d ij ); where i, j∈C, i≠j. d ij is the second category vector h i With h jThe cosine distance between them. Among them, the second threshold can be set and changed as needed, and the preset distance is the sum of the average of multiple maximum values ​​and the second threshold. It should be noted that, in the optimal case, the maximum value of the distance corresponding to each second category feature is equal to the average value of the maximum values ​​corresponding to multiple second category features, that is, multiple second category features are evenly distributed in the hypersphere space. Therefore, by setting the second threshold to limit the maximum value, when the difference between the maximum value and the average value is not greater than the second threshold, it means that the loss of uniform distribution of multiple second category features does not exceed the second threshold. In this way, adjusting the second category features based on the second threshold can ensure that the loss of uniform distribution of multiple second category features is effectively controlled, thereby enabling multiple second category features to be distributed as evenly as possible in the hypersphere space. In the optimal case, the second threshold is 0, then the preset distance is the average value of multiple maximum values, that is, when the maximum value corresponding to each second category feature is equal to the average value of multiple maximum values, multiple second category features are evenly distributed in the hypersphere space. Optionally, the average value of multiple maximum values ​​is obtained by the following formula 1.

[0157] Formula 1:

[0158] in, is the average of multiple maximum values, k is the number of second category features, D i is the maximum value corresponding to the i-th second category feature. In one implementation, since the norm of all second category vectors is 1, that is, they are all unit vectors, the average value of multiple maximum values ​​can be obtained by the following formula 2.

[0159] Formula 2:

[0160] Where X = [h1,…,h i ,…,h K ], is a matrix composed of multiple second category vectors, h1, h i and h K Each represents a second-category vector, and I is the identity matrix. Zi is the i-th row of Z and is an intermediate parameter. In this implementation, the average of multiple maximum values ​​is obtained through matrix operations, which can improve the efficiency of obtaining the average of multiple maximum values.

[0161] In the embodiment of the present application, since the distance between multiple second-category features can effectively characterize the distribution of multiple second-category features in the hyperspherical space, correcting the second-category features based on the distance can achieve effective correction of the second-category features, so that the multiple second-category features are distributed as evenly as possible in the hyperspherical space.

[0162] In the embodiment of the present application, the computer device maps the intent category label to the hypersphere space to obtain the category coding vector of the hypersphere space. Since the coding vector of the hypersphere space is represented in multiple dimensions, it can better represent the current intent category label than the one-hot vector of a single dimension. Figure 4 , Figure 4 This is a comparison chart of label encoding between one-hot vectors and hypersphere space. As can be seen from the figure, one-hot vectors are only represented in a single dimension, while vectors encoded in hypersphere space are represented in multiple dimensions. This indicates that encoding intent category labels in hypersphere space yields more accurate category features. Furthermore, in an embodiment of the present application, the distance between each mapped first category feature is no greater than a preset distance, ensuring that the multiple first category features are distributed as evenly as possible, thereby obtaining denser category features. This ensures that the distribution of multiple intent category labels is more consistent with actual conditions, resulting in more accurate representations of the multiple first category features.

[0163] Step 303: The computer device determines a calibration direction and calibration scale of the initial intent recognition model based on the plurality of intent sample data and the plurality of first category features.

[0164] In one implementation, a computer device determines an intent sample feature of each intent sample data; the computer device performs a dot product on multiple intent sample features and multiple first category features to obtain a calibration direction; and the computer device determines the sum of the moduli of the multiple first category features to obtain a calibration scale.

[0165] In one implementation, the computer device uses a text encoder to extract the intent sample features of each intent sample data, and obtains each intent sample data Q i Intent sample features; Optionally, the text encoder is a pre-trained BERT (Bidirectional Encoder Representation from Transformers, a bidirectional encoder) model, and the intent sample feature is an H-dimensional encoding vector E i ; For example, the encoding vector E i is a CLS (classification) vector. The dimension of the intent sample vector corresponding to the intent sample feature is H, which is the same as the dimension of the first category feature. This facilitates subsequent calculations based on the intent sample feature and the first category feature to obtain the calibration direction and calibration scale. In one implementation, the computer device obtains the calibration direction and calibration scale using a hypersphere decoder.

[0166] Optionally, the computer device performs dot product on multiple intent sample features and multiple first category features through the following formula three to obtain a calibration direction.

[0167] Formula 3: Cali D =E·X T

[0168] Where E∈R N×H represents the matrix composed of all intent sample features, N represents the number of intent sample features, H represents the dimension of intent sample features, X T is the transposed matrix of the matrix composed of multiple second category vectors, Cali D is a calibration direction. Optionally, the calibration direction is a multi-dimensional matrix.

[0169] In an embodiment of the present application, since the intent recognition model recognizes intent from intent sample data based on intent category labels, a calibration direction is obtained by performing a dot product between the first category features representing the intent category labels and the intent sample features representing the intent sample data. This allows the obtained calibration direction to fully consider the overall distribution of the intent category labels and the intent sample data, thereby ensuring high accuracy of the obtained calibration direction. Optionally, the computer device determines the sum of the moduli of multiple first category features using the following formula 4 to obtain a calibration scale.

[0170] Formula 4: Cali S =‖X‖

[0171] Among them, Cali S is the calibration scale, X is a matrix consisting of multiple second-category vectors, and ‖X‖ represents the norm of X. In this embodiment of the present application, the calibration scale is obtained by summing the moduli of multiple first-category features. This calibration scale comprehensively considers the feature representations of multiple intent category labels, thereby ensuring high representativeness and accuracy.

[0172] Step 304: The computer device determines category prediction results of the plurality of intent sample data based on the calibration direction and the calibration scale.

[0173] In one implementation, the computer device determines a first probability feature based on a calibration direction and a calibration scale. The first probability feature includes a predicted probability for each first category feature corresponding to each piece of intent sample data. For each piece of intent sample data, the computer device determines the maximum probability among multiple predicted probabilities corresponding to the intent sample data, and uses the first category feature corresponding to the maximum probability as the category prediction result for the intent sample data. Optionally, the computer device determines the first probability feature based on the calibration direction and the calibration scale using the following formula 5.

[0174] Formula 5: L = Cali S ×Cali D

[0175] Among them, Cali D To calibrate the direction, S is the calibration scale; L is the first probability feature, is a matrix, and each element in the matrix represents the predicted probability of each intent sample data corresponding to each first category feature.

[0176] In an embodiment of the present application, since the maximum value among multiple predicted probabilities corresponding to multiple first category features for each intention sample data can indicate which intention category the intention sample data is most likely to belong to, in this way, the first category feature corresponding to the maximum probability among the multiple predicted probabilities is used as the category prediction result, so that the first category feature can more accurately represent the intention recognition result of the intention sample data.

[0177] Step 305: The computer device corrects the initial intent recognition model based on a first loss value between the category prediction result of each intent sample data and the preset category result of the intent sample data to obtain a target intent recognition model.

[0178] Among them, the preset category result is the first category feature corresponding to the real intent category label corresponding to the intent sample data, and the category prediction result is the first category feature corresponding to the intent category label predicted by the intent sample data. The initial intent recognition model is a model with initial training parameters. The first loss value represents the accuracy corresponding to the initial intent recognition model. The first loss value is involved in backpropagation to adjust the initial intent recognition model, thereby achieving correction of the initial intent recognition model. In one implementation, the computer device continuously iteratively corrects the initial intent recognition model based on the first loss value until the first loss value no longer decreases after the initial intent recognition model has been iteratively corrected for multiple times. The training is then terminated, and the intent recognition model corresponding to the minimum first loss value is determined as the target intent recognition model. In another implementation, the computer device continuously iteratively corrects the initial intent recognition model based on the first loss value. When the first loss value is not greater than the first loss threshold, the intent recognition model corresponding to the first loss value is determined as the target intent recognition model. In an embodiment of the present application, by determining the category prediction results of multiple intent sample data, and then correcting the initial intent recognition model based on the first loss value of the category prediction results and the preset category results, since the first loss value can characterize the corresponding accuracy of the initial intent recognition model, and then correcting the initial intent recognition model based on the first loss value, the accuracy of the corrected target intent recognition model can be improved.

[0179] In an embodiment of the present application, by mapping the intent category label to the hypersphere space and making the distance between multiple first category features no greater than the preset distance, the multiple first category features are distributed as evenly as possible in the hypersphere space, and a denser category feature is obtained, so that the intent category label can be represented in multiple dimensions, avoiding the situation where the category feature obtained by the one-hot vector is only represented in a single dimension and cannot accurately characterize the intent category label; and then the calibration direction and calibration scale are obtained by using the category feature obtained in the hypersphere space to correct the initial intent recognition model, so that a target intent recognition model with high accuracy of intent recognition results can be obtained.

[0180] The present application provides a training method for an intent recognition model. Since the method maps intent category labels to a hypersphere space and controls the distances between multiple category features corresponding to multiple intent category labels, the multiple category features are distributed as evenly as possible in the hypersphere space, thereby obtaining category features that represent the intent recognition labels more finely, so that the category features can more accurately represent the intent category labels; and then, based on such category features and intent sample data, the calibration direction and calibration scale for correcting the initial intent recognition model are determined, thereby improving the accuracy of the intent recognition results of the target intent recognition model obtained after correcting the initial intent recognition model.

[0181] This application embodiment provides a method for training an intent recognition model. Figure 5 , methods include:

[0182] Step 501: A computer device obtains a plurality of intent sample data and an intent category label of each intent sample data.

[0183] Step 502: The computer device maps the intent category label of each intent sample data to the hypersphere space to obtain multiple first category features.

[0184] Step 503: The computer device determines the calibration direction and calibration scale of the initial intent recognition model based on the multiple intent sample data and the multiple first category features.

[0185] Step 504: The computer device determines category prediction results of the plurality of intent sample data based on the calibration direction and the calibration scale.

[0186] Steps 501-504 are similar to steps 301-304 and will not be repeated here.

[0187] Step 505: The computer device obtains a second loss value.

[0188] The second loss value is obtained based on the first parameter and the second parameter of the plurality of intent sample data, and the first parameter and the second parameter of each intent sample data are respectively used to indicate whether the category prediction result corresponding to the intent sample data is accurate and whether it is certain. The computer device obtains the second loss value, including the following steps (1)-(3):

[0189] (1) The computer device obtains the first parameter and the second parameter of a plurality of intent sample data.

[0190] Among them, the process of the computer device obtaining the second parameters of multiple intention sample data includes: the computer device obtains the predicted probability of the intention sample data corresponding to each first category feature for each intention sample data, and constitutes a second probability feature; the computer device determines the second parameter based on the product of the second probability feature of the intention sample data and the logarithm of the second probability feature.

[0191] Optionally, the second parameter represents the uncertainty of the category prediction result corresponding to the intention sample data; the computer device obtains the second parameter through the following formula six based on the product of the second probability feature of the intention sample data and the logarithm of the second probability feature.

[0192] Formula 6: u i =-p i logp i

[0193] Among them, u i is the second parameter, p i is the second probability feature, and is a vector, each element of which represents the predicted probability that the intent sample data corresponds to each first category feature. In this implementation, the second parameter is determined based on the second probability feature. Since the second probability feature represents the predicted probability that the intent sample data corresponds to each first category feature, the second parameter representing the uncertainty is determined based on the predicted probability, so that the second parameter matches the predicted probability, which is more consistent with the actual situation and has high accuracy.

[0194] In one implementation, the first parameter is the confidence level corresponding to each intent sample data point, that is, the predicted probability corresponding to the category prediction result. Whether the category prediction result is accurate depends on whether the category prediction result of the intent sample data point is equal to its preset category result, that is, whether it is equal to its true category result. Optionally, the first parameter of each intent sample data point is obtained using the following formula 7.

[0195] Formula 7: a i =max(p i )ifT′ i =T i , a i =1-max(pi )otherwise

[0196] Among them, a i is the first parameter, p i is the second probability feature obtained based on the softmax function, which is a vector. Each element in the vector represents the predicted probability of the intent sample data corresponding to each first category feature. T′ i Represents the category prediction result, T i Represents the preset category result. If the category prediction result is equal to the preset category result, the first parameter is the maximum probability among the multiple prediction probabilities. If the category prediction result is not equal to the preset category result, the first parameter is the difference between the value 1 and the maximum probability. Therefore, when the category prediction result is accurate, the first parameter is close to 1, and when the category prediction result is inaccurate, the first parameter is close to 0.

[0197] (2) The computer device classifies the plurality of intent sample data based on the first parameters and the second parameters of the plurality of intent sample data to obtain a plurality of sample sets.

[0198] The multiple sample sets include a first set, a second set, a third set, and a fourth set, and the numbers of intent sample data corresponding to the first set, the second set, the third set, and the fourth set are the first number, the second number, the third number, and the fourth number, respectively. The first parameter of each intent sample data in the first set indicates that its category prediction result is accurate, and the second parameter is not greater than the first threshold; the first parameter of each intent sample data in the second set indicates that its category prediction result is accurate, and the second parameter is greater than the first threshold; the first parameter of each intent sample data in the third set indicates that its category prediction result is inaccurate, and the second parameter is not greater than the first threshold; the first parameter of each intent sample data in the fourth set indicates that its category prediction result is inaccurate, and the second parameter is greater than the first threshold. That is, the first set is a set of intent sample data with accurate and certain category prediction results; the second set is a set of intent sample data with accurate and uncertain category prediction results; the third set is a set of intent sample data with inaccurate and certain category prediction results; and the fourth set is a set of intent sample data with inaccurate and uncertain category prediction results.

[0199] The process of obtaining the first threshold value includes: the computer device determines the average value of the second parameter of multiple intention sample data to obtain the first threshold value. θ ∈[0,1], that is, the first threshold value ranges from 0 to 1. In the embodiment of the present application, the first threshold value is determined by the average value of the second parameter of multiple intent sample data, so that the first threshold value is more consistent with the actual distribution of the second parameter, thereby making the first threshold value more accurate and more representative.

[0200] It should be noted that a well-calibrated intent recognition model provides low uncertainty for accurate category prediction results and high uncertainty for inaccurate category prediction results. Therefore, a well-calibrated intent recognition model should produce a high AVU value, where the AVU value is the ratio of the sum of the first quantity of the first set and the fourth quantity of the fourth set to the total number of intent sample data included in the four sets, AVU∈[0,1]. In some implementations, the AVU function corresponding to the AVU value can calibrate the model. To make the AVU function differentiable with respect to the model parameters, the first quantity, second quantity, third quantity, and fourth quantity corresponding to the four sets are expressed by the following formula 8.

[0201] Formula 8:

[0202]

[0203] Among them, n AC Indicates the first quantity, n AU Represents the second quantity, n IC Represents the third quantity, n IU represents the fourth quantity, i represents the i-th intention sample data, T′ i Represents the category prediction result, T i Indicates the preset category result, u i Indicates the second parameter, u θ represents the first threshold, a i Indicates the first parameter.

[0204] (3) The computer device determines a second loss value based on the number of intent sample data in each sample set.

[0205] Accordingly, the computer device determines a second loss value based on the number of intent sample data in each sample set, including the following steps: the computer device determines a first ratio between the second number and the first sum, and determines a second ratio between the fourth number and the second sum, the first sum being the sum of the first number and the second number, and the second sum being the sum of the third number and the fourth number; the computer device determines the logarithm of the sum of the first ratio, the second ratio and a preset constant to obtain a second loss value.

[0206] Optionally, the preset constant is 1, and the computer device obtains the second loss value based on the number of intent sample data in each sample set through the following formula 9. Formula 9:

[0207] in, Represents the second loss value, n AU Indicates the first quantity, n AU Represents the second quantity, nIC Represents the third quantity, n IU It should be noted that when n AU and n IC When optimized to close to zero, the second loss value is close to zero, indicating that the intent sample data in the second and third sets are getting smaller and smaller, while the intent sample data in the first and fourth sets are getting larger and larger, that is, the intent sample data with accurate and certain category prediction results and the intent sample data with inaccurate and uncertain category prediction results in multiple intent sample data are increasing, while the intent sample data with accurate and uncertain category prediction results and the intent sample data with inaccurate and certain category prediction results are decreasing. This indicates that the intent recognition model accurately predicts the intent sample data with accurate category prediction results, and does not make overly confident predictions for the intent sample data with inaccurate category prediction results, which indicates that the recognition results of the intent recognition model are more accurate. The initial intent recognition model is continuously corrected by the second loss value, so that the recognition results of the initial intent recognition model continue to tend to be accurate and certain, thereby improving the accuracy of the output results of the corrected intent recognition model.

[0208] In an embodiment of the present application, since the first parameter and the second parameter of each intent sample data can respectively represent the accuracy and uncertainty of the category prediction result, and then the multiple intent sample data are classified based on these two parameters, it is possible to achieve effective classification of the multiple intent sample data based on accuracy and uncertainty, and then determine the second loss value based on the intent sample data in each category, so that the second loss value can effectively represent the accuracy and uncertainty of the multiple intent sample data, and then subsequently correct the initial intent recognition model based on the second loss value, which can achieve effective correction of the initial intent recognition model.

[0209] In an embodiment of the present application, the second loss value is determined based on the ratio of the second quantity to the first sum value and the ratio of the third quantity to the second sum value, so that the smaller the second quantity and the third quantity are, the smaller the second loss value is, so that the second loss value can effectively characterize the accuracy and determination of the classification in multiple intent sample data, and then the initial intent recognition model is subsequently corrected based on the second loss value, which can achieve effective correction of the initial intent recognition model.

[0210] Step 506: The computer device performs a weighted summation of the average of the multiple first loss values ​​corresponding to the multiple intent sample data and the second loss value to obtain a third loss value; based on the third loss value, the initial intent recognition model is corrected to obtain a target intent recognition model.

[0211] Each first loss value is the loss value between the category prediction result of each intent sample data and the preset category result of the intent sample data. The weight ratio corresponding to the first loss value and the second loss value can be set and changed as needed, and is not specifically limited in the embodiments of the present application.

[0212] In one implementation, the computer device continuously iteratively corrects the initial intent recognition model based on the third loss value until the third loss value of the initial intent recognition model no longer decreases after multiple iterations of iterative correction. Training is terminated, and the intent recognition model corresponding to the minimum third loss value is determined as the target intent recognition model. In another implementation, the computer device continuously iteratively corrects the initial intent recognition model based on the third loss value. When the third loss value is no greater than a second loss threshold, the intent recognition model corresponding to the third loss value is determined as the target intent recognition model.

[0213] In some embodiments, the intent recognition model provided by the present invention achieved superior recognition performance compared to pre-trained models such as BERT on three publicly available intent recognition datasets. While the accuracy and other metrics of multiple models decreased, the confidence calibration metric, ECE (expected calibration error), of the intent recognition model provided by the present invention decreased by an average of 10.50%, indicating a high confidence level for the intent recognition model, indicating reliable category prediction results. Using the TNEWS (news information) dataset as an example, the F1 value (representing accuracy) improved by 1.21%, while the ECE decreased by 29.67%.

[0214] In some embodiments, the intent recognition model provided based on the embodiments of the present application can be applied to projects such as intelligent customer service voice robots and precise routing in the voice interaction department. For example, some inbound or outbound calls currently rely on IVR (Interactive Voice Response) answer button routing and manual agents. IVR routing has low accuracy, poor user experience, and high cost. The AI ​​voice robot that uses the intent recognition model provided by the embodiments of the present application can quickly copy the words of the manual agents, and has the characteristics of low cost, high concurrency, high stability, and high endurance. Using AI outbound call robots to replace manual agents to complete outbound call tasks will greatly save operating costs.

[0215] See also Figure 6 , Figure 6A training flow chart of the intent recognition model provided in an embodiment of the present application. After the computer device obtains the intent sample data, the text encoder is used to extract the intent sample features of the intent sample data. A hypersphere encoder is used to extract the first category features of the intent category label corresponding to the intent sample data; multiple intent sample features and multiple first category features are then input into the hypersphere decoder to obtain a calibration direction and a calibration scale; then based on the first parameter and the second parameter of each intent sample data, a second loss value for representing accuracy and uncertainty is obtained; the first loss value and the second loss value obtained by the calibration direction and the calibration scale are input into an optimizer for correcting the intent recognition model, and the initial intent recognition model is iteratively corrected based on the optimizer to obtain a target intent recognition model.

[0216] In an embodiment of the present application, by mapping the intent category label to the hypersphere space and obtaining evenly distributed and more dense category features to represent the intent category label, the overfitting problem of the intent category label caused by representing the intent category label in sparse one-hot form of vector is alleviated, making the representation of the intent category label more accurate; then, based on the category features and intent sample data of the hypersphere space, the calibration direction and calibration scale for correcting the initial intent recognition model are determined to obtain a first loss value, and the first loss value is combined with the second loss value representing the accuracy parameter and uncertainty parameter of the category prediction result to correct the initial intent recognition model, which can achieve effective correction of the initial intent recognition model and reduce the situation where the intent sample recognition model makes overconfident erroneous predictions, thereby improving the accuracy of the intent recognition result of the target intent recognition model obtained after correcting the initial intent recognition model.

[0217] The present application also provides a training device for an intent recognition model, see Figure 7 , the device comprises:

[0218] An acquisition module 701 is configured to acquire a plurality of intent sample data and an intent category label for each intent sample data;

[0219] A mapping module 702 is configured to map the intent category label of each intent sample data to a hypersphere space to obtain a plurality of first category features, wherein the distance between each first category feature and any other first category feature is no greater than a preset distance;

[0220] A first determination module 703 is configured to determine a calibration direction and a calibration scale of an initial intent recognition model based on a plurality of intent sample data and a plurality of first category features;

[0221] The correction module 704 is used to correct the initial intention recognition model based on the calibration direction and calibration scale to obtain a target intention recognition model.

[0222] In one implementation, the correction module 704 includes:

[0223] A first determining unit, configured to determine category prediction results of a plurality of intent sample data based on a calibration direction and a calibration scale;

[0224] The first correction unit is used to correct the initial intent recognition model based on a first loss value between the category prediction result of each intent sample data and the preset category result of the intent sample data to obtain a target intent recognition model.

[0225] In one implementation, the first determining unit is configured to:

[0226] Determine a first probability feature based on the calibration direction and the calibration scale, where the first probability feature includes a predicted probability of each intent sample data corresponding to each first category feature;

[0227] For each intention sample data, determine the maximum probability among multiple prediction probabilities corresponding to the intention sample data, and use the first category feature corresponding to the maximum probability as the category prediction result of the intention sample data.

[0228] In one implementation, the first correction unit includes:

[0229] an acquisition subunit, configured to obtain a second loss value, the second loss value being obtained based on first and second parameters of a plurality of intent sample data, the first and second parameters of each intent sample data being respectively used to indicate whether a category prediction result corresponding to the intent sample data is accurate and certain;

[0230] a determination subunit, configured to perform a weighted summation of an average of a plurality of first loss values ​​corresponding to a plurality of intent sample data and a second loss value to obtain a third loss value;

[0231] The correction subunit is used to correct the initial intent recognition model based on the third loss value to obtain a target intent recognition model.

[0232] In one implementation, a subunit is obtained for:

[0233] Obtaining first and second parameters of multiple intent sample data;

[0234] Classifying the plurality of intent sample data based on first parameters and second parameters of the plurality of intent sample data to obtain a plurality of sample sets;

[0235] A second loss value is determined based on the amount of intent sample data in each sample set.

[0236] In one implementation, the multiple sample sets include a first set, a second set, a third set, and a fourth set. The quantities of intent sample data corresponding to the first set, the second set, the third set, and the fourth set are respectively a first quantity, a second quantity, a third quantity, and a fourth quantity. The acquisition subunit is configured to:

[0237] determining a first ratio between the second quantity and a first sum, and determining a second ratio between the fourth quantity and the second sum, the first sum being the sum of the first quantity and the second quantity, and the second sum being the sum of the third quantity and the fourth quantity;

[0238] determining a logarithm of the sum of the first ratio, the second ratio, and a preset constant to obtain a second loss value;

[0239] Among them, the first parameter of each intention sample data in the first set indicates that its category prediction result is accurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the second set indicates that its category prediction result is accurate, and the second parameter is greater than the first threshold; the first parameter of each intention sample data in the third set indicates that its category prediction result is inaccurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the fourth set indicates that its category prediction result is inaccurate, and the second parameter is greater than the first threshold.

[0240] In one implementation, the apparatus further includes:

[0241] The second determination module is used to determine the average value of the second parameter of multiple intention sample data to obtain a first threshold.

[0242] In one implementation, the acquisition subunit is configured to obtain, for each intention sample data, a predicted probability of the intention sample data corresponding to each first category feature, to form a second probability feature;

[0243] A second parameter is determined based on a product of the second probability feature of the intent sample data and a logarithmic value of the second probability feature.

[0244] In one implementation, the mapping module 702 includes:

[0245] An encoding unit, configured to encode the plurality of intent category labels respectively to obtain a plurality of second category features;

[0246] A second determining unit is configured to determine, for each second category feature, a distance between the second category feature and other second category features;

[0247] The second correction unit is configured to correct the second category feature based on a plurality of distances corresponding to the second category feature, and obtain the first category feature corresponding to the second category feature when the plurality of distances are not greater than a preset distance.

[0248] In one implementation, the second correction unit is configured to:

[0249] determining an average between a maximum value among the plurality of distances and a maximum value corresponding to each second category feature;

[0250] Correcting the second category features based on the maximum and mean values;

[0251] When the difference between the maximum value and the average value corresponding to the corrected second category feature is not greater than the second threshold, it is determined that the multiple distances are not greater than the preset distance, and the corrected second category feature is determined as the first category feature.

[0252] In one implementation, the first determining module 703 is configured to:

[0253] Determine the intent sample features of each intent sample data;

[0254] Perform dot product on multiple intent sample features and multiple first category features to obtain a calibration direction;

[0255] The sum of the moduli of the plurality of first category features is determined to obtain a calibration scale.

[0256] Figure 8 The following is a block diagram of a terminal 800 according to an exemplary embodiment of the present application. Terminal 800 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 800 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0257] Typically, the terminal 800 includes a processor 801 and a memory 802 .

[0258] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0259] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one program code, which is used to be executed by the processor 801 to implement the training method of the intent recognition model provided in the method embodiment of the present application.

[0260] In some embodiments, terminal 800 may optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 803 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.

[0261] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0262] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 804 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.

[0263] Display screen 805 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. When display screen 805 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 805. These touch signals can be input as control signals to processor 801 for processing. Display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 805, located on the front panel of terminal 800. In other embodiments, there can be at least two display screens 805, located on different surfaces of terminal 800 or in a foldable design. In still other embodiments, display screen 805 can be a flexible display, located on a curved or foldable surface of terminal 800. Display screen 805 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 805 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0264] The camera assembly 806 is used to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0265] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 801 for processing, or input into the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 807 may also include a headphone jack.

[0266] Positioning component 808 is used to locate the current geographic location of terminal 800 to implement navigation or LBS (Location Based Service). Positioning component 808 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, or Russia's Galileo system.

[0267] Power supply 809 is used to power various components in terminal 800. Power supply 809 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 809 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0268] In some embodiments, the terminal 800 further includes one or more sensors 810 , including but not limited to: an acceleration sensor 811 , a gyroscope sensor 812 , a pressure sensor 813 , a fingerprint sensor 814 , an optical sensor 815 , and a proximity sensor 816 .

[0269] The accelerometer 811 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the terminal 800. For example, the accelerometer 811 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 811. The accelerometer 811 can also be used to collect game or user motion data.

[0270] The gyroscope sensor 812 can detect the orientation and rotation angle of the terminal 800. It can work in conjunction with the accelerometer 811 to collect the user's 3D movements of the terminal 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0271] The pressure sensor 813 can be set on the side frame of the terminal 800 and / or the lower layer of the display screen 805. When the pressure sensor 813 is set on the side frame of the terminal 800, it can detect the user's grip signal of the terminal 800, and the processor 801 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is set on the lower layer of the display screen 805, the processor 801 controls the operational controls on the UI interface based on the user's pressure operation on the display screen 805. The operational controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0272] The fingerprint sensor 814 is used to collect the user's fingerprint. The processor 801 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 814, or the fingerprint sensor 814 identifies the user's identity based on the collected fingerprint. When the user's identity is recognized as a trusted identity, the processor 801 authorizes the user to perform relevant sensitive operations, such as unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 814 can be set on the front, back, or side of the terminal 800. When a physical button or manufacturer logo is provided on the terminal 800, the fingerprint sensor 814 can be integrated with the physical button or manufacturer logo.

[0273] The optical sensor 815 is used to detect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity detected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity detected by the optical sensor 815.

[0274] Proximity sensor 816, also known as a distance sensor, is typically located on the front panel of terminal 800. Proximity sensor 816 is used to detect the distance between the user and the front of terminal 800. In one embodiment, when proximity sensor 816 detects that the distance between the user and the front of terminal 800 is gradually decreasing, processor 801 controls display screen 805 to switch from the screen-on state to the screen-off state. When proximity sensor 816 detects that the distance between the user and the front of terminal 800 is gradually increasing, processor 801 controls display screen 805 to switch from the screen-off state to the screen-on state.

[0275] Those skilled in the art will understand that Figure 8 The structure shown in the figure does not constitute a limitation on the terminal 800, and the terminal 800 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0276] Figure 9 It is a block diagram of a server provided by an embodiment of the present disclosure. The server 900 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 901 and one or more memories 902, wherein the memory 902 is used to store executable program code, and the processor 901 is configured to execute the above executable program code to implement the training method of the intent recognition model provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described here.

[0277] In an exemplary embodiment, a storage medium including a program code is further provided, such as a memory 902 including the program code, and the program code can be executed by the processor 901 of the server 900 to complete the training method of the above-mentioned intent recognition model. Optionally, the storage medium can be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0278] An embodiment of the present application also provides a computer-readable storage medium, in which at least one program code is stored, and the at least one program code is loaded and executed by a processor to implement the training method of the intent recognition model of any of the above-mentioned implementation methods.

[0279] An embodiment of the present application also provides a computer program product, which includes computer program code, the computer program code is stored in a computer-readable storage medium, the processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the training method of the intent recognition model of any of the above-mentioned implementation methods.

[0280] In some embodiments, the computer program product involved in the embodiments of the present application can be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. Multiple computer devices distributed at multiple locations and interconnected through a communication network can constitute a blockchain system.

[0281] The present application provides a training method for an intent recognition model. Since the method maps intent category labels to a hypersphere space and controls the distances between multiple category features corresponding to multiple intent category labels, the multiple category features are distributed as evenly as possible in the hypersphere space, thereby obtaining category features that represent the intent recognition labels more finely, so that the category features can more accurately represent the intent category labels; and then, based on such category features and intent sample data, the calibration direction and calibration scale for correcting the initial intent recognition model are determined, thereby improving the accuracy of the intent recognition results of the target intent recognition model obtained after correcting the initial intent recognition model.

[0282] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the intent sample data involved in this application are all obtained with full authorization.

Claims

1. A training method for an intent recognition model, characterized in that: The method comprises: Obtain multiple intent sample data and the intent category label of each intent sample data; Mapping the intent category label of each intent sample data to a hypersphere space to obtain a plurality of first category features, wherein the distance between each first category feature and each other first category feature is no greater than a preset distance; Determining a calibration direction and a calibration scale of an initial intent recognition model based on the plurality of intent sample data and the plurality of first category features; Based on the calibration direction and the calibration scale, the initial intent recognition model is corrected to obtain a target intent recognition model; and determining the calibration direction and calibration scale of the initial intent recognition model based on the plurality of intent sample data and the plurality of first category features includes: Determine the intent sample features of each intent sample data; Performing a dot product on the plurality of intent sample features and the plurality of first category features to obtain the calibration direction; The sum of the moduli of the plurality of first category features is determined to obtain the calibration scale.

2. The method according to claim 1, characterized in that The step of correcting the initial intent recognition model based on the calibration direction and the calibration scale to obtain a target intent recognition model includes: Determining category prediction results of the plurality of intent sample data based on the calibration direction and the calibration scale; Based on a first loss value between a category prediction result of each intent sample data and a preset category result of the intent sample data, the initial intent recognition model is corrected to obtain the target intent recognition model.

3. The method according to claim 2, characterized in that The determining, based on the calibration direction and the calibration scale, category prediction results of the plurality of intent sample data includes: Determining a first probability feature based on the calibration direction and the calibration scale, wherein the first probability feature includes a predicted probability of each intent sample data corresponding to each first category feature; For each intention sample data, determine the maximum probability among multiple prediction probabilities corresponding to the intention sample data, and use the first category feature corresponding to the maximum probability as the category prediction result of the intention sample data.

4. The method according to claim 2, characterized in that The first loss value between the category prediction result of each intent sample data and the preset category result of the intent sample data is used to correct the initial intent recognition model to obtain the target intent recognition model, including: Obtaining a second loss value, where the second loss value is obtained based on first parameters and second parameters of the plurality of intent sample data, where the first parameter and second parameter of each intent sample data are respectively used to indicate whether a category prediction result corresponding to the intent sample data is accurate and certain; Obtain a third loss value by weightedly summing an average of the plurality of first loss values ​​corresponding to the plurality of intent sample data and the second loss value; Based on the third loss value, the initial intent recognition model is corrected to obtain the target intent recognition model.

5. The method according to claim 4, characterized in that The obtaining of the second loss value includes: Acquire first parameters and second parameters of the plurality of intent sample data; Classifying the plurality of intent sample data based on the first parameter and the second parameter of the plurality of intent sample data to obtain a plurality of sample sets; The second loss value is determined based on the amount of intent sample data in each sample set.

6. The method according to claim 5, characterized in that The multiple sample sets include a first set, a second set, a third set, and a fourth set, the numbers of intent sample data corresponding to the first set, the second set, the third set, and the fourth set are respectively a first number, a second number, a third number, and a fourth number, and determining the second loss value based on the number of intent sample data in each sample set includes: determining a first ratio between the second amount and a first sum value, which is the sum of the first amount and the second amount, and determining a second ratio between the fourth amount and a second sum value, wherein the first sum value is the sum of the first amount and the second amount, and the second sum value is the sum of the third amount and the fourth amount; determining a logarithm of the sum of the first ratio, the second ratio, and a preset constant to obtain the second loss value; Among them, the first parameter of each intention sample data in the first set indicates that its category prediction result is accurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the second set indicates that its category prediction result is accurate, and the second parameter is greater than the first threshold; the first parameter of each intention sample data in the third set indicates that its category prediction result is inaccurate, and the second parameter is not greater than the first threshold; the first parameter of each intention sample data in the fourth set indicates that its category prediction result is inaccurate, and the second parameter is greater than the first threshold.

7. The method according to claim 6, characterized in that The process of obtaining the first threshold includes: Determine an average value of the second parameter of the plurality of intention sample data to obtain the first threshold.

8. The method according to claim 4, characterized in that The process of obtaining the second parameter of the plurality of intent sample data includes: For each intention sample data, obtain the predicted probability of the intention sample data corresponding to each first category feature to form a second probability feature; The second parameter is determined based on the product of the second probability feature of the intention sample data and the logarithm of the second probability feature.

9. The method according to claim 1, characterized in that The intent category label of each intent sample data is mapped to a hypersphere space to obtain a plurality of first category features, including: Encode multiple intent category labels separately to obtain multiple second category features; For each second category feature, determining the distance between the second category feature and other second category features respectively; The second category feature is corrected based on a plurality of distances corresponding to the second category feature, and when none of the plurality of distances is greater than the preset distance, the first category feature corresponding to the second category feature is obtained.

10. The method according to claim 9, characterized in that The correcting the second category feature based on the multiple distances corresponding to the second category feature, and obtaining the first category feature corresponding to the second category feature when none of the multiple distances is greater than the preset distance, includes: determining an average value between a maximum value among the plurality of distances and a maximum value corresponding to each second category feature; Correcting the second category feature based on the maximum value and the average value; When the difference between the maximum value and the average value corresponding to the corrected second category feature is not greater than the second threshold, it is determined that the multiple distances are not greater than the preset distance, and the corrected second category feature is determined as the first category feature.

11. A training device for an intent recognition model, characterized in that: The device comprises: An acquisition module is used to obtain multiple intent sample data and the intent category label of each intent sample data; a mapping module configured to map the intent category label of each intent sample data to a hypersphere space to obtain a plurality of first category features, wherein the distance between each first category feature and any other first category feature is no greater than a preset distance; a first determination module configured to determine a calibration direction and calibration scale of an initial intent recognition model based on the plurality of intent sample data and the plurality of first category features; A correction module, configured to correct the initial intent recognition model based on the calibration direction and the calibration scale to obtain a target intent recognition model, wherein determining the calibration direction and calibration scale of the initial intent recognition model based on the plurality of intent sample data and the plurality of first category features comprises: Determine the intent sample features of each intent sample data; Performing a dot product on the plurality of intent sample features and the plurality of first category features to obtain the calibration direction; The sum of the moduli of the plurality of first category features is determined to obtain the calibration scale.

12. A computer device, characterized in that: The computer device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories. The at least one program code is loaded and executed by the one or more processors to implement the training method of the intent recognition model as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that The storage medium stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the training method of the intent recognition model as described in any one of claims 1 to 10.

14. A computer program product, characterized in that The computer program product includes a computer program code, which is stored in a computer-readable storage medium. The processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the training method of the intent recognition model as described in any one of claims 1 to 10.