Device and method for end-to-end user query correction and slot filling
The proposed solution addresses the challenges of user query correction and slot filling in domain-specific languages by employing an end-to-end processing method that corrects user queries and identifies slots using advanced normalization and encoder-decoder techniques, achieving improved accuracy and reduced human intervention across multiple languages and domains.
Patent Information
- Application Number
- PCT/EP2023/080425
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
Existing user query correction and slot filling solutions struggle to handle queries that are not part of natural languages, particularly in domain-specific languages like application software names, and fail to manage distorted inputs such as misspellings, abbreviations, and space-related issues.
A data processing device and method that performs end-to-end user query correction and slot filling by receiving a user query, processing it to correct errors, and performing slot filling operations to identify and classify slots within the corrected query, using a combination of normalization, encoder-decoder architectures, and probabilistic graphical models to handle variations in input and support multiple languages and domains.
The solution effectively corrects user queries and identifies slots across various languages and domains, including application-specific languages, improving the accuracy of query processing and slot detection, and reducing human intervention through active learning and model boosting techniques.
Smart Images

Figure EP2023080425_08052025_PF_FP_ABST
Abstract
Description
[0001] DEVICE AND METHOD FOR END-TO-END USER QUERY CORRECTION AND SLOT FILLING
[0002] FIELD OF THE INVENTION
[0003] This invention relates to natural language processing operations, in particular to user query correction and slot filling.
[0004] BACKGROUND
[0005] The operations of slot filling and input query correction have evolved to the point where many research and commercial solutions are readily available for different natural language processing (NLP) applications. Although this has led to some promising results, the performance of such methods may be degraded when applied to non-NLP data, such as specific mobile network vendor data, bioinformatics domain data, network protocols data, mobile applications data and physical circuit specifications.
[0006] Input query correction is used in various applications and tasks, such as search engines, sentiment analysis and text summarization. The aim of input query correction processes is to detect and correct different errors in the user input, such as misspellings, autocomplete errors, abbreviations and space-related issues. In real-world scenarios, it is common to deal with queries having typos. However, in the “application domain” world, autocomplete errors, abbreviations and space-related issues also need to be commonly dealt. For example, a user wanting to search for “huawei” could mistype the word it in different ways, as shown in the examples in Figure 1 , where different manipulations of characters have been performed. It is estimated that at least 22% of the queries to the Huawei AppGallery are wrongly entered by users.
[0007] The main goal of slot filling is to identify different slots from a running dialog or static input text, where the slots correspond to different contiguous parts of the user’s query. Slot filling is a natural language understanding concept that effectively saves an extracted entity to an object. However, in for example Power Virtual Agents, slot filling means placing the extracted entity value into a variable. Every slot value has its own associated slot type that is created to capture and save the required parameters. Thus, the main challenge in the slot filling task is to extract the target entity. One example of an input query and corresponding slots’ values and types can be observed in Figure 2. The field called “corrected_value” is used to denote the corrected version of the original slot value element. In this example, there are two slot value entities “plain blue wallpapers” of app_content slot type and “free” of “app_features” slot type. Since both slot values are correctly spelled in this example, the corresponding “corrected_value” is the same as the corresponding “slot_value”.
[0008] The main problem of existing input user query correction solutions is that they cannot manage queries that are not part of the natural languages (i.e. language that occurs naturally in a human community, such as English, French or Spanish), either in full or partially. Thus, the existing libraries and approaches for spelling correction work well for general English, French, Spanish, Italian, etc. but are not adapted for a particular domain-specific languages (i.e. a language specialised to a particular domain) such as the application software name domain. General spelling correction might correct “Spotify” to “Notify”, etc. which is incorrect in application domain language.
[0009] Similarly, existing slot filling approaches generally do not reflect the domain classes of application distribution platforms (for example, the Huawei AppGallery) and also do not deal with distorted inputs (i.e. inputs that contain abbreviations, misspellings, autocomplete, spaces, etc.). Current solutions usually target to the natural language scenarios to support other domains that are purely based on natural language. Moreover, current solutions usually detect only obvious slots while more challenging slots that don’t belong to the natural language are frequently missed.
[0010] Some solutions are based on Large Language Models (LLMs). Those solutions generally employ prominent LLMs, but generally do not successfully manage to detect the slots and obtain the user intent when tested on different datasets. The solutions based on LLMs cannot fully solve the problem, either for distorted input queries or for slot filling problem as well for application domain target scenarios. Most such LLMs are originally trained on, and thus more appropriate for, different tasks such as sentence-pair classification, structured prediction, question answering, sentence retrieval, etc. Those LLMs are generally not able to handle noncorrect input and application domain languages.
[0011] It is desirable to develop an approach for user query correction and slot filling that can overcome at least some of the above problems.
[0012] SUMMARY OF THE INVENTION
[0013] According to one aspect there is provided a data processing device for performing end-to-end user query correction and slot filling, the data processing device having one or more processors configured to: receive a user query; process the user query to determine whether the user query is correct with respect to an application domain language or a natural language; based on the processing of the user query, form a corrected user query; perform a slot filling operation for the corrected user query to identify one or more slots for the corrected user query, each identified slot corresponding to a part of the corrected user query; and classify each identified slot of the corrected user query into one of a plurality of classes.
[0014] This approach can handle variations of non-correct input in multilingual settings in addition to the slot filling task and can handle application domain language in addition to many different natural languages.
[0015] The one or more processors may be configured to implement a user query correction block and a slot filling block end-to-end. The user query correction block may be configured to output the corrected user query. The slot filling block may be configured to directly receive the corrected user query output from the user query correction block as input.
[0016] The user query correction block may comprise a normalization block, a pre-encoder block, an encoder block, a decoder block and a post-decoder block. The user query correction block comprising a normalization block, pre-encoder block and post-decoder block in addition to the encoder-decoder can allow for corrections on both natural language in addition to application domain language input (also given in corresponding natural languages).
[0017] The normalization block may be trainable to model character-level manipulations of the user query to mimic the input of an erroneous user query to the encoder block. This may improve the accuracy of the input user query correction block.
[0018] The pre-encoder block may be trainable to select an input to the encoder block by: receiving pairs of character-level manipulations of the user query from the normalization block; and selecting one of each pair of character-level manipulations of the user query for input to the encoder block by determining a similarity between each of the pair of character-level manipulations of the user query and a corresponding ground truth user query. This may allow inputs to the encoder block to be determined during training of a model implemented by the one or more processors for the end-to-end user query correction and slot filling so that the preencoder block can learn to select inputs to the encoder.
[0019] The encoder block and the decoder block may have transformer architectures. This may allow for compatibility with existing language models and encoder-decoders that are the current state-of-the-art for encoder-decoder architectures. This may allow for optimal performance. The post-decoder block may comprise a regulator block for regulating the output of the decoder. The regulator block may regulate the output of the decoder block by applying a verification stage for the tokenized output of the decoder block, which may be bounded by beginning and end tokens. The regulator block may be configured to receive multiple predictions output by the decoder block and respective probability scores for each of the predictions.
[0020] The post-decoder block may be configured to: receive multiple predictions and respective associated probability scores output by the decoder block; select a prediction of the multiple predictions; compare the selected prediction with a dictionary of classes and corresponding prominent tokens; and output a result from the dictionary as the corrected user query. The dictionary may be configured to store a plurality of possible corrected user queries in the application domain language and / or one or more natural languages. This may allow the postdecoder block to select the most suitable result from the dictionary as the corrected user query.
[0021] The regulator block of the post-decoder block may be configured to perform the steps of receiving multiple predictions and respective associated probability scores output by the decoder block and selecting a prediction of the multiple predictions. The regulator block may be configured to select a prediction of the multiple predictions by taking a difference between pairs of predictions output by the decoder and their respective probability values. For example, the regulator block may be configured to receive three predictions and their respective probability values. The difference between the first and the second decoder outputs (diff12), as well as the difference between the second and the third decoder outputs (diff23) probability value predictions can be taken. The regulator block may learn for which of those the first predicted decoder output is the correct one (i.e. when the difference between the first and the second is reliably high, i.e. beyond some threshold T1). The same determination can then be performed for the second ranked prediction, with a threshold T2 for which the second ranked is the correct one. In some examples, the following logic can be employed:
[0022] 1) Take the first ranked output if diff12 >= T1
[0023] 2) If not, take the second output if diff23 >= T2
[0024] 3) If neither of 1) and 2) above are true, take the original first ranked predicted output. The selected output can then be compared with the dictionary of classes.
[0025] The post-decoder block may comprise a second block for comparing the selected prediction output by the regulator block with the dictionary of classes. The one or more processors may be configured to compare the selected prediction with the dictionary by determining a Levenshtein distance or a Jaro-Winkler distance between the selected prediction and one or more classes in the dictionary. This may allow distances between the selected prediction and classes in the dictionary to be determined so that the closest class in the dictionary can be chosen.
[0026] The slot filling block may be configured to implement a multi-layer directional transformer encoder. This may allow state-of-the-art performance by the encoder in performing the slot filling task.
[0027] The one or more processors may be configured to perform the slot filling operation by implementing a probabilistic graphical model to detect one or more additional slots in the corrected user query and / or to adjust one or more of the identified slots in the correct user query. This may advantageously allow additional slots to be detected in the corrected user query.
[0028] The one or more processors may be configured to form a feature representation for each of the identified slots. The feature representations may be input to a hierarchical and discriminative feature representation block and used to refine the slot prediction output. This may further improve the performance of the slot filling operation and can detect brand new slots (for example, at the application domain language level).
[0029] The slot prediction output by the slot filling block may be refined using a B-l-0 refiner. The refiner may comprise a double binary logistic regression classifier. The feature representations for each of the identified slots may be classified by the classifier to determine the final class of each part of the corrected user query in B-l-0 notation, which is a suitable tagging format for tagging tokens in computational linguistics.
[0030] The one or more processors may be configured to tag each part of the corrected user query with a slot label. This may allow the identified slots, which may correspond to contiguous parts of the corrected user query, to be used in subsequent processes.
[0031] The one or more processors may be configured to process the received user query to determine whether there is an error in the user query relating to one or more of: a misspelling, an abbreviation, an autocomplete error and a space-related error. This may allow the approach to handle a range of non-correct inputs, for example with character-level mutations to the intended user query, in multilingual settings in addition to the slot filling operation. This can allow for corrections on both natural language in addition to application domain language input (also given in corresponding natural languages).
[0032] The one or more processors may be configured to process the user query in dependence on multiple natural languages and / or multiple application domain languages. The framework can support multilingualism (for example, supporting up to 100 different languages), as well as the application domain language (in respective natural languages).
[0033] The one or more processors may be configured to input the classifications for each slot to a binary logistic regression classifier to modify the notation of the class for each slot. This may improve the accuracy of the slot filling task.
[0034] The slot filling operation may employ a deep learning model implemented by the one or more processors. This may allow the deep learning model to be trained to optimize the performance of the end-to-end user query correction and slot filling operations.
[0035] The deep learning model may be configured to output a classification for each slot of the corrected user query and a corresponding confidence value for each classification. The one or more processors may be configured to update the deep learning model in dependence on the confidence values. This may allow the deep learning model to be updated to improve the performance of the end-to-end user query correction and slot filling tasks.
[0036] The one or more processors may be configured to label respective sets of slots and their corresponding classifications and periodically retrain the deep learning model using the labelled sets of slots and their corresponding classifications. This may allow the model to be continuously improved using active labelling and / or active learning. This may improve the accuracy of the end-to-end user query correction and slot filling task. Model boosting based on active learning and labelling can employ a tuneable model selection approach and use refinement, manual verification and retraining block-steps. This may significantly reduce human intervention and improve the trained model by relying more on a heuristic that employs the model output.
[0037] The received user query and the corrected user query may each be in a string format. This may be an appropriate format for processing the queries. Each part of the corrected user query may be a contiguous word subsequence of the corrected user query. This may allow slots to be identified for each contiguous part of the corrected user query.
[0038] The data processing device may be communicatively connected to an application distribution platform. The corrected user query may be used as a search input to the application distribution platform. This may allow the application distribution platform to search for applications relating to incorrectly-entered user queries and retrieve the intended application from an application library.
[0039] According to another aspect, there is provided a method for performing end-to-end user query correction and slot filling, the method comprising: receiving a user query; processing the user query to determine whether the user query is correct with respect to an application domain language ora natural language; based on the processing of the user query, forming a corrected user query; performing a slot filling operation for the corrected user query to identify one or more slots for the corrected user query, each identified slot corresponding to a part of the corrected user query; and classifying each identified slot of the corrected user query into one of a plurality of classes.
[0040] This method can handle variations of non-correct input in multilingual settings in addition to the slot filling task and can handle application domain language in addition to many different natural languages.
[0041] According to another aspect, there is provided a computer program which, when executed by a computing device, causes the computing device to perform the method described above.
[0042] According to a further aspect, there is provided a data carrier storing in non-transient form the computer program described above.
[0043] BRIEF DESCRIPTION OF THE FIGURES
[0044] The present invention will now be described by way of example with reference to the accompanying drawings.
[0045] In the drawings:
[0046] Figure 1 illustrates examples of mistyping of the input user query “huawei”; Figure 2 illustrates an example of an input query and corresponding slot values and types;
[0047] Figure 3 schematically illustrates an overview of exemplary component blocks of the approach described herein;
[0048] Figure 4 schematically illustrates the input to and the output from the input user query correction block;
[0049] Figure 5 schematically illustrates exemplary sub-blocks of the input user query correction block;
[0050] Figure 6 schematically illustrates an exemplary normalisation block of the input user query correction block;
[0051] Figure 7 schematically illustrates an exemplary pre-encoder block of the input user query correction block;
[0052] Figure 8 schematically illustrates an exemplary post-decoder block of the input user query correction block;
[0053] Figure 9(a) schematically illustrates the use of Levenshtein crosschecking to compare selected predictions with a dictionary of classes;
[0054] Figure 9(b) schematically illustrates the use of Jaro-Winkler crosschecking to compare selected predictions with a dictionary of classes;
[0055] Figure 10 schematically illustrates the operation of an example of an optional hierarchical and discriminative feature representation block used to refine the slot prediction output;
[0056] Figure 11 schematically illustrates an approximate overall structure of one branch of the output of the optional hierarchical and discriminative feature block, showing multiple clusters;
[0057] Figure 12 schematically illustrates how a refined slot filler class output can be obtained from contiguous slot features;
[0058] Figure 13 schematically illustrates the input and output to the B-l-0 refiner block; Figure 14 depicts the structure of the output in a step-like manner to illustrate the in-depth execution of the method;
[0059] Figure 15 schematically illustrates the end-to-end architecture of the B-l-0 refiner block;
[0060] Figures 16 and 17 schematically illustrate the optional use of model boosting using active labelling and active learning approaches;
[0061] Figure 18 shows an example of a data processing method in accordance with embodiments of the present invention;
[0062] Figure 19 schematically illustrates an apparatus for performing the method and some of its associated components;
[0063] Figures 20(a) and 20(b) show results for the approach described herein using different pretrained models: XML-R-XXL in Figure 20(a) and XML-R-XL in Figure 20(b);
[0064] Figures 21(a)-21 (c) show results for other competitive approaches. Figure 21(a) shows results for an OPT-6.7B based model, Figure 21(b) shows results for a CAPSULE-NLU approach, Figure 21 (c) shows results obtained using a STIL method;
[0065] Figure 22 shows a table with performance metrics for various implementations of the method described herein; and
[0066] Figure 23 shows a graph with the f1 score vs training dataset size for a fixed testing set for English language only using the XLM-R-XXL model.
[0067] DETAILED DESCRIPTION
[0068] Embodiments of the present invention relate to end-to-end input query correction and slot filing tasks. The process receives as input a user query and outputs one or more slots and their corresponding slot types, each slot corresponding to a part of the corrected input query. The slot filling method is based on multilingual models combined with machine learning techniques to excel in the slot filling detection task across different data types. The output slots can be classified into one of a plurality of classes and used in a range of applications. In some implementations, a compact hierarchical and discriminative feature representation (HIDIFER) can additionally be used to refine the slot prediction output (i.e. the slot value and its corresponding slot type / class) of each part of the corrected user query. The approach also introduces an optional B-l-0 refiner for the predicted slots based on a double binary logistic regression classifier to further improve the slot prediction performance. The HIDIFER and B- 1-0 refiner blocks can be employed to detect brand new slots and to improve the overall detection process. These aspects will be discussed in greater detail herein. A method for model boosting based on the active learning / labelling with a viable tuneable model selection approach is also described.
[0069] As shown in Figure 3, a user query 301 is input to the input user query correction (IIIQC) block 302. The output of the input user query correction block 302 is a corrected user query, which is then input to a slot filling block 303. The output of the slot filling block 303 can optionally be input to a HIDIFER block 304 and the output of this block 304 can be input to a B-l-0 refiner 305. The final output is a list of slot types and corresponding slot values 306.
[0070] The input and output to the IIIQC block 302 are shown in more detail in Figure 4. The IIIQC block 302 receives as input the input user query 301 outputs a corrected user query 350. The input user query 301 and the corrected user query 350 may each be in a string format. The slot filling block 303 is configured to directly receive the corrected user query output from the user query correction block as input.
[0071] The IIIQC block 302 obtains corrections to input user queries at the natural language level in addition to the application domain level (in corresponding natural languages). The main goal of the IIIQC block is to correct the user query and make it searchable and usable and to send it further down the process to the slot filling block, which will be described in the subsequent section. Therefore, at the IIIQC block, the input user query is verified and corrected.
[0072] The IIIQC block 302 will now be further described with reference to Figure 5. The IIIQC block 302 employs a preprocessing and data augmentation approach. As shown in Figure 5, the IIIQC block 302 comprises a normalization block 310, a pre-encoder block 311 , an encoder block 312, a decoder block 313 and a post-decoder block 314. In the normalisation and preencoder blocks, character-level mutations are modelled in order to mimic spelling and other erroneous input to the encoder. In this example, the encoder and decoder blocks have a transformer architecture (see M. Lewis et al., “BART: Denoising Sequence-to-Sequence Pretraining for Natural Language Generation, Translation and Comprehension”, Proceedings of the 58thAnnual Meeting of the Association for Computational Linguistics, pages 7871-7880, 2020, Association for Computational Linguistics). However, the encoder and decoder blocks may have other suitable architectures. In an exemplary implementation, the encoder-decoder part is based on the BART model (see above reference). The post-decoder block 314 will be described in more detail later with reference to Figure 8. This block takes care of joint predictions and handles application domain data.
[0073] The IUQC block 302 can correct different kind of input distorted text, such as text to be autocompleted, abbreviations, misspellings and space-related issues. These defects are outlined below.
[0074] The user of a device such as a smartphone may expect autocompletion of input text to be performed and may adjust the text that they input to the device accordingly. For example, when the user starts to enter the following, the input may be autocompleted as follows: ‘f’ may be corrected as facebook, ‘in’ may be corrected as instagram, and ‘tikt’ may be corrected as tiktok.
[0075] A user may accidentally input or intentionally misspell inputs in different ways. There may be different manipulations of characters of the user input. For example, an input user query of ‘facbook’ may be corrected as facebook, and an input user query of ‘maseger’ may be corrected as messenger.
[0076] Users may also widely use abbreviations which also might be ambiguous. For example, an input user query of ‘cod’ may be corrected as either call of duty or cape cod times. An input user query of ‘ml’ may be corrected as mobile legends. An input user query of ‘fb’ may be corrected as facebook.
[0077] Many applications use multiple English words together as a name without a space. For example, a user input of ‘baby bus’ may be corrected as babybus for searching an application distribution platform. A user input of ‘ice cream’ may be corrected as icecream with respect to application domain language.
[0078] The pre-processing and data augmentation steps are performed by the IUQC block 302 in order to tackle this problem. The input query correction problem is addressed as a sequence- to-sequence (s2s) problem that corrects a distorted text into its normal form. Also, if typos are considered noises in text, input query correction can be considered as a denoising process that converts spoiled texts into the unspoiled ones. The overall architecture of the input query correction block 302 is shown in Figure 5. The normalization block 310 aims to mimic various different errors in the user input. The errors may be character-level manipulations. This can help to make the model more resilient in the application domain and natural language worlds. The operation of the normalisation block 310 comprises several steps, as shown in Figure 6.
[0079] In step 1 , shown at 601 , different incorrect forms of input user query are generated together with their respective original forms. In addition, aliases of tokens are generated that form the user query if those exist. The alias is dictionary based (e.g. rational logical, verified; honest ethical, moral).
[0080] At step 2, shown at 602, grammar corrections (at singular and plural level both on natural and abstract application language) are performed. Autocomplete may be treated using, for example, a Neuspell decoder previously finetuned on autocomplete queries.
[0081] At the decision step shown at 603, the word order is shuffled only (in this example, if shorter than m=5 tokens) and the process proceeds to step 3, shown at 604. If there are repeated tokens’ root form, one can be kept in its original lemma form. Lemmatization considers the context and converts the word to its meaningful base form, which is called lemma.
[0082] At step 4, shown at 605, extra spaces / special characters are removed from tokens. Special characters are non-query inclusive, and they are combined based on versatility of the input user query. Also based on given input languages, a union of special characters is applied.
[0083] At the step shown at 606, digits are removed from within the token (regarded as a typo) if larger than p=1 token. Since typos can (re)appear after removing extra spaces and special characters the process can be repeated until error rate falls beyond some reliable thresholds.
[0084] The output at step 5, shown at 607, creates “clean slate” with respect to the data entry input at step 1 , 601. The steps may be applied in the order shown in Figure 6 to train the normalisation block. This block 310 may be skipped in the inference stage.
[0085] In order to make the model learn profoundly, but also to eliminate over-confusing inputs, the pre-encoder block 311 may not receive all of the possible character-level manipulations of the input user query at one time (i.e. to happen simultaneously).
[0086] In the preferred implementation, manipulations of the input user query may be taken in pairs, with two manipulations taken simultaneously by the pre-encoder block 311 . Figure 7 shows an exemplary illustration of how pairs of character-level manipulations of a user query (i.e. incorrect user queries) 750 may be selected by the pre-encoder block 311 for input to the encoder block 312.
[0087] At each time step, a different combination of manipulations can be chosen based on the multiplex logic. In Figure 7, Auto 701 , Miss 702, Abbr 703 and Spac 704 denote autocomplete, misspelling, abbreviations and space-related noise functions respectively.
[0088] L 720 denotes the following logic of the selector 730. Once a pair of manipulations is chosen, its similarity score with respect the ground truth is obtained. The ranking from the most to the least similar pair of manipulations is taken. In case of equally similar pairs, the one that has the lower word mover’s distance (WMD) is chosen to proceed with first. This method has been found to help the model learns better compared to the user of a random or partial selection procedure. This block is skipped in the inference stage.
[0089] The encoder and decoder blocks 312 and 313 may be based on the BART model. This is a denoising word-level seq2seq autoencoder pretrained for natural language comprehension, generation, translation, etc. Thus, the spelling correction is observed and employed as a character-level seq2seq denoising autoencoder problem. Pretraining data with various character-level mutations may be used, previously created, in order to “mimic” spelling errors. To train the model, different noise functions combined or used on their own may be created in order to generate common errors that frequently occur in application distribution platforms.
[0090] The post-decoder block 313 is the final block of the input user query correct block 302 described earlier with reference to Figure 5. The components of the post-encoder block 313 are shown in Figure 8. The post-decoder block 313 comprises an output regulator 801 and an applications integrator 802. The post-decider block 313 outputs the corrected user query 350.
[0091] The outp.ut regulator 801 regulates the output of the decoder block by applying a verification stage for the tokenized output bounded by <beginning> and <end> tokens. This firstly involves receiving predictions including respective probability scores associated with them. In one implementation, three outputs (predictions) of the decoder 312 are observed togetherwith their respective probability scores.
[0092] The difference between the first and the second decoder outputs (diff12), as well as the difference between the second and the third decoder outputs (diff23) probability value predictions are taken and it can be learned for which of those the first predicted decoder output is the correct one (i.e. when the difference between the first and the second is reliably high, i.e. beyond some threshold T1). The same is then performed for the second ranked prediction: the threshold T2 for which the second ranked is the correct one is obtained. Introducing multiplication and / or addition functions under some conditions can improve precision even more (or at least average precision). Since T1 is more important than T2, the following logic can be employed:
[0093] 1) Take the first ranked output if diffl 2 >= T1
[0094] 2) If not, take the second output if diff23 >= T2
[0095] 3) If neither of 1) and 2) above are true, take the original first ranked predicted output.
[0096] Since application names, developer names, characters and others represent indigenous names in the application space, those usually differ significantly from the natural language. In the apps integrator module 802, the following method can be employed.
[0097] There may be collected a dictionary of all application names, developer names, characters and others in its inflected (lemmatic) form. The tokenized output is taken and those tokens are crosschecked with the dictionary forms.
[0098] Crosschecking may be performed by employing a Levenshtein distance. This may include some enhancements (for example, it may employ transpositions as well as insertions, deletions and substitutions). The Levenshtein distance may be employed for strings longer than, for example, 15 characters. For strings shorter than, for example, 15 characters, the Jaro-Winkler distance may be employed. In the case of the Levenshtein distance, each substring in the sequence may be allowed to be edited up to 4 times including “the same character” scenario.
[0099] Instead of using inherited rules for substring matching, the approach may employ the Longest Common Subsequence (LCSS) approach. Furthermore, the approach may allow swapping of characters (symbols) that are placed next to each other. Therefore, the most appropriate word forms may be obtained from the dictionaries mentioned previously. This supports inter- and intra-multilinguality.
[0100] Two typical examples can be seen in Figures 9(a) and 9(b). In Figure 9(a), for the input user query ‘FoOx Sprt Wastch Liivee’ 901 , the selected prediction is compared with the dictionary by determining a Levenshtein distance between the selected prediction and one or more classes in the dictionary, shown at 902. The output is corrected to ‘FOX Sports Watch Live’, shown at 903.
[0101] In Figure 9(b), for the input user query ‘Color bok’ 904, the selected prediction is compared with the dictionary by determining a Jaro-Winkler distance between the selected prediction and one or more classes in the dictionary, shown at 905. The output is corrected to ‘Color book’, shown at 906.
[0102] The output of the IIIQC block 302, the corrected user query 350, is then directly input to the slot filling block 303.
[0103] An exemplary operation of the slot filling block will now be described. In this example, the XLM- RoBERTa-XXL model (see N. Goyal et al., “Larger-Scale Transformers for Multilingual Masked Language Modeling”, arXiv:2105.00572 [cs.CL], 2021) is used. The XLM-RoBERTa-XXL model is based on a multi-layer bidirectional Transformer encoder. XLM-RoBERTa-XXL is pretrained on 2.5TB of filtered CommonCrawl data containing 100 languages in a self-supervised fashion with the Masked language modeling (MLM) objective. Taking a sentence, the model randomly masks 15% of the words in the input then run the entire masked sentence through the model and predicts the masked words.
[0104] In this implementation, those models are extended to jointly support intent recognition and slot filling tasks. The input token sequence is given by x = (xi, . . ., XT). The output of MLM is H = (hi, . . ., h-r). Based on the hidden state of the first special token ([CLS]), denoted hi, the intent is predicted by applying a softmax function as below: yl= softmax(\Nlh + bl) where W‘ is a weight matrix, blis a bias, is the output value and h±is the hidden state vector. For slot filling, the final hidden states of other tokens h2, . . ., hTare fed into a softmax layer to perform classification on the slot filling labels according to the following: where hndenotes the hidden state corresponding to the first sub-token of word xn. To jointly model intent classification and slot filling, the following is used:
[0105] The learning objective is to maximize the conditional probability p(y‘,ys|x).
[0106] In the present approach, conditional random fields (CRF) can be used. The CRF can be applied on top of the XLM-RoBERTa-XXL model. The CRF can be used with a dictionary of prominent queries in addition to viable post processing and adaptive filtering, which together with multilingual models such as XLM-RoBERTa-XXL is employed to detect brand new slots (as opposed to other competitive approaches, which do not detect new slots).
[0107] Using features and their corresponding weights and a dependency graph, the weights functions can be parameterized by employing a set of probabilistic rules to structure how the newly derived (intermediate) labels depend on the observed labels. Instead of using learnt weights in a classical manner, where the first order of Markov assumption has been employed, a two-step method can be employed that firstly associates feature weights to (0,1) and [1 , +oO) intervals in case of similarity to O-class and non-0 class label respectively. In the second step, an expectation-maximization (EM) algorithm is applied to find the parameters’ estimation of the maximum likelihood of the model given a set of observed feature vectors for the [1 , +oO) interval for all the inclusive classes. An additional variable z can be put between m and n to avoid treating P(m|n) as a traditional CRF would have, giving the following:
[0108] P(m |n)=zP(m\z,ri)P(z\ri)
[0109] A method which deals with prominent queries and tokens can optionally be added to further boost of the performance of the method. This heuristic takes advantage of prevalent class(es) (and is related to application name or content), meaning that if a testing query contains at least one token which is not present in the training set (disregarding stop words) then that query is predicted as the prevalent one. The intuition behind this is that the prevalent class contributes notably to the performance increments, while it minorly spoils performance in case of the wrong predictions. The whole query does not need to be observed (at once). The queries used may consist of up to 70% of original tokens (at the same time). Therefore, this block not only detects new slots, but also improves the overall performance of the slot filling task. As mentioned above, a HIDIFER can optionally be employed, which has been found to enhance performance of the slot filling task. A fine-grained feature representation can be created for the slots, by applying hierarchical k-means approach repeatedly, and by deriving heuristics to fix and employ a cluster centers structure to obtain predictions. By using this approach, new slots can be correctly detected and an improved slot filling performance can be obtained.
[0110] Figure 10 illustrates an example of this process. The features 1001 can be chosen and partitioned into different clusters based on their slot relevance, for example their distance in feature space. To achieve this, hierarchical k-means clustering can be applied by repeatedly applying k-means clustering until there are single element clusters. These are taken as the final features. Each feature is related to the slot type it belongs to.
[0111] To speed up the overall process, a two-branch hierarchical k-means trees (O and non-0 i.e. for 16 non-0 classes) can be used, as shown in Figure 10. Thus, a fine-grained feature representation is created for the slots that manages to correctly detect new slots and also obtains better slot filling performance.
[0112] In some cases, this method may only be applied for lower confidence slots, for example ones whose confidence is below T = 0.83.
[0113] As shown in Figure 10, the features 1001 can be split into two groups 1002 and 1003 based on O and non-0 classes respectively. This may halve the search time. For each group, a hierarchical k-means tree 1006 and 1007 respectively can be created by clustering the feature vectors using the extended k-means algorithm repeatedly. This is shown in Figure 10 for the O-features at 1004 and the non-0 features at 1005. This partitioning of II features into k disjoint subsets Sj each containing llj features, minimizes the sum-of-squares criterion, where luis a vector representing the uthdata point and rrij is the geometric centroid of all data points in Sj:
[0114] To properly fix the cluster centers, the following method may be employed. This approach employs a re-estimation procedure. Initially, the features are assigned at random to the sets. For step 1 , the centroid is computed for each set. In step 2, every feature is assigned to the cluster whose centroid is closest to that feature. Then these two steps are used one after another until a stopping criterion is met (based on threshold obtained or sometimes based on number of iterations reached). Stabilizing cluster centers is not a trivial task, especially not in a high-dimensional feature space. Moreover, their number creates an additional challenge due to more complicated “separability” using somewhat similar slot values. Thus, in order to make the trees more robust, the following procedure can be applied.
[0115] For the first three levels of the hierarchical tree k cluster, centers are found by calculating the weighted linear mean (WLM) of several previously created cluster centers (in repeated tests). The weights can be calculated using an inverse asymptotic rule (i.e. reciprocal distance to the clusters). Thus, an emphasis is put on the first level clusters, then the second level clusters, third level and so on. In other words, for i iterations there are i cluster centers vectors, each of length k. Then the weighted linear mean (WLM) for each dimension is calculated resulting in a vector of length i representing k cluster centers. For the higher tree levels this process may not be necessary, as cluster centers may be already properly fixed. This approach only uses linear memory, O(k + II), in the number of cluster centers k and feature points II. The approximate overall structure of one branch, showing multiple clusters 1101 , 1102, 1103, 1104, 1105 and 1106, is schematically illustrated in Figure 11 .
[0116] In each hierarchical k-means tree part, an algorithm assigns “votes” and sorts features based on the number of received votes. The tree coefficients may be based on weighted normated functions of the position of the cluster centers. Each “path” may have different coefficients based on the calculated clusters of the tree. A filtering scheme was found to significantly increase the performance, as this kind of tree allows for better differentiation between the features.
[0117] Voted cluster center locations (CCLs) of the feature vectors are ranked from the most likely to the least likely. A verification stage can be employed by using the double-directional matching to reorder the top five previously ranked clusters.
[0118] As shown in Figure 12, first, the ranking obtained for the feature vectors 1201 is weighted by normalized double-directional matching location scores, shown at 1202 and 1203, and again normalized, thus associating normalized votes with each CCL. A confidence is assigned for each CCL and defined as the ratio of normalized votes associated with that CCL and total number of (normalized) votes. After each matching, a reranking approach, shown at 1204 and 1205, is employed based on number of the votes each class “receives”. Finally, for each contiguous part of user query, the cluster element whose feature element has a class associated to it is obtained. Thus, this class determines the final class of the contiguous part of the query in B-l-0 notation (see e.g. https: / / en.wikipedia.org / wiki / lnside-outside- beginning_(tagging), where B = beginning, I = inside and O = outside, e.g. B-app_name or I- app_content), as shown at 1206 in Figure 12.
[0119] As mentioned above, the approach may optionally use a refiner for the predicted slots based on a double binary logistic regression classifier to further improve the slot prediction performance.
[0120] A B-l-0 refiner may utilize generic features, such as user query length, the number of tokens for the user query and their order of appearance in the user query, either query or filler being placed firstly. This set of features is based on the segmentation of the user query according to the position of tokens, resulting for most user queries in three segments i.e. token being before, between or after slot filler. Then, two binary logistic regression classifiers are appended to the end where they can refine wrongly classified B-l-0 predictions. The classifiers may be trained on data features that already contain ground-truth for B-l-0 that exactly match its tokens. Thus, the approach aims at correcting the previous block’s output and performance of the B-l-0 extractor i.e. the performance of the slot filling task. This means to modify and adapt the notation of the class that refers either to B, I or O notation slots (as opposed to other competitive approaches) without changing the inflicting base slot filling class (e.g. app_name or app_content).
[0121] The architecture diagram is shown schematically in Figure 13. Features 1301 are input to the B-l-0 refiner 305. The output of the B-l-0 refiner is a refined B-l-0 output 1303.
[0122] In Figure 14, the structure of the output is depicted in a step-like manner to illustrate the in- depth execution of this particular method. For example, for the input use query “omegle chat - live video chat”, the input query eventually comprises several contiguous parts of the user query which could be independently identified as following: a) windows lite: B-app_name l-app_name b) - : O c) live: B-app_features d) audio chat: O O
[0123] Since “audio chat” and do not belong to any of the predefined classes, they are eventually classified as outside classes. For “live”, this belongs to the app_features class, as it describes the app qualities i.e. features of the app itself - thus it is associated to app_features class, “windows lite” refers to the name of the app and thus denote app_name. It consists of two tokens, the first one, “windows”, being B-app_name and “lite” being l-app_name, where B and I denote beginning and inside respectively. Figure 14 shows some additional examples for the input user queries “seamloc” and “last version”.
[0124] The probability p(X) is modelled using a function that gives outputs between 0 and 1 for all values of X, where:
[0125] This represents the probability of having the correct B-l-0 notation introduced previously (i.e. a class can either be B-class_name, l-class_name or output (O)). This notation is modelled using this probabilistic function.
[0126] To fit the model in the above equation, maximum likelihood estimation is used:
[0127] Taking the log of this equation, in the logistic regression model, an increase of X by one changes the log odds by 01, or equivalently it multiplies the odds by .
[0128] 0o and 0- are estimated using the training data and are chosen to maximize the likelihood function. Based on the data, the exact functional form of the cumulative density function (CDF) of a variable whose value is represented by X is not known. Therefore, the CDF can be approximated using the empirical cumulative density function (ECDF) by applying a discrete sampling. Another approximation is performed by using modifying chain rule to include both 0 and 1 values simultaneously. In order to evaluate ECDF, one can evaluate the function at a series of points. This suggests an approximation in which instead of calling the ECDF function many times, one can employ the cumulative sum function (CUSIIM) by modifying the trapezoidal function to do a sum of the sliced areas that approximate the normalized area under the samples. In other words, since the CUSIIM is exactly zero for x < 0, one can add the CUSIIM value at 0 to the values returned by the trapezoidal function. where (p( ))"is the ECDF at a value X, n is the number of samples, and x_i is the value at the j-th sample. By applying the Newton-Raphson method for approximation, v and < can be obtained. Figure 15 schematically illustrates an exemplary end-to-end architecture of the B-l-0 refiner block 305. The input features 1301 are input to a first binary logistic regression classifier 1501 where they are classified into O and non-0 features 1502. The features 1502 are then input to a second binary logistic regression classifier 1502 where the non-0 features are further classified into l-class and B-class features to give the B-l-0 refined output 1303.
[0129] The two-step binary logistic regression classifier (i.e. B-l-0 refiner) therefore corrects B-l-0 prediction types, leaving class’ values unchanged. As described above, in the first step, it is determined whether a feature is O or non-O. Then, if it is non-O, the non-0 can be classified as either B-class or l-class. This means to modify and adapt the notation of the class that refers either to B, I or O notation slots, as opposed to other competitive approaches. The approach can further improve the performance of the slot filling task.
[0130] Optionally, the models described herein can be boosted using a combination of active labelling (ALB) and active learning (ALR), with a tuneable model selection approach. This may further improve the performance of the slot filling task. Current slot filling algorithms’ performance heavily relies and directly depends on rich and abundant sources of data. The main trait of ALR is speed rather than performance, i.e., its ability to significantly reduce human intervention by relying more on a heuristic that includes the model output, as for some user queries one does not need additional manual verification. Thus, data creation is obtained at much lower time cost. In addition to random sampling, which is present in the annotation process, the speed boost up came from current algorithms’ output in addition to a heuristic eventually supported by a manual verification process. Thus, the annotation speed was increased from 23.2 data samples per hour to 68.2 data samples per hour (on average).
[0131] This approach is a combination of ALB and ALR and is used to further boost the slot filling model. Its architecture is schematically illustrated in Figure 16.
[0132] The user query 1600 is input to the model at 1601 . The output of the end-to-end user query collection and slot filling model is obtained in addition to corresponding confidence value scores.
[0133] In the next step shown at 1602, the confidence scores are checked; mainly the absolute difference between the first and the second ranked output (from the whole ranked list). In the case that this difference is reliably high (for example, the difference between the values exceeds a threshold), the process proceeds on to a manual verification block 1604. In the case that the difference is not reliably high (for example, the difference between the values does not exceed a threshold), the model can be refined, as shown at block 1603.
[0134] In this block 1603, it is checked whether the slot value output from the model 1601 contains a prominent (based on frequency of occurrence) token (occurring once) that relates to a class such a “app name”, “app content”, “developer name”, “game type”, “app features”, etc. If so, that slot type is output. If more than one class occurs, the following actions can be taken: i. It is checked whether the first ranked or the second ranked belongs to the “app name” class: if only one, the rank of that one can be taken as the correct one. If none of them, the first ranked one can be taken as the correct one. ii. Then if any of the combination includes either the “app name” or “developer name” classes (with excluding “game type”), that corresponding slot type is output. If “developer name” and “game type” occur together, “game type” can be taken as the final slot type prediction. iii. In the case that there is a combination of “app content” and “game type” classes, one of the “game type” classes is taken as the correct one. Similarly, for the combination of “app features” and “game type”, precedence of “game type” is chosen here also. iv. In case there is a combination of “app content” and “app features” or “app content” or some different class (“time version” or “other”), the “app name” class is chosen that most of “app content” relate to. v. Lastly in case of any combination of “app features”, “other” and “time version” lasses, similarly to the earlier step, the class “app features” is taken. For “other” and “time version”, “app name” is chosen that most of “other” relate to. “character” is observed in the same manner as “app content” class
[0135] The process then proceeds to the manual verification block 1604. After the manual verification block 1604, the model can be retrained at 1605 (if needed).
[0136] The output of the process, show at 1606, is a set of slot type and slot value elements. Active labelling has been found to introduce a significant improvement in speed during the dataset annotation process i.e. , an ability to reduce human intervention by relying more on its output, and relationships between the classes on interest.
[0137] Active learning and labelling can therefore be used to proactively strengthen the trained models by continuously checking their performances and discarding those that perform more poorly than the best current model. Thus, the retraining process can be repeated indefinitely in theory, depending on thresholds that are set in the quality assessment checker. Practically, once the performance does not increase the retraining process can be stopped.
[0138] The whole process is depicted in Figure 17. Input user query data 1701 is input to a reduced training model 1702. The previous model parameters 1703 are retained. At 1704, it is evaluated whether the model 1702 is better than the model with the previous parameters 1703. If yes, the model 1702 is used as a newly created model 1705 and is used to produce predictions 1706. At 1707, a quality assessment is performed as described above. If the quality is determined to be sufficient, the result is output as the final output. If not, the model parameters and hyperparameters are adjusted at 1708 and used as the parameters of the reduced training model 1702 in the next iteration of the process. In this example, the quality assessment is in terms of statistical significance based on the confidence scores. The statistical significance is measured based on a p-value (p=0.05) and if within those boundaries the quality is considered sufficient. The p-value is the probability of obtaining results at least as extreme as the observed results of a statistical hypothesis test, assuming that the null hypothesis is correct. The p-value serves as an alternative to rejection points to provide the smallest level of significance at which the null hypothesis would be rejected. A smaller p-value means that there is stronger evidence in favour of the alternative hypothesis.
[0139] The calculations are based on the assumed or known probability distribution of the specific statistic tested. P-values are calculated from the deviation between the observed value and a chosen reference value, given the probability distribution of the statistic, with a greater difference between the two values corresponding to a lower p-value. Mathematically, the p- value may be calculated using integral calculus from the area under the probability distribution curve for all values of statistics that are at least as far from the reference value as the observed value is, relative to the total area under the probability distribution curve.
[0140] The null hypothesis, also known as the conjecture, states that the predicted values are given in specific order and with particular class values for a particular user input. The alternative hypothesis states whether the predicted values differ from the value of the “population parameter” stated in the conjecture. For different inputs, different null hypotheses are formed. There can also be corresponding alternative hypotheses. For this specific value of p, parameters and hyperparameters can be optimized using a grid search engine for which 90% of the null-hypotheses were accepted (which was found in experiments to provide a good quality amount of the new data samples obtained at faster speed). The output of the overall process is contiguous (mutually non-overlapping and exclusive) parts of the corrected user query. For each of the parts, its corresponding class value (i.e. slot type, one out of K classes) is also obtained.
[0141] Figure 18 shows an exemplary method for performing end-to-end user query correction and slot filling. At step 1801 , the method comprises receiving a user query. At step 1802, the method comprises processing the user query to determine whether the user query is correct with respect to an application domain language or a natural language. At step 1803, the method comprises, based on the processing of the user query, forming a corrected user query. At step 1804, the method comprises performing a slot filling operation for the corrected user query to identify one or more slots for the corrected user query, each identified slot corresponding to a part of the corrected user query. At step 1805, the method comprises classifying each identified slot of the corrected user query into one of a plurality of classes.
[0142] Figure 19 shows an example of a data processing apparatus 1900 configured to implement the method described herein. The apparatus comprises a device 1901. The device 1901 comprises a processor 1902 and a memory 1903.
[0143] The device 1901 may also comprise a transceiver 1904 is capable of communicating over a network with other entities 1910, 1911. Those entities may be physically remote from the device 1901. The network may be a publicly accessible network such as the internet. The entities 1910, 1911 may be based in the cloud. Entity 1910 is a computing entity. Entity 1911 is a command and control entity. These entities are logical entities. In practice they may each be provided by one or more physical devices such as servers and data stores, and the functions of two or more of the entities may be provided by a single physical device. Each physical device implementing an entity comprises a processor and a memory. The devices may also comprise a transceiver for transmitting and receiving data to and from the transceiver 1904 of device 1901. The memory stores in a non-transient way code that is executable by the processor to implement the respective entity in the manner described herein.
[0144] Therefore, the method may be deployed in multiple ways, for example in the cloud, on the device, or alternatively in dedicated hardware. As indicated above, the cloud facility could perform training to develop new algorithms or refine existing ones. Depending on the compute capability near to the data corpus, the training could either be undertaken close to the source data, or could be undertaken in the cloud, e.g. using an inference engine. The method may also be implemented at the device, in a dedicated piece of hardware, or in the cloud. The tables in Figures 20(a) and 20(b) show results for the approach described herein using different pretrained models: XML-R-XXL in Figure 20(a) and XML-R-XL in Figure 20(b). The model was tested on different application domain language datasets (denoted as lang), here in English (EN), German (DE), French (FR), Italian (IT) and Spanish (ES) languages. The following evaluation metrics are used for the results: precision (prec), recall and flscore.
[0145] In Figures 21 (a)-(c), results are shown for some competitive approaches such as an OPT-6.7B based model (“OPT: Open Pre-trained Transformer Language Models”, S. Zhang et al., https: / / arxiv.org / abs / 2205.01068), the CAPSULE-NLU approach ("Joint Slot Filling and Intent Detection via Capsule Neural Networks", C. Zhang et al., Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5259-5267, Florence, Italy, Association for Computational Linguistics), and the STIL method ("STIL - Simultaneous Slot Filling, Translation, Intent Classification, and Language Identification: Initial Results using mBART on MultiATIS++“, J. Fitzgerald et al., Proceedings of the 1st Conference of the Asia- Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 576-581 , Suzhou, China, Association for Computational Linguistics).
[0146] Regarding the table in Figure 22, the first row exemplifies the performance repeatability for English data when re-training the same model five times for 75 epochs from scratch using the same English training and testing splits. An average [mean + / - SD] metric is used.
[0147] The second row shows the 5-fold cross-validation for English data (4 / 5 data for training and 1 / 5 for testing that are different in each 5 folds). The model was trained for 75 epochs form scratch each time. The average [mean + / - SD] metric is used.
[0148] For another English dataset subset training / testing split of “clean” data was fixed and the obtained precision / recall / f1 scores were 92.01 / 92.89 / 92.45% respectively. The training size was 1400 data samples. If the training set is enriched with 300 "noisy" (misspelled) data samples and the model is retrained, poorer performance is obtained when tested on the same testing split, as shown in the third row of Figure 22.
[0149] In the fourth row, the same conditions as for the third row were kept, but instead of “noisy”, 300 “clean" data samples were added and the model was retrained. The performance was then obtained using the same testing dataset and is improved relative to the noisy dataset. The graph in Figure 23 shows the f1 score in % on the y-axis vs training dataset size on the x- axis for a fixed testing set for English language only using XLM-R-XXL.
[0150] The approach described herein for end-to-end input user query correction and slot filling can handle more variations of non-correct input (misspellings, abbreviations, autocomplete, space- related issues) in multilingual setting in addition to slot filling task. The approach can advantageously handle application domain language in addition to many different natural languages. Different misspellings can be handled and the approach can be used for inputs to multilingual application distribution platforms, such as but not included to AppGallery. Since the above approach has been achieved in an end-to-end manner, this approach can obtain predictions for different kinds of domains and natural language data.
[0151] The input user query correction block comprising a normalization block, pre-encoder block and post-decoder block in addition to the encoder-decoder can allow for corrections on both natural language in addition to application domain language input (also given in corresponding natural languages).
[0152] The slot filling can use an approach based on adapting and extending CRFs. The block can advantageously detect brand new slots from application domain language corrected inputs.
[0153] The HIDIFER block further improves the slot filling performance and can detect brand new slots (for example, at the application domain language level).
[0154] The B-l-0 refiner block can correct the notation of the class that refers either to B, I or O class type by applying a more sophisticating refiner based on approximated discrete sampling. It can modify and improve the notation of the class that refers either to B, I or O notation slot (as opposed to other competitive approaches). The approach also aims at improving the slot filling task performance.
[0155] Model boosting based on the active learning and labelling employs a viable tuneable model selection approach and uses refinement, manual verification and retraining block-steps. It has the ability to significantly reduce human intervention and improve the trained model by relying more on a heuristic that employs the model output, as for some user queries one does not need additional manual verification.
[0156] The system is capable of obtaining slot predictions from an input user query previously corrected from misspellings, autocomplete shortcuts, abbreviations, and space related issues. This can enable slot filling supported by multiple query corrections for intent recognition applications and has not only improved dedicated independent components but also provides a functional end-to-end system.
[0157] The slot filling approach can detect brand new slots in addition to the standard ones and also to refine the predictions using a dedicated approach. Model performance boosting has been additionally obtained by using active learning and labelling with a viable tuneable model selection approach.
[0158] The framework can support multilingualism (for example, supporting up to 100 different languages).
[0159] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.
Claims
CLAIMS1. A data processing device (1901) for performing end-to-end user query correction and slot filling, the data processing device having one or more processors (1902) configured to: receive (1801) a user query (301); process (1802) the user query (301) to determine whether the user query is correct with respect to an application domain language or a natural language; based on the processing of the user query, form (1803) a corrected user query (350); perform (1804) a slot filling operation for the corrected user query to identify one or more slots (306) for the corrected user query, each identified slot corresponding to a part of the corrected user query; and classify (1805) each identified slot of the corrected user query into one of a plurality of classes.
2. The data processing device (1901) as claimed in claim 1 , wherein the one or more processors are configured to implement a user query correction block (302) and a slot filling block (303) end-to-end, wherein the user query correction block (302) is configured to output the corrected user query (350) and wherein the slot filling block (303) is configured to directly receive the corrected user query (350) output from the user query correction block (302) as input.
3. The data processing device (1901) as claimed in claim 2, wherein the user query correction block (302) comprises a normalization block (310), a pre-encoder block (311), an encoder block (312), a decoder block (313) and a post-decoder block (314).
4. The data processing device (1901) as claimed in claim 3, wherein the normalization block(310) is trainable to model character-level manipulations of the user query (301) to mimic the input of an erroneous user query to the encoder block.
5. The data processing device (1901) as claimed in claim 4, wherein the pre-encoder block(311) is trainable to select an input to the encoder block (312) by: receiving pairs of character-level manipulations of the user query (301) from the normalization block; and selecting one of each pair of character-level manipulations of the user query (301) for input to the encoder block (312) by determining a similarity between each of the pair of character-level manipulations of the user query and a corresponding ground truth user query.
6. The data processing device (1901) as claimed in any of claims 3 to 5, wherein the encoder block (312) and the decoder block (313) have transformer architectures.
7. The data processing device (1901) as claimed in any of claims 3 to 6, wherein the postdecoder block (314) is configured to: receive multiple predictions and respective associated probability scores output by the decoder block (313); select a prediction of the multiple predictions; compare the selected prediction with a dictionary of classes and corresponding prominent tokens; and output a result from the dictionary as the corrected user query (350).
8. The data processing device (1901) as claimed in claim 7, wherein the one or more processors (1902) are configured to compare the selected prediction with the dictionary by determining a Levenshtein distance ora Jaro-Winkler distance between the selected prediction and one or more classes in the dictionary.
9. The data processing device (1901) as claimed in any of claims 2 to 8, wherein the slot filling block (303) is configured to implement a multi-layer directional transformer encoder.
10. The data processing device (1901) as claimed in any preceding claim, wherein the one or more processors (1902) are configured to perform the slot filling operation by implementing a probabilistic graphical model to detect one or more additional slots in the corrected user query (350) and / or to adjust one or more of the identified slots in the correct user query (350).
11. The data processing device (1901) as claimed in any preceding claim, wherein the one or more processors (1902) are configured to form a feature representation for each of the identified slots.
12. The data processing device (1901) as claimed in any preceding claim, wherein the one or more processors (1902) are configured to tag each part of the corrected user query (350) with a slot label.
13. The data processing device (1901) as claimed in any preceding claim, wherein the one or more processors (1902) are configured to process the received user query (301) to determine whether there is an error in the user query relating to one or more of: a misspelling, an abbreviation, an autocomplete error and a space-related error.
14. The data processing device (1901) as claimed in any preceding claim, wherein the one or more processors (1902) are configured to process the user query (301) in dependence on multiple natural languages and / or multiple application domain languages.
15. The data processing device (1901) as claimed in any preceding claim, wherein the one or more processors (1902) are further configured to input the classifications for each slot to a binary logistic regression classifier to modify the notation of the class for each slot.
16. The data processing device (1901) as claimed in any preceding claim, wherein the slot filling operation employs a deep learning model implemented by the one or more processors (1902).
17. The data processing device (1901) as claimed in claim 16, wherein the deep learning model is configured to output a classification for each slot of the corrected user query (350) and a corresponding confidence value for each classification and wherein the one or more processors (1902) are configured to update the deep learning model in dependence on the confidence values.
18. The data processing device (1901) as claimed in claim 17, wherein the one or more processors (1902) are configured to label respective sets of slots and their corresponding classifications and periodically retrain the deep learning model using the labelled sets of slots and their corresponding classifications.
19. The data processing device (1901) as claimed in any preceding claim, wherein the received user query (301) and the corrected user query (350) are each in a string format.
20. The data processing device (1901) as claimed in any preceding claim, wherein each part of the corrected user query (350) is a contiguous word subsequence of the corrected user query.
21. The data processing device (1901) as claimed in any preceding claim, wherein the data processing device is communicatively connected to an application distribution platform and wherein the corrected user query (350) is used as a search input to the application distribution platform.
22. A method (1800) for performing end-to-end user query correction and slot filling, the method comprising:receiving (1801) a user query (301); processing (1802) the user query (301) to determine whether the user query (301) is correct with respect to an application domain language or a natural language; based on the processing of the user query (301), forming (1803) a corrected user query (350); performing (1804) a slot filling operation for the corrected user query (350) to identify one or more slots (306) for the corrected user query (350), each identified slot corresponding to a part of the corrected user query (350); and classifying (1805) each identified slot of the corrected user query (350) into one of a plurality of classes.
23. A computer program which, when executed by a computing device (1901), causes the computing device to perform the method (1800) of claim 22.
Citation Information
Patent Citations
Misspelling correction based on deep learning architecture
US11017167B1