Create an intent recognition model from the proximity of randomized intent vectors

By generating and verifying candidate intent vectors and using the LSTM model to process noise vectors in parallel, the existing intent recognition model is solved in the creation of difficult problems when dealing with natural language variants, achieving more accurate and fast intent recognition model creation.

CN113597602BActive Publication Date: 2025-05-23INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080021344.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-22
Filing Date
2020-04-02
Publication Date
2025-05-23
Estimated Expiration
2040-04-02

AI Technical Summary

Technical Problem

Existing intention recognition models have difficulty creating when dealing with variations of natural language, requiring a large number of high-quality sentence/phrase input changes to train the model, but in practice it is difficult to obtain these inputs, resulting in the model being inaccurate enough when identifying user intentions.

Method used

By generating candidate intent vectors and verifying them, a vector semantically similar to the input intent vector is selected as the effective intent vector, and a long short-term memory (LSTM) model is used to process it in parallel with the noise vector to improve the accuracy of the intent recognition model.

Benefits of technology

This method can automatically deduce a robust intent recognition model with increased intent inference accuracy while relying on less sentence/phrase input, improving the speed and accuracy of the model creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113597602B_ABST
    Figure CN113597602B_ABST
Patent Text Reader

Abstract

A plurality of candidate intent vectors are generated from the input intent vector. A verification of the plurality of candidate intent vectors is performed, the verification selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The field of the present invention relates to cognitive intent recognition applications, such as chatbots, which utilize intent recognition models to infer the intent of a user who utters a sentence or command. For example, a user may ask "Where can I find coffee?" and the cognitive intent recognition application attempts to infer from the intent recognition model that the user wants to get directions to a coffee shop.

[0002] However, due to the complexity of natural language (NL) variants that exist in almost every different language and across different regions where the language is used to express ideas and intentions, the conventional creation of intent recognition models is problematic. For example, some traditional methods generate "semantically similar sentences" by determining "a set of corresponding semantically similar words" for "each word in the input utterance". However, for any given word, there may be multiple such semantically similar words (e.g., synonyms), and quite different sentences or phrases may be used to express similar meanings. The result of these types of language complexities is a large number of permutations of word and phrase combinations that can be used to express ideas. As a further result, a large amount of training data in the form of high-quality sentence / phrase input variations is required to accurately train traditional intent recognition models. Therefore, the creation of traditional intent recognition models is a manual intensive process of collecting and filtering input variations from a variety of different input sources. Then, language experts are used to generate reasonable and meaningful sentence / phrase variations using synonyms as seed text for intent recognition model generation and training. The amount of data provided and the phrase variations generated determine the quality of the intent recognition model. However, few traditional intent recognition models are prepared using the necessary number of high-quality sentence / phrase input variations. As a result, traditional intent recognition models are poorly created and poorly trained, so they are unable to effectively correctly identify users' intent across all natural language variants and sentence formats to express thoughts and intentions in any given language.

[0003] Long Short-Term Memory (LSTM) is an artificial recurrent neural network (RNN) architecture used in computer-based deep learning processing. In addition to processing single data points, LSTM also has feedback connections that allow processing of data sequences (such as speech or video).

[0004] Sentence vectors are formed using word embeddings to represent words. Sentence vectors are constructed by using a neural network that recursively combines word embeddings in a generative model, such as a recurrent / recurrent neural network or using some other non-neural network algorithm (e.g., doc2vec). Sentence vectors generally have a similar shape compared to word embeddings.

[0005] Gaussian noise is a form of statistical noise that has a probability density function (PDF) equal to a normal distribution, alternatively referred to as a Gaussian distribution. Gaussian noise is modeled using random variables in areas such as telecommunications and computer networks to simulate the effects of noise from natural sources that affect telecommunications and computer network systems. These natural sources that affect those telecommunications and computer network systems include thermal vibrations of atoms in conductors and radiation from celestial bodies (e.g., from the Earth, from the Sun, etc.). Summary of the invention

[0006] A computer-implemented method includes generating a set of candidate intent vectors from an input intent vector. The computer-implemented method includes performing a validation of the set of candidate intent vectors, the validation selecting any one of the set of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector. This embodiment has the advantage of improving intent recognition model creation and validation by relying on less sentence / phrase input and by autonomously deriving a robust intent recognition model with increased intent inference accuracy.

[0007] A system for performing the computer-implemented method and a computer program product for causing a computer to perform the computer-implemented method are also described.

[0008] The licensed embodiment involves generating a set of candidate intent vectors by iteratively processing an input intent vector in parallel with a noise vector, so that each candidate intent vector in the set of candidate intent vectors is randomly distributed relative to the input intent vector, which has the advantage of creating an intent recognition model faster than using traditional techniques.

[0009] The permitted embodiments involve generating a set of candidate intent vectors by processing an input intent vector in parallel with a noise vector within a long short-term memory (LSTM) model, with the advantage of creating a more accurate intent recognition model than using traditional techniques.

[0010] The licensed embodiment involves determining a multidimensional intention vector distance between an input intention vector and a corresponding candidate intention vector and selecting a candidate intention vector within the multidimensional intention vector distance configured for the input intention vector as a valid intention vector that is semantically similar to the input intention vector, with the advantage of identifying a candidate vector that is closest to a sample intention vector in a multidimensional vector space, thereby identifying similar intention vector encodings.

[0011] A licensing embodiment involves at least one of generating and performing validation as a service in a cloud environment, with the advantages of rapid deployment and serviceability of the intent recognition model creation techniques described herein.

[0012] The permitted embodiments involve creating an intent recognition model from a valid intent vector that is semantically similar to an input intent vector, with the advantage that the intent recognition model can be quickly created from the verified intent vector.

[0013] The licensed embodiments involve, in response to determining during verification that a generated candidate intent vector is not semantically similar to an input intent vector, performing feedback from verification to the generated candidate intent vector generation, adjusting vector generation parameters for generating the candidate intent vector, and iteratively adjusting the vector generation parameters by using the candidate intent vector generation feedback until the generated candidate intent vector is semantically similar to the input intent vector, the advantage of which is that the percentage of high-quality candidate intent vectors generated by the computer is increased, the computational speed of intent recognition model creation is improved, and the accuracy of the generated intent recognition model is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Describes a cloud computing environment according to an embodiment of the present invention;

[0015] Figure 2 Depicting the abstract model layers according to an embodiment of the present invention;

[0016] Figure 3 is a block diagram of an example of an implementation of a system for creating an intent recognition model from randomized intent vector proximities according to an embodiment of the present subject matter;

[0017] Figure 4 is a block diagram of an example of an implementation of a core processing module capable of performing intent recognition model creation from randomized intent vector proximity according to an embodiment of the present subject matter;

[0018] Figure 5 is a flow chart of an example of an implementation of a process for creating an automatic intent recognition model from randomized intent vector proximities according to an embodiment of the present subject matter;

[0019] Figure 6 is a flow chart of an example of an implementation of a process for creating an automatic intent recognition model from randomized intent vector proximities according to an embodiment of the present invention, showing additional details and certain additional / alternative operations. DETAILED DESCRIPTION

[0020] The embodiments listed below represent information necessary to enable those skilled in the art to implement the present invention and to illustrate the best mode for implementing the present invention. After reading the following description according to the accompanying drawings, those skilled in the art will understand the concept of the present invention and will recognize the application of these concepts not particularly mentioned herein. It should be understood that these concepts and applications fall within the scope of the present disclosure and the appended claims.

[0021] The subject matter described herein provides intent recognition model creation from randomized intent vector proximity. The present technology solves the recognized intent recognition model creation problem by providing a technique including a new form of autonomous computer-controlled creation and verification of sentence / phrase vectors that improves the speed of intent recognition model creation and the accuracy of intent recognition model functionality relative to traditional intent recognition models. The technology described herein is based on computer-controlled processing of limited sentence / phrase input variations, and therefore requires less input data to create a more accurate intent recognition model relative to traditional intent recognition models.

[0022] Some embodiments of the present invention may include one or more of the following features, feature operations, and / or advantages: (i) combining noise with sentence (sample) vectors to generate many candidate vectors; (ii) using a long short-term memory (LSTM) model with feedback to facilitate computer generation of a large number of candidate vectors that are jittered around an initial sample vector; (iii) using post-Gaussian vectors that are close in the resulting vector space to identify spoken phrases that are similar in intent to those encoded into the initial sample vector (once also converted to vectors), without using traditional synonym-based vector encoding; (iv) taking into account variations in language use, including variations in dialects and regional languages; and / or (v) improving intent recognition in at least this respect, as traditional synonym-based vector encoding may not be able to recognize intent from phrases spoken using these types of variations in spoken / encoded words.

[0023] The techniques described herein operate based on the observation that vector distances between intent vectors encoded with the same or similar meanings are very close in a multidimensional intent vector space. As a result of this observation, it was determined that the distances between intent vectors can be used to represent the similarity of the meanings of sentences or phrases encoded by different vectors, and by randomly varying the intent vectors themselves, a more accurate intent recognition model is created relative to traditional input text variations. The techniques described herein completely transform the creation of intent recognition models, and the resulting / created intent recognition models significantly improve computer-based recognition of user intent.

[0024] Based on using the distance between intent vectors, the technology described herein computer generates a large number of validated encoded intent vectors to train and form an intent recognition model. Therefore, the technology described herein provides for the creation of large-scale, faster and more accurate intent recognition models using less input data than the traditional manual creation of semantically similar sentences using synonyms and word choice variations and then encoding those similar sentences.

[0025] The technical process described herein takes as input a version of a sample intent text in the form of a sentence or phrase, which represents the intent to be encoded into the intent recognition model. Then, the process converts and encodes the sample intent text into a sample intent vector using a digital encoding method that is convenient for the processing platform. The conversion of the sample intent text to the sample intent vector can be performed by any processing suitable for a given implementation, such as frequency domain encoding (e.g., fast Fourier transform, etc.) or other forms of encoding, and each word or speech is represented in digital form, which can be used to uniquely identify the corresponding word or speech. A computer-based intent vector generator automatically generates a large number of candidate intent vectors from the sample intent vector for evaluating the multidimensional vector distance relative to the sample intent vector. In order to improve accuracy, a computer-based vector verifier filters valid and meaningful candidate intent vectors that are close to each other in the multidimensional vector distance (e.g., matrix dimension cosine similarity, where the number of dimensions evaluated for cosine similarity is equal to the vector matrix dimension) into a final set of similar intent vectors, while discarding any candidate intent vectors that are not similar in the multidimensional vector distance to the sample intent vector (e.g., not close in the matrix dimension cosine similarity). The last set of similar intent vectors is used to train the intent recognition model. As such, the techniques described herein improve the speed and accuracy of creation of computer-based intent recognition models, and additionally improve the speed and accuracy of computer-based intent recognition when the intent recognition models are deployed. As a further result, the techniques described herein operate at scale, with significantly reduced data input requirements, relative to traditional forms of intent recognition model creation. Specifically, in contrast to traditional synonym-based approaches, the techniques described herein utilize a single sample input text to create an intent recognition model that is capable of determining intent for various variations of a user-input phrase that express an intent similar to that of the sample input text.

[0026] To encode the sample intent vector, the sample intent text phrase is digitized according to a configurable / specified dimensionality. For purposes of example, a dimension of three hundred may be specified (dimension=300), but other dimensions such as two hundred and fifty-six (dimension=256) may be used as appropriate for a given implementation. For purposes of example, continuing with a dimension of three hundred (dimension=300), each word of the sample input text may be converted into a 300-dimensional vector comprising 300 digital values ​​representing the particular word. Using a sample input text phrase of four words, the digital representation of the sample input text phrase produces a matrix of four such 300-dimensional vectors (e.g., a matrix of 4×300 digital values) that encode the phrase.

[0027] The obtained digital value matrix is ​​then input into the "Convolutional Neural Network" model, and the intermediate result after the maximum pooling layer processing is selected as the sample intent vector. The sample intent vector thus obtained has the semantics of the sample intent text phrase digitally represented in the multi-dimensional intent vector space.

[0028] To further elaborate on the maximum pooling process in the convolutional neural network algorithm, the representation of the entire sample input text / phrase is abstracted into a single / result sample intent vector. This single vector can be used as a semantic representation of the sample input text / phrase in the semantic vector space. Typically, the type (n) of the convolution kernel and the number (m) of the convolution kernels (n*m) are determined together. In the example above, the dimension of the resulting sample intent vector is also three hundred dimensions (300), where the type (n) has been set to three (n=3) and the number of convolution kernels (m) has been set to one hundred (m=100). However, it should be noted that the number of convolution kernels can alternatively be set to four hundred (400), five hundred (500), or any other value suitable for a given implementation.

[0029] A computer-based intent vector generator generates a configurable number of candidate intent vectors from a sample intent vector by varying individual elements of the sample intent vector. The candidate intent vectors are then evaluated / validated for multi-dimensional vector distances relative to the sample intent vector. For purposes of example, the number of candidate intent vectors to be generated and validated may be set to one thousand and twenty-four (1,024), although the number of configured candidate intent vectors to be generated and validated may be set to any number suitable for a given implementation (e.g., 2,048, 4,096, etc.).

[0030] To further improve the probabilistic accuracy of the generated candidate intent vectors, a computer-based intent vector generator utilizes a randomly generated noise vector, each element of which has a Gaussian (e.g., normal) distribution across frequencies for each element (e.g., word / utterance) of a sample intent vector. Adding Gaussian noise to the sample vector randomly distributes each generated vector in a semantic space around the sample vector. In this way, although the generated vectors do not directly correspond to vocabulary in a real language, the generated vectors have tangible meanings in the semantic space, which can be used to formulate improved and more accurate intent recognition models as described herein. In this embodiment, the computer-based intent vector generator repeatedly and in parallel processes the sample intent vectors with the randomly generated Gaussian noise vectors within a "long short-term memory" (LSTM) model, and obtains a post-Gaussian (candidate) intent vector for each iteration. In this way, the LSTM model processes a large and configurable number of post-Gaussian candidate intent vectors.

[0031] Continuing with this embodiment, a computer-based intent vector verifier receives the generated post-Gaussian candidate intent vectors and inputs each post-Gaussian candidate intent vector into an intent recognition algorithm. A given post-Gaussian candidate intent vector is considered to be accepted as a valid / final intent vector for the intent recognition model by first measuring the difference between the sample input vector and the corresponding post-Gaussian candidate intent vector to obtain a vector distance as a measure of loss or change relative to the sample input vector in the form of error (e.g., a 15% loss). The measure of loss can then be inverted to obtain a confidence level about the similarity of the intent of the corresponding post-Gaussian candidate intent vector to the sample intent vector. For purposes of example, a multidimensional vector distance of zero point eighty-five (0.85) results in a confidence level of eighty-five percent (85%). In addition, a confidence level of eighty-five percent (85%) or higher (e.g., a multidimensional vector distance close to 1.0) can be used to treat the corresponding post-Gaussian candidate intent vector as representing a similar intent that is valid to the sample. Although it should be understood that any other confidence level may be applicable to a given implementation.

[0032] The intent vector generation feedback loop from the intent vector verifier to the intent vector generator and LSTM model is used to train the intent vector generator based on the determined error / loss measurement. The LSTM parameters are iteratively updated throughout the feedback process to improve the intent vector generation for each candidate intent vector generated and additional candidate intent vector processing over time. This iteration and feedback training process is repeated until the output post-Gaussian candidate intent vector generated by the intent vector generator consistently passes the intent vector verifier and is considered a valid / final intent vector for the intent recognition model.

[0033] Thus, the process described herein programmatically self-adjusts intent vector generation to ensure that a higher percentage of valid intent vectors are produced by the computer controlled and automated intent vector generation process. This higher percentage of valid intent vectors is then used to quickly create and train an intent recognition model with a high degree of confidence that, when deployed, accurately determines user intent for a variety of phrases that represent intent similar to that of sample input text.

[0034] Specifically, a large number of valid / final intent vectors that are programmatically generated and verified in the deployed intent recognition model are considered to form an intent vector space or random distribution that statistically spans a multidimensional vector space region around the original sample intent vector. In this way, when a given user of the deployed intent recognition model speaks a phrase, the phrase is digitized, and if the digitized representation of the phrase spoken by the user is close to or located within the multidimensional vector space region created by the set of valid / final intent vectors, the phrase spoken by the user is mapped back to the sample intent vector, and then mapped back to the original sample intent text to determine the user's intent. It should be noted that due to the proximity of intent vectors representing similar intents as described above, even if a specific phrase spoken by the user is not encoded as a final intent vector in the intent recognition model, an accurate interpretation of the user's intent can be determined quickly and in real time. In this way, the technology described herein expands the set of phrases for which intent can be determined without the need for separate encoding of each possible variation of the phrase as required by traditional techniques to recognize a single phrase. In this way, the technology described herein improves the accuracy of the deployed intent recognition model with limited sentence / phrase input variations, and is therefore considered to be a significant improvement over traditional intent recognition model technology.

[0035] Some terms used to describe embodiments of the present invention will now be explained. "Natural language" is defined herein as a recognizable spoken language or dialect spoken by people in one or more regions, and may include formal languages ​​such as Chinese, English, French, German and other languages ​​and dialect variants of these languages. "Convolutional neural network" is defined herein as a computational model that can be executed to convert spoken text into one or more data element sequences. The data element digitally represents a phrase or sentence of the spoken text as a multidimensional matrix, and is processed in matrix form within a digital computing device. "Intention vector element" is defined herein as a spoken word or utterance that has been digitally encoded into a multidimensional vector, which represents and distinguishes spoken words or utterances for processing as digital data elements. "Intention vector" is optionally referred to herein as "sentence vector", and is therefore defined as a sequence of intention vector elements that encode a sentence or phrase of speech into a multidimensional matrix and process as a matrix / unit of data as described herein. "Sample intention vector" is alternatively referred to herein as "sample sentence vector", and is therefore defined as an intention vector generated from a sample text by recursively combining the word embedding in the generative model into a vector format used as the initial / starting input of the intention / sentence vector generation and verification process as described herein. "Gaussian noise" is defined herein as noise that has a standard / normal distribution in frequency relative to a mean or set of mean values. "Gaussian noise intention vector" is defined herein as a matrix of data elements, each data element representing a distribution of Gaussian noise relative to a set of elements represented within a particular intention vector. "Post-Gaussian intention vector" is optionally referred to herein as a "candidate intention vector" and is therefore defined as a resultant matrix of encoded data elements output from parallel processing of an intention vector and a Gaussian noise intention vector, wherein each resulting encoded data element encodes the statistical variation of the corresponding / paired intention vector elements of the intention vector according to the normal distribution provided by the Gaussian noise represented within the sequence of data elements of the Gaussian noise intention vector. "Long short-term memory" (LSTM) is defined herein as a computer processing network capable of processing an intention vector and a Gaussian noise intention vector in parallel as a single input unit and providing an output of a post-Gaussian intention vector. "Multidimensional vector distance" is defined herein as the vector distance between a sample intention vector and a candidate intention vector calculated based on the number of dimensions of the intention vector matrix of each intention vector. "Matrix-dimensional cosine similarity" is defined herein as a form of multidimensional vector distance that identifies the cosine similarity computed over the number of dimensions of the intent vector matrix for each intent vector. "Verified intent vector" is defined herein as a candidate intent vector that has been verified to be within a configured multidimensional vector distance of a given sample sentence vector from which the corresponding post-Gaussian / candidate intent vector was generated. "Verified intent recognition model" is defined herein as an intent recognition model defined using automatically generated and verified intent vectors as defined and described herein.

[0036] The technology described herein operates by generating a set of candidate intent vectors from an input intent vector, and performing validation of the set of candidate intent vectors, which selects any set of candidate intent vectors that are semantically similar to the input intent vector as valid intent vectors. This embodiment has the advantage of improving intent recognition model creation and validation by relying on less sentence / phrase input, and by autonomously deriving a robust intent recognition model with increased intent inference accuracy.

[0037] Some advantages of some embodiments of the present subject matter will now be described. An embodiment of performing the generation of a set of candidate intent vectors by iteratively processing an input intent vector in parallel with a noise vector, so that each candidate intent vector of the set of candidate intent vectors is randomly distributed relative to the input intent vector, which has the advantage of being able to create an intent recognition model faster than using conventional techniques. An embodiment of performing the generation of a set of candidate intent vectors by processing an input intent vector in parallel with a noise vector within a long short-term memory (LSTM) model has the advantage of being able to create an intent recognition model more accurately than using conventional techniques. An embodiment of performing verification of a set of candidate intent vectors involves determining a multidimensional intent vector distance between an input intent vector and a corresponding candidate intent vector and selecting a candidate intent within a multidimensional intent vector configured with the input intent vector as a valid intent vector that is semantically similar to the input intent vector, with the advantage of identifying candidate vectors that are adjacent to the sample intent vector in the multidimensional vector space and then encoding similar intent vectors. An embodiment involving generating and performing verification as at least one of the services in a cloud environment has the advantage of rapid deployment and serviceability of intent recognition model creation. An embodiment involving creating an intent recognition model from a valid intent vector that is semantically similar to the input intent vector has the advantage of quickly creating an intent recognition model from a verified intent vector. An embodiment relates to adjusting vector generation parameters for generating candidate intent vectors by performing candidate intent vector generation feedback from verification, in response to determining during verification that the generated candidate intent vector is not semantically similar to the input intent vector, and iteratively adjusting the vector generation parameters by using the candidate intent vector generation feedback until the resulting generated candidate intent vector is semantically similar to the input intent vector, having the advantage of increasing the percentage of high-quality candidate intent vectors generated by the computer, which improves the computational speed of intent recognition model creation and improves the accuracy of the resulting intent recognition model.

[0038] It should be noted that the concepts of the present subject matter stem from the identification of certain limitations associated with the creation of intent recognition models. For example, it has been observed that creating a robust intent recognition model requires a large amount of data in the form of verified high-quality sentence / phrase input, but the personnel responsible for creating the intent recognition model generally do not have access to verified high-quality sentence / phrase input. It has been further observed that due to the limited available input data, the intent recognition model has a high error rate with respect to the intent inferred from the user's statement / question relative to the user's actual intent. Based on these observations, it was determined that new computational techniques are needed to improve the creation and verification of intent recognition models that rely on fewer sentence / phrase inputs and autonomously derive robust intent recognition model creation with increased intent reasoning accuracy relative to traditional forms of intent recognition. The present subject matter described herein improves intent recognition model creation and verification by providing intent recognition model creation from randomized intent vector proximities, as described in more detail above and below. In this way, improved intent recognition model creation and verification is obtained by using the present technology.

[0039] Creation of intent recognition models based on the randomized intent vector proximities described herein can be performed in real time to allow rapid creation and validation of intent recognition models. For the purposes of this specification, real-time shall include any time frame of sufficiently short duration to provide a reasonable response time for information processing acceptable to users of the subject matter. In addition, the term "real-time" shall include what is commonly referred to as "near real-time" - generally meaning any time frame of sufficiently short duration to provide a reasonable response time (e.g., within a fraction of a second or within a few seconds) for on-demand information processing acceptable to users of the subject matter described. These terms, while difficult to define precisely, are well understood by those skilled in the art.

[0040] Table 1 below shows an example of a portion of an encoded 300-dimensional sample intent vector based on the input text / phrase "I can't breathe, is this a symptom of chronic rhinitis?" Table 1 below omits many intermediate values ​​by using ellipsis points to reduce the length of the example. However, it should be understood that the encoded 300-dimensional sample intent vector forming the example of Table 1 includes three hundred (300) elements because, as described above and for the purpose of example, the type (n) has been set to three (n=3) and the number of convolution kernels (m) is set to one hundred (m=100) in the convolutional neural network algorithm used to encode the sample intent vector.

[0041] 0.06217742 -0.09291203 -0.02840577 0.00374683 -0.10506968 -0.13434719 -0.15676261 0.04667582 0.19579822 0.04030076 -0.21957162 0.04010129 0.07424975 0.01369099 0.07016561 0.14306846 -0.08841907 -0.28088206 -0.23514667 -0.01831863 -0.00475325 -0.0811624 -0.37781501 -0.17635982 -0.00429389 0.05205948 -0.03414214 -0.11769973 -0.02660859 0.16670637 0.07024501 0.25753206 0.01059075 -0.12026082 -0.22058713 0.06767957 … … … … … … 0.10361771 0.09943719 0.32700372 -0.0283169 0.12805387 0.01565604

[0042] Table 1 Part of the 300-dimensional sample intent vector

[0043] Based on the example of Table 1 above, it should also be noted that each of the multiple candidate intent vectors generated according to the techniques described herein will also be represented as a 300-dimensional candidate intent vector, where the individual elements of the corresponding candidate intent vectors can be randomly changed relative to the corresponding individual elements of the sample intent vector. The example of Table 1 above provides sufficient detail regarding the formation of the sample intent vectors and candidate intent vectors as described herein, and additional examples of digital encoding patterns of intent vectors are omitted for the sake of brevity. As described above, a multi-dimensional vector distance calculation is applied to determine the relative vector space distance between the sample intent vector and each generated candidate intent vector to perform the verification described herein.

[0044] Additional details of algorithmic processing and computational efficiency are provided further below.The following portion of this specification provides an example of a high-level computing platform in which the present technology may be implemented, followed by further details of creating an intent recognition model based on the randomized intent vector proximity described herein.

[0045] First, it should be understood that although the present disclosure includes a detailed description about cloud computing, the implementation of the technical solutions recorded therein is not limited to a cloud computing environment, but can be implemented in combination with any other type of computing environment now known or later developed.

[0046] Cloud computing is a service delivery model for convenient, on-demand network access to a shared pool of configurable computing resources. Configurable computing resources are resources that can be quickly deployed and released with minimal management cost or interaction with the service provider, such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0047] Features include:

[0048] On-demand self-service: Cloud consumers can unilaterally and automatically deploy computing capabilities such as server time and network storage on demand without human interaction with the service provider

[0049] Broad network access: Computing power can be accessed over the network through standard mechanisms that facilitate the use of the cloud through different types of thin-client or thick-client platforms (e.g., mobile phones, laptops, personal digital assistants (PDAs)).

[0050] The computing resources of the provider are grouped into resource pools and serve multiple consumers through a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated on demand. In general, consumers cannot control or even know the exact location of the provided resources, but can specify the location (such as country, state, or data center) at a higher level of abstraction, so it is location-independent.

[0051] Rapid elasticity: The ability to quickly and elastically (sometimes automatically) deploy computing power to scale up quickly, and quickly release it to scale down quickly. To consumers, the computing power available for deployment often appears to be unlimited, and any amount of computing power can be accessed at any time.

[0052] Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0053] The service model is as follows:

[0054] Software as a Service (SaaS): The capability provided to consumers is to use the provider's applications running on a cloud infrastructure. Applications can be accessed from a variety of client devices through a thin client interface such as a web browser (e.g., web-based email). Aside from limited user-specific application configuration settings, consumers neither manage nor control the underlying cloud infrastructure including networks, servers, operating systems, storage, or even individual application capabilities.

[0055] Platform as a Service (PaaS): The capability provided to consumers is to deploy applications created or acquired by consumers on the cloud infrastructure, which are created using programming languages ​​and tools supported by the provider. Consumers neither manage nor control the underlying cloud infrastructure including networks, servers, operating systems or storage, but have control over the applications they deploy and may also have control over the configuration of the application hosting environment.

[0056] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, network, and other basic computing resources on which consumers can deploy and run arbitrary software including operating systems and applications. Consumers neither manage nor control the underlying cloud infrastructure, but have control over the operating system, storage, and applications they deploy, and may have limited control over selected network components (such as host firewalls).

[0057] The deployment model is as follows:

[0058] Private Cloud: Cloud infrastructure is run solely for an organization. The cloud infrastructure can be managed by the organization or a third party and can exist inside or outside the organization.

[0059] Community Cloud: A cloud infrastructure is shared by several organizations and supports a specific community with common interests (e.g., mission, security requirements, policies, and compliance considerations). A community cloud can be managed by multiple organizations within the community or by a third party and can exist inside or outside the community.

[0060] Public cloud: Cloud infrastructure is provided to the public or to large industry groups and is owned by an organization that sells cloud services.

[0061] Hybrid cloud: A cloud infrastructure consisting of two or more clouds (private, community or public) in a deployment model that remain distinct entities but are bound together by standardized or proprietary technologies (such as cloud burst sharing for load balancing between clouds) that enable data and application portability.

[0062] The cloud computing environment is service-oriented, with characteristics centered on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure that includes a network of interconnected nodes.

[0063] Reference now Figure 1 , which shows an exemplary cloud computing environment 50 according to an embodiment of the present invention. As shown in the figure, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which a local computing device used by a cloud computing consumer can communicate. The local computing device can be, for example, a personal digital assistant (PDA) or a mobile phone 54A, a desktop computer 54B, a laptop computer 54C and / or an automobile computer system 54N. The cloud computing nodes 10 can communicate with each other. The cloud computing nodes 10 can be physically or virtually grouped in one or more networks including but not limited to private clouds, community clouds, public clouds or hybrid clouds, or a combination thereof as described above (not shown in the figure). This allows cloud consumers to provide infrastructure as a service, platform as a service and / or software as a service without having to maintain resources on local computing devices. It should be understood that Figure 1 The various types of computing devices 54A-N shown are merely illustrative, and the cloud computing node 10 and the cloud computing environment 50 may communicate with any type of computing device on any type of network and / or network-addressable connection (eg, using a web browser).

[0064] Reference now Figure 2 , which shows a cloud computing environment 50 ( Figure 1 ) provides a set of functional abstraction layers. First of all, it should be understood that Figure 2The components, layers, and functions shown are only exemplary and the embodiments of the present invention are not limited thereto. Figure 2 As shown, the following layers and corresponding functions are provided:

[0065] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: host 61; server 62 based on RISC (Reduced Instruction Set Computer) architecture; server 63; blade server 64; storage device 65; network and network components 66. Examples of software components include: network application server software 67 and database software 68.

[0066] The virtualization layer 70 provides an abstraction layer that can provide examples of the following virtual entities: virtual servers 71 , virtual storage 72 , virtual networks 73 (including virtual private networks), virtual applications and operating systems 74 , and virtual clients 75 .

[0067] In one example, the management layer 80 may provide the following functions: Resource provisioning function 81: providing dynamic acquisition of computing resources and other resources for performing tasks in a cloud computing environment; Metering and pricing function 82: tracking the cost of resource usage within the cloud computing environment and providing bills and invoices for this purpose. In one example, the resource may include an application software license. Security function: providing identity authentication for cloud consumers and tasks, and providing protection for data and other resources. User portal function 83: providing access to the cloud computing environment for consumers and system administrators. Service level management function 84: providing allocation and management of cloud computing resources to meet the required service levels. Service level agreement (SLA) planning and fulfillment function 85: providing pre-arrangement and provision for future demand for cloud computing resources predicted according to the SLA.

[0068] The workload layer 90 provides examples of functionality that can take advantage of a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analytics processing 94; transaction processing 95; and creating intent recognition models from randomized intent vector proximities.

[0069] In the above examples, the cloud computing environment illustrates several types of computing devices 54A-N on which virtual agents can be deployed and managed. Figure 3 and 4 It should be understood that for a given embodiment, various alternatives may be combined with or substituted for the above-described embodiment options.

[0070] Figure 31 is a block diagram of an example of an implementation of a system 100 for creating an intent recognition model from randomized intent vector proximity. Computing device_1 102 to computing device_N 104 communicate with several other devices via a network 106. Computing device_1 102 to computing device_N 104 represent user devices that utilize one or more intent recognition models for user interface and device control processing. However, it should be noted that computing device_1 102 to computing device_N 104 can additionally / alternatively implement automatic intent recognition model creation from the randomized intent vector proximity described herein. As a further alternative, an intent recognition model (IRM) server 108 can provide services to computing device_1 102 to computing device_N 104 by generating and verifying one or more intent recognition models based on automatic intent recognition model creation of randomized vector proximity described herein. As appropriate for a given implementation, the IRM server 108 or one or more of computing device_1 102 through computing device_N 104 may store one or more validated intent recognition models locally or within an intent recognition model (IRM) database 110 for access and use within the system 100 .

[0071] In view of the above implementation scheme, the present technology can be implemented in a cloud computing platform, on a user computing device, at a server device level, or by a combination of these platforms and devices suitable for a given implementation. There are various possibilities for implementing the present subject matter, and all of these possibilities are considered to be within the scope of the present subject matter.

[0072] The network 106 includes any form of interconnection suitable for the intended purpose, including private or public networks such as an intranet or the Internet, direct module-to-module interconnection, dial-up, wireless, or any other interconnection mechanism capable of interconnecting various devices, respectively.

[0073] IRM server 108 comprises any device capable of providing data for consumption by devices such as computing device_1 102 (e.g., computing device_1 110) via a network such as network 106. Thus, IRM server 108 may comprise a network, server, application server, or other data server device.

[0074] IRM database 110 includes a relational database, an object database, or any other storage type of device. In this way, IRM database 110 can appropriately implement a given implementation.

[0075] Figure 4is a block diagram of an example of an implementation of a core processing module 200 capable of performing intent recognition model creation from randomized intent vector proximities. As appropriate for a given implementation, the core processing module 200 may be associated with computing device_1 102 to computing device_N 104, or with an IRM server 108, with devices within a cloud computing environment 50. Thus, the core processing module 200 is generally described herein, but it should be understood that many variations of the implementation of components within the core processing module 200 are possible, and all such variations are within the scope of the present subject matter. Furthermore, as appropriate for a given implementation, the core processing module 200 may be implemented as an embedded processing device having circuitry specifically designed to perform the processing described herein.

[0076] The core processing module 200 can provide different and complementary processing for intent recognition model creation from the randomized intent vector proximity associated with each implementation. Thus, for any of the following examples, it should be understood that any aspect of the functionality described with respect to any one device described in conjunction with another device (e.g., sending / transmitting, etc.) should be understood to simultaneously describe the functionality of the other corresponding device (e.g., receiving / receiving, etc.).

[0077] Central processing unit (CPU) 202 ("processor" or "application specific" processor) provides the hardware for performing computer instruction execution, computations, and other functions within core processing module 200. Display 204 provides visual information to a user of core processing module 200 and input devices 206 provide input capabilities for the user.

[0078] Display 204 includes any display device, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED), an electronic ink display, a projection, a touch screen, or other display element or panel. Input device 206 includes a computer keyboard, a keypad, a mouse, a pen, a joystick, a touch screen, a voice command processing unit, or any other type of input device by which a user can interact with and respond to information on display 204.

[0079] It should be noted that the display 204 and the input device 206 may be optional components of the core processing module 200 for certain implementations / devices, or may be located remotely from the respective device and hosted by another computing device located therein that communicates with the respective device. Thus, the core processing module 200 may operate as a fully automated embedded device without direct user configurability or feedback. However, as appropriate for a given implementation, the core processing module 200 may also provide user feedback and configurability via the display 204 and the input device 206, respectively.

[0080] As appropriate for a given implementation, the communication module 208 provides hardware, protocol stack processing, and interconnect functionality that allows the core processing module 200 to communicate with other modules within the system 100 or within the cloud computing environment 50. As appropriate for a given implementation, the communication module 208 includes any electrical, protocol, and protocol conversion capabilities that can be used to provide interconnect capabilities. Thus, the communication module 208 represents a communication device capable of communicating with other devices. As appropriate for a given implementation, the communication module 208 may also include one or more wireless communication capabilities.

[0081] The memory 210 includes an intent vector processing and storage area 212, which stores intent vector processing information within the core processing module 200. As will be described in more detail below, the intent vector processing information 212 stored in the intent vector processing and storage area is used to quickly (in real time) generate high-quality intent recognition models based on limited sentence / phrase input changes. In this way, the intent vector processing and storage area 212 stores sample intent vectors, Gaussian noise vectors, one or more groups of post-Gaussian / candidate intent vectors, and verified verification intent vectors within the multi-dimensional vector distance of the configuration of a given sample intent vector. Processing for creating one or more verified intent recognition models can be performed within the memory 210. The verified intent recognition model storage area 214 provides storage for one or more verified intent recognition models generated from the verified intent vectors. The verified intent recognition model can be deployed / distributed or used within the verified intent recognition model storage area 214. Suitable for a given implementation, many other forms of intermediate information can be stored in the verified intent recognition model storage area 214 in association with the creation and verification of the verified intent recognition model.

[0082] It should be understood that the memory 210 includes any combination of volatile and non-volatile memories that are suitable for the intended purpose, appropriately distributed or located, and may include other memory portions that are not shown in this example for ease of illustration. For example, the memory 210 may include a code storage area, an operating system storage area, a code execution area, and a data area without departing from the scope of the present subject matter.

[0083] Also shown is an intent vector creation and proximity verification module 216. The intent vector creation and proximity verification module 216 includes an intent vector generator module 218 that provides randomized Gaussian / candidate intent vector generation. The intent vector creation and proximity verification module 216 also includes an intent vector verifier module 220 that provides verification of the proximity of the generated candidate intent vectors to the sample intent vectors to the core processing module 200, as described in more detail above and below. The intent vector creation and proximity verification module 216 implements automatic intent recognition model creation from the randomized vector proximity of the core processing module 200 by performing sample intent vector encoding, candidate intent vector generation, verification of candidate intent vectors, generation of verified intent recognition, and model and other related processing suitable for a given implementation.

[0084] It should also be noted that the intent vector creation and proximity verification module 216 may form part of the other circuits described without departing from the scope of the present subject matter. The intent vector creation and proximity verification module 216 may form part of an interrupt service routine (ISR), part of an operating system, or part of an application program without departing from the scope of the present subject matter. The intent vector creation and proximity verification module 216 may also include an embedded device having circuitry specifically designed to perform the processing described herein as appropriate for a given implementation.

[0085] Again in Figure 4 , an IRM database 110 is shown associated with the core processing module 200. Thus, as appropriate for a given implementation, the IRM database 110 can be operably coupled to the core processing module 200 without using a network connection, and can be used in association with the core processing module 200 to create or use one or more validated intent recognition models.

[0086] The CPU 202, display 204, input device 206, communication module 208, memory 210, intent vector creation and proximity verification module 216, and IRM database 110 are interconnected via interconnect 222. Interconnect 222 includes a system bus, a network, or any other interconnect capable of providing appropriate interconnection for the respective components for their respective purposes.

[0087] Although for ease of illustration and description purposes, Figure 4The different modules shown in the are shown as component-level modules, but it should be noted that these modules include any hardware, programmed processors and memories for performing the following operations. The functions of the various modules are described above and in more detail below. For example, the modules may include application-specific integrated circuits (ASICs), processors, antennas and / or additional controller circuits in the form of discrete integrated circuits and components for performing communication and electrical control activities associated with the various modules. In addition, the modules may include appropriate interrupt-level, stack-level and application-level modules. In addition, the modules may include any memory components for storage, execution and data processing for performing processing activities associated with the various modules. The modules may also form part of the other circuits described, or may be combined without departing from the scope of the present subject matter.

[0088] In addition, although the core processing module 200 is shown and has certain components described, other modules and components may be associated with the core processing module 200 without departing from the scope of the present subject matter. In addition, it should be noted that although the core processing module 200 is described as a single device for ease of illustration, the components within the core processing module 200 may be co-located or distributed and interconnected via a network without departing from the scope of the present subject matter. Many other possible arrangements of the components for the core processing module 200 are possible, and all of these arrangements are considered to be within the scope of the present subject matter. It should also be understood that although the IRM database 110 is shown as a separate component for illustrative purposes, the information stored in the IRM database 110 may also / optionally be stored in the memory 210 without departing from the scope of the present subject matter. Therefore, the core processing module 200 can take many forms and can be associated with many platforms.

[0089] Described below Figures 5 and 6 An example process is represented that may be performed by a device such as the core processing module 200 to perform automatic intent recognition model creation from randomized intent vector proximity associated with the present subject matter. Many other variations on the example process are possible, and all variations are considered to be within the scope of the present subject matter. The example process may be performed by modules associated with these devices, such as the intent vector creation and proximity verification module 216 and / or by the CPU 202. It should be noted that for ease of illustration, timeout processes and other error control processes are not shown in the example processes described below. However, it should be understood that all of these procedures are considered to be within the scope of the present subject matter. In addition, the described processes may be combined, the sequence of the described processes may be changed, and additional processes may be added or removed without departing from the scope of the present subject matter.

[0090] Figure 55 is a flow chart of an example of an implementation of a process 500 for creating an automatic intent recognition model from randomized intent vector proximities. The process 500 represents a computer-implemented method for performing the computer-based validation intent recognition model creation described herein. At block 502, the process 500 generates a plurality of candidate intent vectors from an input intent vector. At block 504, the process 500 performs validation of the plurality of candidate intent vectors, selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector.

[0091] Figure 6 6 is a flowchart of an example of an implementation of a process 600 for creating an automatic intent recognition model from randomized intent vector proximities, which shows additional details and certain additional / alternative operations. Process 600 represents a computer-implemented method for performing the computer-based validation intent recognition model creation described herein. At decision point 602, process 600 determines whether to generate an intent recognition model (IRM). In response to determining to generate an intent recognition model, process 600 receives sample text at box 604. The sample text is a phrase or sentence to be encoded, and an intent recognition model is generated for it. At box 606, process 600 encodes the sample text into a sample intent vector. As described above, the sample intent vector is generated from the received sample text as a multidimensional matrix, where each word corresponds to a dimensional vector (e.g., dimension = 300). The sample intent vector as a multidimensional matrix of dimensional vectors is used as an input for the intent vector generation and validation process. The sample intent vector includes a sequence of data elements representing the sample text. In this way, the sample intent vector represents a sequence of digitized words / utterances as a sequence of encoded and distinguishable elements.

[0092] At block 608, process 600 generates a Gaussian noise vector based on the sample intent vector. As described above, the Gaussian noise vector includes a sequence of data elements, each of which represents a Gaussian noise distribution relative to the individual elements forming the sample intent vector, and is also a matrix of noise elements. Thus, there is a mapping between the elements of the sample intent vector and the Gaussian noise distribution elements of the Gaussian noise vector.

[0093] At block 610, process 600 configures long short-term memory (LSTM) parameters for use by the LSTM model in generating candidate intent vectors. As described above, the LSTM model generates candidate intent vectors by processing corresponding sample intent vectors and Gaussian noise vectors in parallel. The output of the LSTM processing produces a set of candidate intent vectors.

[0094] The process 600 begins an iterative process to generate a set of candidate intent vectors. At block 612, the process 600 again generates candidate intent vectors by processing the sample intent vector in parallel with the Gaussian noise vector. Each element of the sample intent vector as a matrix is ​​randomly processed according to the corresponding element of the Gaussian noise vector to create a unique candidate intent vector.

[0095] Then, process 600 begins candidate intent vector validation processing. At box 614, process 600 processes the candidate intent vector using an intent recognition algorithm. At box 616, process 600 determines the vector distance and confidence level of the candidate intent vector relative to the sample intent vector. The vector distance and confidence of the candidate intent vector are determined by measuring the multidimensional vector distance of the candidate intent vector from the sample intent vector, for example, by determining the matrix-dimensional cosine similarity of two multidimensional vectors / matrices. An acceptable multidimensional vector distance or threshold for accepting a candidate intent vector as a valid intent vector is established and configured using any suitable vector distance algorithm. For purposes of example and not limitation, a multidimensional vector distance or other distance measurement technique may be utilized, and the threshold multidimensional vector distance is zero point eighty-five (0.85), or expressed as a confidence level of eighty-five. The percentage (85%) may be configured so that any candidate intent vector having a multidimensional vector distance that meets or exceeds a confidence level (85%) of eighty-five percent is considered to be a valid intent vector, while any candidate intent vector having a multidimensional vector distance with a confidence level below eighty-five percent (85%) is not considered to be a valid intent vector. It should be noted that many techniques for measuring vector distances are possible, and all of these techniques are considered to be within the scope of the present subject matter.

[0096] The process 600 then begins filter-based processing to determine the matching quality of the candidate intent vector relative to the sample intent vector. At decision point 618, the process 600 determines whether the candidate intent vector passes verification based on the measurement of the loss and confidence level of the candidate intent vector. As described above, if the candidate intent vector is close to the sample intent vector in the multi-dimensional intent vector space (e.g., the matrix dimension cosine similarity is greater than or equal to 0.85), the candidate intent vector is considered to be a valid vector and passes verification).

[0097] In response to determining at decision point 618 that the candidate intent vector has not passed verification (e.g., the cosine similarity for the matrix dimensions is less than 0.85), process 600 begins an iterative feedback process to adjust the candidate intent vector generation parameters of the LSTM model used to generate the candidate intent vector. At box 620, process 600 discards the candidate intent vector that has not passed verification. At box 622, process 600 updates the LSTM parameters used to generate the candidate intent vector from the sample intent vector and the Gaussian noise vector. The process iterates and returns to box 612 to generate new candidate intent vectors using the updated LSTM parameters and iterates as described above.

[0098] Returning to the description of decision point 618, in response to determining that the candidate intent vector passes validation (e.g., cosine similarity for the matrix dimensions is greater than or equal to 0.85), process 600 assigns the candidate intent vector as a final intent recognition model (IRM) vector by adding the candidate intent vector to the intent recognition model at box 624. At box 626, process 600 stores the LSTM parameters used during creation of the validated intent vector and the confidence level determined with the final IRM vector.

[0099] At decision point 628, process 600 determines whether to iterate to generate and verify another candidate intent vector. In response to determining to generate and verify another candidate intent vector, process 600 returns to box 610 to configure new LSTM parameters to enable the LSTM model to generate a new / unique candidate intent vector based on the sample intent vector and the Gaussian noise vector. In this way, process 600 iteratively adjusts LSTM parameters to quickly generate multiple verified intent vectors within the intent recognition model.

[0100] The process 600 iterates as described above to generate and validate candidate intent vectors until sufficient final IRM vectors have been added to the intent recognition model to implement intent recognition across the various automatically generated intent vectors of the intent recognition model during deployment. In response to determining at decision point 628 not to generate and validate another candidate intent vector, the process 600 stores the final IRM vector and mapping to the received and encoded sample text as the intent recognition model vector set at box 630. At box 632, as appropriate for a given implementation, the process 600 deploys the generated intent recognition model to the IRM database 110 or. The process 600 returns to decision point 602 and iterates as described above.

[0101] Thus, process 600 receives sample text and encodes it into a sample intent vector, and processes the sample intent vector in parallel with the Gaussian noise vector using an LSTM model that has been parameterized to adjust the effects of elements of the Gaussian noise vector on elements of the sample intent vector. Process 600 iteratively generates and verifies candidate intent vectors within a feedback loop that adjusts the LSTM parameters until a high-quality (close / minimum distance) final IRM vector is produced. Process 600 iteratively creates the final IRM vector until a sufficient number of final IRM vectors span the natural (random) variation of the original sample text. Process 600 stores a verified set of final IRM vectors as an intent recognition model and deploys the computer-generated intent recognition model.

[0102] Some embodiments of the present invention improve computer technology in one or more of the following ways: (i) lower input data requirements for large-scale intent recognition model generation; (ii) faster generation of intent recognition models using less data; (iii) more accurate deployment of intent recognition from the generated intent recognition models.

[0103] The present invention is not abstract because it particularly relates to computer operations and / or hardware, including for the following reasons: (i) computer-controlled processing that creates a more accurate intent recognition model using fewer inputs than traditional methods; (ii) computer-controlled final / validation intent vector distribution that programmatically changes sample intent vectors to improve the creation of intent recognition models; and (iii) computer-controlled processing that improves the computer's accuracy in recognizing user intent from the user's spoken phrases.

[0104] As above combined Figures 1 to 6 As described, example systems and processes provide for creating an intent recognition model from randomized intent vector proximities. Many other variations and additional activities associated with creating an intent recognition model from randomized intent vector proximities are possible, and all are considered within the scope of the present subject matter.

[0105] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0106] Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, - but not limited to - an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. The computer readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.

[0107] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0108] The computer program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of a computer-readable program instruction, and the electronic circuit may execute a computer-readable program instruction, thereby realizing various aspects of the present invention.

[0109] Various aspects of the present invention are described herein with reference to the flow charts and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each box of the flow chart and / or block diagram and the combination of each box in the flow chart and / or block diagram can be implemented by computer-readable program instructions.

[0110] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0111] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0112] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present invention. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.

[0113] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the present invention. As used herein, the singular forms "a", "an", "this", "that" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "includes", when used in this specification, specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0114] The corresponding structures, materials, actions, and equivalents of all means or step-plus-function elements in the claims are intended to include any structure, material, or action for performing a function in combination with other claimed elements as specifically claimed. The description of the present invention is presented for the purpose of illustration and description, but is not intended to be exhaustive or limit the present invention in the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art based on the teachings herein without departing from the scope of the present invention. The subject matter is described in order to explain the principles and practical applications of the present invention and to enable other persons of ordinary skill in the art to understand the various embodiments of the present invention and various modifications suitable for the intended specific use.

Claims

1. A computer-implemented method, include: By the processor: generating a plurality of candidate intent vectors from an input intent vector; performing a validation of the plurality of candidate intent vectors, the validation selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector, and The input intent vector is iteratively processed in parallel with the noise vector such that each candidate intent vector in the plurality of candidate intent vectors is randomly distributed relative to the input intent vector.

2. The computer-implemented method of claim 1 , wherein the processor generates a plurality of candidate intent vectors from an input intent vector include: The processor, for each candidate intent vector in the plurality of candidate intent vectors: The input intent vector and the noise vector are processed in parallel in the Long Short-Term Memory (LSTM) model.

3. The computer-implemented method of claim 1 , wherein the processor performs a validation of the plurality of candidate intent vectors, the validation selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector include: The processor, for each candidate intent vector in the plurality of candidate intent vectors: determining a multi-dimensional intent vector distance between an input intent vector and a corresponding candidate intent vector; and A candidate intent vector that is within the multi-dimensional intent vector distance configured by the input intent vector is selected as a valid intent vector that is semantically similar to the input intent vector.

4. The computer-implemented method of claim 1 , further comprising: In response to determining during verification that the generated candidate intent vector is not semantically similar to the input intent vector, performing feedback from the verification to the generated candidate intent vector to adjust vector generation parameters used to generate the candidate intent vector; and The vector generation parameters are iteratively adjusted by using the candidate intent vector generation feedback until the generated candidate intent vector is semantically similar to the input intent vector.

5. The computer-implemented method of claim 1 further comprises the processor creating an intent recognition model from valid intent vectors that are semantically similar to the input intent vector. 6 . The computer-implemented method of claim 1 , further comprising providing at least one of processor generation and processor execution verification as a service of a cloud environment.

7. A computer system, include: Memory; and At least one processor set programmed to: generating, in a memory, a plurality of candidate intent vectors from an input intent vector; performing verification of the plurality of candidate intent vectors, the verification selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector; as well as The input intent vector is iteratively processed in parallel with the noise vector such that each candidate intent vector in the plurality of candidate intent vectors is randomly distributed relative to the input intent vector.

8. The system of claim 7, wherein the processor is programmed to generate a plurality of candidate intent vectors from the input intent vector in the memory, the at least one processor set being programmed to, for each candidate intent vector in the plurality of candidate intent vectors: The input intent vector and the noise vector are processed in parallel in the Long Short-Term Memory (LSTM) model.

9. The system of claim 7, wherein the at least one processor set is programmed to perform validation of a plurality of candidate intent vectors, the validation selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector, the at least one processor set being programmed to, for each of the plurality of candidate intent vectors: determining a multi-dimensional intent vector distance between an input intent vector and corresponding candidate intent vectors; and A candidate intent vector that is within the multi-dimensional intent vector distance configured by the input intent vector is selected as a valid intent vector that is semantically similar to the input intent vector.

10. The system of claim 7, wherein the at least one processor set is further programmed to: In response to determining during verification that the generated candidate intent vector is not semantically similar to the input intent vector, performing feedback from the verification to the generated candidate intent vector to adjust vector generation parameters used to generate the candidate intent vector; and The vector generation parameters are iteratively adjusted by using the candidate intent vector generation feedback until the generated candidate intent vector is semantically similar to the input intent vector.

11. The system of claim 7, wherein the at least one processor set is further programmed to create an intent recognition model from valid intent vectors that are semantically similar to the input intent vector.

12. A computer program product, include: A computer readable storage medium containing computer readable program code, wherein the computer readable storage medium itself is not a transient signal, and the computer readable program code when executed on a computer causes the computer to: generating a plurality of candidate intent vectors from an input intent vector; performing verification of the plurality of candidate intent vectors, the verification selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector; as well as The input intent vector is iteratively processed in parallel with the noise vector such that each candidate intent vector in the plurality of candidate intent vectors is randomly distributed relative to the input intent vector.

13. The computer program product of claim 12, wherein when causing a computer to generate a plurality of candidate intent vectors from an input intent vector, the computer readable program code when executed on the computer causes the computer to, for each candidate intent vector in the plurality of candidate intent vectors: The input intent vector and the noise vector are processed in parallel in the Long Short-Term Memory (LSTM) model.

14. The computer program product of claim 12, wherein when causing a computer to perform a validation of a plurality of candidate intent vectors, the validation selecting any one of the plurality of candidate intent vectors that is semantically similar to the input intent vector as a valid intent vector, the computer readable program code when executed on the computer causes the computer to, for each of the plurality of candidate intent vectors: determining a multi-dimensional intent vector distance between an input intent vector and corresponding candidate intent vectors; and A candidate intent vector that is within the multi-dimensional intent vector distance configured by the input intent vector is selected as a valid intent vector that is semantically similar to the input intent vector.

15. The computer program product of claim 12, wherein the computer readable program code when executed on a computer further causes the computer to: In response to determining during verification that the generated candidate intent vector is not semantically similar to the input intent vector, performing feedback from the verification to the generated candidate intent vector to adjust vector generation parameters used to generate the candidate intent vector; and The vector generation parameters are iteratively adjusted by using the candidate intent vector generation feedback until the generated candidate intent vector is semantically similar to the input intent vector.

16. The computer program product of claim 12, wherein the computer readable program code, when executed on a computer, further causes the computer to create an intent recognition model from valid intent vectors that are semantically similar to the input intent vector.

17. The computer program product of claim 12, wherein at least one of the computer-generated and computer-executed validation is provided as a service of a cloud environment.

Citation Information

Patent Citations

  • Neural network-based dialogue semantic intention prediction method and learning training method

    CN108363690A