Voice response system based on personalized vocabulary and user profiling - A personalized language processing AI engine
The system addresses the challenge of user-specific language interpretation by using IoT data to train AI voice responses, ensuring accurate and contextually relevant interactions through personalized vocabulary analysis and feedback refinement.
Patent Information
- Application Number
- JP2023504649
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-24
- Filing Date
- 2021-06-01
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-06-01
AI Technical Summary
Existing AI voice response systems struggle to understand and respond to user requests accurately due to variations in personal language usage based on emphasis, context, cultural, and demographic influences, leading to misunderstandings of keywords and phrases like homonyms, homographs, and complex vocabulary.
A system that collects user data from IoT sensors to identify a personalized vocabulary, trains a voice response system using Bi-LSTM and TI-CNN modules, and responds to verbal requests by analyzing unknown content against the user's personalized vocabulary, leveraging regional and cultural corpora for assistance.
Enhances the accuracy of voice response systems by providing personalized and contextually relevant responses, improving user understanding and satisfaction through iterative feedback learning.
Smart Images

Figure 0007725156000001 
Figure 0007725156000002 
Figure 0007725156000003
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to the field of computing, and more particularly to data management and analysis. [Background technology]
[0002] Humans may have unique ways of interpreting spoken language, which may be based on applied emphasis and / or context given personal experiences and / or events, or may be based on cultural and / or demographic influences. Human language may be further based on, among other things, pragmatics (e.g., situational context), syntax (e.g., placement of words and phrases in clauses and sentences or both), morphology (e.g., word function, which may be grammatical or lexical or both), semantics (e.g., word meaning), phonology (e.g., classification and / or study of sounds), and phonetics (e.g., study and / or classification of sounds). An artificial intelligence (AI) voice response system can analyze a user's voice request and respond to the user accordingly using a pre-programmed response generator. Summary of the Invention
[0003] Embodiments of the present invention disclose a method, computer system, and computer program product for personalized voice response. The method may include collecting a plurality of user data from Internet of Things (IoT) connected sensors. The method may include identifying a personalized vocabulary based on the collected plurality of user data. The method may include training a voice response system based on the collected plurality of user data and the identified personalized vocabulary. The method may include receiving a verbal request. The method may include utilizing the trained voice response system to respond to the received verbal request.
[0004] These and other objects, features and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in connection with the accompanying drawings, in which various features of the drawings are not to scale as they are for clarity purposes to facilitate understanding of the invention by those skilled in the art in connection with the detailed description. [Brief explanation of the drawings]
[0005] [Figure 1] FIG. 1 illustrates a networked computer environment in accordance with at least one embodiment. [Figure 2] 1 is an operational flowchart illustrating a process for personalized voice responses according to at least one embodiment. [Figure 3] FIG. 1 is a block diagram of a personalized voice response program according to at least one embodiment. [Figure 4] FIG. 2 is a block diagram of internal and external components of the computer and server depicted in FIG. 1 according to at least one embodiment. [Figure 5] 2 is a block diagram of an exemplary cloud computing environment including the computer system depicted in FIG. 1, according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a block diagram of functional layers of the example cloud computing environment of FIG. 5 in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0006] Although detailed embodiments of the claimed structures and methods are disclosed herein, it should be understood that the disclosed embodiments are merely exemplary of the claimed structures and methods, which may be embodied in various forms. However, the present invention may be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0007] The present invention may be an integratable system, method, or computer program product, or a combination thereof, at any level of technical detail. The computer program product may include a computer-readable storage medium having stored thereon computer-readable program instructions for causing a processor to carry out aspects of the present invention.
[0008] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, by way of example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROMs, DVDs, memory sticks, floppy disks, mechanically encoded devices having instructions recorded on punch cards or ridge-in-groove structures, or the like, and suitable combinations thereof. Computer-readable storage devices, as used herein, should not be construed as ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over wires.
[0009] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network may be comprised of copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0010] Computer-readable program instructions for carrying out the operations of the present invention may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language and similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, as a standalone software package, or partially on the user's computer. Alternatively, the computer may be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the computer-readable program instructions in order to carry out aspects of the present invention.
[0011] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0012] These computer-readable program instructions can be provided to a general-purpose computer, a processor of a special-purpose computer, or other programmable data processing apparatus to create a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable storage media can also be stored in computer-readable storage media connectable to a computer, programmable data processing apparatus, or other device, or combination thereof, that functions in a particular way, such that the computer-readable program instructions stored therein configure one of the products containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0013] Computer-readable program instructions, such as instructions to perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams on a computer, other programmable apparatus, or other device, can also be loaded into a computer, other programmable apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device to generate a computer-implemented process.
[0014] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of executable implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which constitute one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
[0015] The exemplary embodiments described below provide systems, methods, and program products for personalized voice responses. As such, the embodiments have the ability to advance the field of data management and analysis by utilizing a user's personalized vocabulary to formulate responses to voice requests. More specifically, the present invention may include collecting a plurality of user data from Internet of Things (IoT)-connected sensors. The present invention may include identifying a personalized vocabulary based on the collected plurality of user data. The present invention may include training a voice response system based on the collected plurality of user data and the identified personalized vocabulary. The present invention may include receiving a verbal request. The present invention may include utilizing the trained voice response system to respond to the received verbal request.
[0016] As previously mentioned, humans may have unique ways of interpreting spoken language, which may be based on applied emphasis and / or context given personal experiences and / or events, or may be based on cultural and / or demographic influences. Human language may be further based on, among other things, pragmatics (e.g., situational context), syntax (e.g., placement of words and phrases in clauses and sentences or both), morphology (e.g., word function, whether grammatical or lexical or both), semantics (e.g., word meaning), phonology (e.g., classification and / or study of sounds), and phonetics (e.g., study and / or classification of sounds). An artificial intelligence (AI) voice response system can analyze a user's voice request and respond to the user accordingly using a pre-programmed response generator.
[0017] In many cases, the user may not understand the keywords and / or phrases used by the AI voice response system. Similarly, the AI voice response system may not understand the keywords and / or phrases used by the user. The keywords and / or phrases may be, among others, homonyms, homographs, place names, smells, lengths, local language corpora, contextual situations, trending words and / or phrases, words or phrases that are official and / or scientific, or words of complex vocabulary, or combinations thereof.
[0018] Therefore, it may be advantageous, among other things, to enable an AI voice response system to utilize a user's personalized vocabulary while responding to the user's voice requests.
[0019] According to at least one embodiment, the present invention can utilize a user's personalized vocabulary to generate responses to received verbal requests.
[0020] According to at least one embodiment, an artificial intelligence (AI) voice response system may analyze one or more characteristics of unknown content in a received verbal request and may verify the unknown content against relatively known content that may be included in the user's personalized vocabulary.
[0021] Referring to FIG. 1, an exemplary networked computing environment 100 according to one embodiment is depicted. The networked computing environment 100 may include a computer 102 having a processor 104 and data storage device 106 capable of executing a software program 108 and a personalized voice response program 110a. The networked computing environment 100 may include a server 112 capable of executing a personalized voice response program 110b, which may interact with a database 114 and a communications network 116. The networked computing environment 100 may include multiple computers 102 and servers 112, only one of which is shown. The communications network 116 may include various types of communications networks, such as a wide area network (WAN), a local area network (LAN), a telecommunications network, a wireless network, a public switched network, or a satellite network, or combinations thereof. It should be understood that FIG. 1 provides only an illustration of one implementation and does not suggest any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made based on design and implementation requirements.
[0022] The client computer 102 may communicate with the server computer 112 via a communications network 116. The communications network 116 may include connections such as wired, wireless communication links, or fiber optic cables. As described with reference to FIG. 4 , the server computer 112 may include internal components 902a and external components 904a, respectively, and the client computer 102 may include internal components 902b and external components 904b, respectively. The server computer 112 may also operate in a cloud computing service model such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). The server 112 may also be located in a cloud computing deployment model such as a private cloud, a community cloud, a public cloud, or a hybrid cloud. The client computer 102 may be, for example, a mobile device, a phone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of running programs, accessing a network, and accessing a database 114. According to various implementations of this embodiment, the personalized voice response programs 110a, 110b may interact with a database 114 that may be embodied in various storage devices, such as, but not limited to, the computer / mobile device 102, a networked server 112, or a cloud storage service.
[0023] According to this embodiment, a user using the client computer 102 or the server computer 112 can use the personalized voice response program 110a, 110b (respectively) to utilize the user's personalized vocabulary to formulate responses to voice requests. The personalized voice response method is described in more detail below with respect to Figures 2 and 3.
[0024] Referring now to FIG. 2, an operational flowchart illustrating an exemplary personalized voice response process 200 used by personalized voice response programs 110a and 110b according to at least one embodiment is depicted.
[0025] At 202, user data is collected. Internet of Things (IoT)-connected sensors (e.g., those embedded in mobile devices, smartphones, smartwatches, home appliances, vehicles, etc.) can be used to collect various information about the user, including, but not limited to, user movement patterns, user movement locations, user preferences, user requests (i.e., requests sent by the user), user responses (i.e., responses received from the user), user identification information, and user activities (i.e., activities performed by the user). IoT-connected sensors (e.g., temperature sensors, pressure sensors, proximity sensors, optical sensors, smoke sensors, etc.) may identify details about the user, and devices housing the IoT-connected sensors may store the identified details in a data store (e.g., a repository for storing and managing collections of data).
[0026] The identified details may include, but are not limited to, steps taken by the user to perform activities such as driving a vehicle and cooking dinner. For example, IoT-connected sensors embedded within the user's vehicle may identify the user's actions and habits as they relate to how the user drives the vehicle (e.g., turning on a turn signal one mile before a turn, using the horn to beep at a vehicle entering the roadway 100 feet ahead of the user's vehicle when that vehicle enters the roadway perpendicular to the user's vehicle).
[0027] A personalized vocabulary is identified in the collected user data at 204. The personalized voice response program 110a, 110b (i.e., the artificial intelligence voice response system) may collect user data as described above with respect to step 202 above, and may identify activities performed by the user, any requests submitted by the user, and any responses to requests received by the user, etc.
[0028] A user's language patterns, vocabulary, or topics, or a combination thereof, may be analyzed by integrating with various external products, including but not limited to social media tools, IoT devices, reading applications, or other systems that allow a user to speak, read, or write content, or a combination thereof. Analyzing communications that a user creates, engages with, or consumes, or a combination thereof, may enable the personalized voice response program 110a, 110b to build a personalized vocabulary for the user.
[0029] A user's personalized vocabulary may also be identified through a contextual understanding of the user's experiences. For example, personalized vocabulary may be identified based on integration with a calendar system, a global positioning system (GPS), or a health monitoring system (e.g., a system capable of making health-related observations and remotely transmitting health-related data), or a combination thereof.
[0030] A user's IoT-connected sensors can collect device content read by the user using data from integrations that may include, but are not limited to, web browsers, web or mobile applications or combinations thereof, and social media plug-ins.
[0031] User-specified topic information may be identified using contextual analysis via Latent Dirichlet Allocation (LDA), a generative statistical model in natural language processing (NLP) that can map a set of observations to dynamic topics when language patterns are similar. LDA may be a method that enables identification of topics within a document and mapping the document to the identified topics. As here, when voice data is mapped to text by the personalized voice response program 110a, 110b, the mapped text may be classified into dynamic topics.
[0032] A user's social network contributions, including text and image information (e.g., shared images, tagged images, etc.), may be analyzed to determine contextual significance and further to identify the user's personalized vocabulary. A user's social network profile may be accessed, for example, via authentication on a connected application programming interface (API) login of the social network account.
[0033] The user's context (i.e., the voice response) delivered by the voice response may also be analyzed in identifying personalized vocabulary by using a tone analyzer API (e.g., the Watson™ Tone Analyzer API). The Tone Analyzer API can collect voice and tone patterns. For example, a Tone Analyzer API, such as the Watson™ Tone Analyzer API (Watson and all Watson-based trademarks are trademarks or registered trademarks of International Business Machines Corporation in the United States and / or other countries), may utilize a database of historical information including past interactions between the user and the personalized voice response program 110a, 110b to determine whether the voice response portrays an intense, light-hearted, serious, whimsical, or witty tone, among many other voice tones.
[0034] The personalized vocabulary may be used to predict content that may be known and content that may be unknown to the user based on trained modules, as described with respect to step 206 below.
[0035] A personalized vocabulary may be identified for each user of the personalized voice response program 110a, 110b. The user's personalized vocabulary may be stored, inter alia, in a cloud environment, in a user profile associated with the personalized voice response program 110a, 110b, cached on the user's device, or both, and may be accessed using authenticated credentials (e.g., set by the user upon starting the personalized voice response program 110a, 110b). There may be a separate knowledge corpus for each user of the personalized voice response program 110a, 110b, each user being identified by the user's voice profile.
[0036] At 206, the personalized voice response programs 110a, 110b are trained based on the collected user data. A bidirectional long-short-term memory (Bi-LSTM) training module with a text and image information-based convolutional neural network (TI-CNN) module may comprise an artificial intelligence voice response system (i.e., the personalized voice response programs 110a, 110b) trained based on the collected user data. The Bi-LSTM training module may employ a recurrent neural network (RNN) architecture (e.g., a deep learning module) in which signals propagate both backward and forward (e.g., data may be read from beginning to end and from end to beginning), which may enable faster learning than a unidirectional approach. The Bi-LSTM training module may continuously collect user data and, together with the text and image information-based convolutional neural network (TI-CNN) module, may use pattern analysis to contextually identify, classify, and learn.
[0037] The LSTM training module can take into account the context of other words in the text or speech or both, which can be useful in text or speech analysis. The forward LSTM training module can predict the next utterance in a voice response conversation, allowing the forward LSTM training module to better predict the user's needs. The backward LSTM training module can predict previous portions in a voice response conversation, allowing the backward LSTM training module to have more context surrounding the user's current question. The Bi-LSTM can do both (e.g., predict the user's needs and consider past portions of the user's conversation), thereby better enabling both the user's future intent and the context of the user's conversation.
[0038] At 208, a verbal request is received. The personalized voice response program 110a, 110b may decompose the received verbal request (e.g., by parsing through the received verbal request, particularly to determine semantic structure) as well as content delivered by previous verbal replies (i.e., previous verbal responses) as described above with respect to step 204 above to form a verbal response to the received verbal request. The personalized voice response program 110a, 110b may analyze unknown content of the verbal response to be delivered (e.g., portions of the verbal request that require a personalized response by the user) and may identify similar content in the user's personalized vocabulary that can be used to respond to the received verbal request. Here, the personalized voice response program 110a, 110b may form a response to the verbal request based on one or more identified patterns in the user's personalized vocabulary.
[0039] The personalized voice response program 110a, 110b may use K-means clustering to perform a similarity analysis using the user's personalized vocabulary. K-means clustering may be an unsupervised machine learning algorithm that can cluster data together based on determined similarities. For example, a received user request may be related to the current weather, and the personalized voice response program 110a, 110b may determine how to convey "It's going to be cold in the morning" based on the user's personalized vocabulary (e.g., use the word "cold" instead of "chilly" and the word "morning" instead of "AM").
[0040] If similar and / or comparative content is not available in the user's personalized vocabulary, the personalized voice response program 110a, 110b may employ external assistance and / or may search the personalized vocabulary of other users found using local stores or databases or both (e.g., similar users, or users located nearby based on determined GPS location, etc., or both) and / or may apply to a particular user, or a cultural or regional subset of the local population, by leveraging a regional language corpus of integrated vocabulary, phrases, or voice personalization, or a combination thereof.
[0041] Semantic similarity can be estimated by defining topological similarity through the use of ontologies to define distances between terms and / or concepts. Ontologies may be used to formalize groups such as definitions, categories, properties, and entities to better group both data-in and data-out. For example, a naive metric for comparing concepts ordered in a partially ordered set and represented as nodes in a directed acyclic graph (e.g., a taxonomy) may be the shortest path connecting two concept nodes.
[0042] IBM's Watson™ Ground Truth Editor for document stemming (Watson and all Watson-based trademarks are trademarks or registered trademarks of International Business Machines Corporation in the United States and / or other countries) may also be used to establish user-understandable ontologies, potentially enhancing the understanding of user content to improve the use of synonyms and correlation of unknown content with the user's knowledge space.
[0043] Document annotation and classification for a target domain through IBM's Watson™ Ground Truth Editor may involve manually classifying and / or annotating portions of the training and / or provided data for the personalized voice response program 110a, 110b. The Watson™ Ground Truth Editor may obtain ground truth, or a collection of vetted data that can be used to adapt Watson™ to a particular domain. A human user can assist in the classification of ground truth.
[0044] If the similar and / or comparative content is not in the user's personalized vocabulary or in the personalized vocabulary of another (e.g., similar) user, the personalized voice response program 110a, 110b may refer to the individual from whom the verbal request was made (e.g., by a verbal request for information from the original requester) and further elaborate on the verbal request.
[0045] For example, the personalized voice response program 110a, 110b can verbally query the original requester whether the user is referring to a type of "wrapping" (e.g., wrapping paper) or a music genre (e.g., rap music).
[0046] The voice response system responds to the received verbal request at 210. The personalized voice response program 110a, 110b may construct a response to the received verbal request by using information gathered from the personalized vocabulary and / or the personalized vocabulary of similar users and may deliver the same verbally to the user.
[0047] In responding to a received verbal request, the personalized voice response program 110a, 110b may analyze historical information about the user (e.g., travel locations, activities performed, content read, known words, etc.) to identify knowledge possessed by the user. In doing so, the personalized voice response program 110a, 110b may utilize the user's identified personalized vocabulary to generate a verbal response to the received verbal request that may be understandable by the user.
[0048] Referring now to FIG. 3, a block diagram 300 of the personalized voice response programs 110a and 110b is depicted, according to at least one embodiment. A user speaks a voice request into an artificial intelligence voice response device 302 equipped with an embedded microphone. The voice request received by the artificial intelligence voice response device 302 may be connected to the personalized voice response programs 110a, 110b, as depicted by 304, which can process the received voice request. The personalized voice response programs 110a, 110b, as depicted by 304, may be capable of implementing natural language processing (NLP) and natural language understanding (NLU) techniques, which may be used to analyze collected data (e.g., the spoken voice request). Other data collected by the artificial intelligence voice response device 302 and / or other connected Internet of Things (IoT) devices may be stored in a data collection module of the personalized voice response programs 110a, 110b, as depicted by 304.
[0049] An embedded deep neural network, also depicted by 304, within the personalized voice response program 110a, 110b can utilize a training module to analyze the user's previous voice responses, and optionally those of similar users, to generate a personalized response to the user's request. As depicted by 304, the corpus of data included in the personalized voice response program 110a, 110b can be the user's personalized linguistic corpus (i.e., personalized vocabulary), which can include pragmatics, syntax, morphology, semantics, phonology, or phonetics, or a combination thereof, specific to a given user.
[0050] Optionally, the two-way feed may be utilized by the personalized voice response program 110a, 110b depicted by 304 to access regional and / or cultural linguistic corpora that do not appear in the user's personalized corpus but that may contain language data that may be useful in responding to received voice requests.
[0051] 2 and 3 provide only an illustration of one embodiment and do not imply any limitations on how different embodiments may be implemented. Many modifications to the depicted embodiment may be made based on design and implementation requirements.
[0052] According to at least one alternative embodiment, an iterative feedback learning module may be included in the personalized voice response program, which may enable monitoring of a user's reactions to responses provided by the personalized voice response program via a close-up camera and / or IoT sensors, which may assist in determining the user's satisfaction. For example, the iterative feedback learning module may employ face scanning technology and / or a tone analyzer to look for indicators of the user's satisfaction, which may range from confusion and / or frustration to happiness. Depending on the content and satisfaction, the iterative feedback learning module may refine and / or modify a corpus of user data to improve the accuracy of the user's personalized vocabulary, improve the accuracy of responses to received verbal requests, and optimize the user's satisfaction with the personalized voice response program.
[0053] Figure 4 is a block diagram 900 of the internal and external components of the computer depicted in Figure 1 in accordance with an exemplary embodiment of the present invention. It should be understood that Figure 4 provides only an illustration of one implementation and is not intended to suggest any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environments may be made based on design and implementation requirements.
[0054] Data processing systems 902, 904 are representative of any electronic device capable of executing machine-readable program instructions. Data processing systems 902, 904 may be representative of a smartphone, computer system, PDA, or other electronic device. Examples of computing systems, environments, or configurations, or combinations thereof, that may be represented by data processing systems 902, 904 include, but are not limited to, personal computer systems, server computer systems, thin client, thick client, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, network PCs, minicomputer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0055] The user client computer 102 and the network server 112 may include respective sets of internal components 902a, b and external components 904a, b illustrated in Figure 4. Each set of internal components 902a, b includes one or more processors 906, one or more computer-readable RAMs 908 and one or more computer-readable ROMs 910 on one or more buses 912, and one or more operating systems 914 and one or more computer-readable tangible storage devices 916. The one or more operating systems 914, software programs 108, and personalized voice response program 110a in the client computer 102 and the personalized voice response program 110b in the network server 112 may be stored on one or more computer-readable tangible storage devices 916 for execution by the one or more processors 906 via one or more RAMs 908 (which typically include cache memory). In the embodiment shown in Figure 4, each of the computer-readable tangible storage devices 916 is a magnetic disk storage device of an internal hard drive. Alternatively, each of the computer-readable tangible storage devices 916 is a semiconductor storage device such as a ROM 910, an EPROM, a flash memory, or any other computer-readable tangible storage device capable of storing computer programs and digital information.
[0056] Each set of internal components 902a,b also includes a R / W drive or interface 918 for reading from and writing to one or more portable computer-readable tangible storage devices 920, such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor storage devices. Software programs, such as software program 108 and personalized voice response programs 110a, 110b, can be stored on one or more of the respective portable computer-readable tangible storage devices 920 and can be read via the respective R / W drive or interface 918 and loaded onto the respective hard drive 916.
[0057] Each set of internal components 902a, b may also include a network adapter (or switch port card) or interface 922, such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G or 4G wireless interface card, or other wired or wireless communication link. The software program 108 and personalized voice response program 110a in the client computer 102 and the personalized voice response program 110b in the network server computer 112 may be downloaded from an external computer (e.g., a server) via a network (e.g., the Internet, a local area network, or other wide area network) and their respective network adapters or interfaces 922. From the network adapters (or switch port adapters) or interfaces 922, the software program 108 and personalized voice response program 110a in the client computer 102 and the personalized voice response program 110b in the network server computer 112 are loaded onto their respective hard drives 916. The network may include copper wire, optical fiber, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof.
[0058] Each of the set of external components 904 a,b can include a computer display monitor 924, a keyboard 926, and a computer mouse 928. The external components 904 a,b can also include touch screens, virtual keyboards, touchpads, pointing devices, and other human interface devices. Each of the set of internal components 902 a,b also includes a device driver 930 for interfacing to the computer display monitor 924, the keyboard 926, and the computer mouse 928. The device driver 930, the R / W drive or interface 918, and the network adapter or interface 922 include hardware and software (stored in the storage device 916 or the ROM 910, or both).
[0059] Although this disclosure includes detailed descriptions related to cloud computing, it is understood in advance that implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the present invention may be practiced in conjunction with any other type of computing environment now known or later developed.
[0060] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.
[0061] The characteristics are as follows: On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time or network storage, automatically as needed, without the need for human interaction with the service provider. Broad network access: Computing power is available over the network and can be accessed through standard mechanisms, facilitating use by heterogeneous thin or thick client platforms (e.g., cell phones, laptops, PDAs). Resource Pooling: Computing resources from a provider are pooled and offered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated based on demand. Consumers generally have no control or knowledge of the exact location of the resources they are provided with, resulting in a sense of location independence. However, consumers may be able to determine location at a higher level of abstraction (e.g., country, state, data center). Rapid Elasticity: Computing capacity can be provisioned quickly and elastically, sometimes automatically, to instantly scale out and quickly release to instantly scale in. To the consumer, the computing power available for provisioning often appears unlimited, and can be purchased at any time and in any quantity. Metered Services: Cloud systems leverage measurement capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported to provide transparency to both providers and consumers of utilized services.
[0062] The service model is as follows: Software as a Service (SaaS): The functionality offered to the consumer is the availability of a provider's applications running on a cloud infrastructure that can be accessed from a variety of client devices through a thin client interface such as a web browser (e.g., webmail). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functionality, except for limited user-specific application configuration settings. Platform as a Service (PaaS): The capability offered to consumers is to deploy applications they create or acquire using programming languages and tools supported by the provider onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of their hosting environment. Infrastructure as a Service (IaaS): The functionality offered to consumers is the provisioning of processors, storage, networking, and other basic computing resources on which they can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating system, storage, and deployed applications, and in some cases partial control over some network components (e.g., host firewalls).
[0063] The deployment model is as follows: Private Cloud: This cloud infrastructure is dedicated to a specific organization and can be managed by that organization or a third party, and can exist on-premise or off-premise. Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common interests (e.g., mission, security requirements, policy, and compliance considerations). This cloud infrastructure can be managed by those organizations or a third party and can exist on-premises or off-premises. Public cloud: This cloud infrastructure is available to the general public or large industry organizations and is owned by an organization that sells cloud services. Hybrid cloud: This cloud infrastructure combines two or more cloud models (private, community, or public), each of which retains its inherent nuances but is bound by standards or specific technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0064] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0065] Referring now to FIG. 5 , an exemplary cloud computing environment 1000 is depicted. As shown, the cloud computing environment 1000 includes one or more cloud computing nodes 100, to which local computing devices used by cloud consumers (e.g., PDAs or cell phones 1000A, desktop computers 1000B, laptop computers 1000C, or automobile computer systems 1000N, or combinations thereof) can communicate. The nodes 100 can communicate with each other. The nodes 100 can be physically or virtually grouped (not shown) in one or more networks, such as, for example, private, community, public, or hybrid clouds, or combinations thereof, as described above. This enables the cloud computing environment 1000 to provide infrastructure, platform, or software as a service, or combinations thereof, for which cloud consumers do not need to maintain resources on their local computing devices. It should be understood that the types of computing devices 1000A-N shown in FIG. 5 are for illustrative purposes only, and that the computing node 100 and cloud computing environment 1000 can communicate with any type of electronic device via any type of network or network-addressable connection (e.g., using a web browser) or both.
[0066] Referring now to Figure 6, there is shown a set of functional abstraction model layers 1100 provided by the cloud computing environment 1000. It should be understood in advance that the components, layers, and functions shown in Figure 6 are merely exemplary, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0067] Hardware and software layer 1102 includes hardware and software components. Examples of hardware components include mainframe 1104, reduced instruction set computer (RISC) architecture-based server 1106, server 1108, blade server 1110, storage device 1112, and network and network components 1114. In some embodiments, software components include network application server software 1116 and database software 1118.
[0068] The virtualization layer 1120 provides an abstraction layer from which the following virtual entities can be provided, for example: virtual servers 1122, virtual storage 1124, virtual networks including virtual private networks 1126, virtual applications and operating systems 1128, and virtual clients 1130.
[0069] By way of example, management layer 1132 may provide the following functionality: Resource provisioning 1134 enables dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 1136 enables cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. By way of example, these resources may include application software licenses. Security enables identification and verification of cloud consumers and tasks, as well as protection for data and other resources. User portal 1138 provides consumers and system administrators with access to the cloud computing environment. Service level management 1140 enables allocation and management of cloud computing resources so that requested service levels are met. Service level agreement (SLA) planning and fulfillment 1142 enables advance arrangement and procurement of anticipated future cloud computing resources required according to SLAs.
[0070] The workload layer 1144 provides examples of functionality available to a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include mapping and navigation 1146, software development and lifecycle management 1148, virtual classroom instruction delivery 1150, data analytics processing 1152, transaction processing 1154, and personalized voice response 1156. The personalized voice response programs 110a, 110b provide a way to create responses to voice requests using a user's personalized vocabulary.
[0071] The description of various embodiments of the present invention is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications or technical improvements to technology found in the market, or to enable those skilled in the art to understand the embodiments described herein.
Claims
1. A method for personalized voice response by a computer, said method comprising: The computer collects a plurality of user data from Internet of Things (IoT) connected sensors; identifying a personalized vocabulary based on the collected plurality of user data; training a voice response system based on the collected plurality of user data and the identified personalized vocabulary; receiving a verbal request by the computer; responding to the received verbal request using the trained voice response system and personalized vocabularies of similar or nearby users, the similar or nearby users being identified using a regional language corpus employing a lexical similarity ontology; A method comprising:
2. The method of claim 1 , wherein the collected plurality of user data is selected from the group consisting of user movement patterns, user movement locations, user preferences, user requests, user responses, user identification information, and user activities.
3. The computer identifying the personalized vocabulary based on the collected plurality of user data includes: The computer uses Latent Dirichlet Allocation (LDA) contextual analysis to identify topic information from the collected user data; determining, by the computer, a contextual significance of the collected plurality of user data based on social network contributions; The method of claim 1 further comprising:
4. The method of claim 1, wherein the computer's training of the voice response system based on the collected plurality of user data and the identified personalized vocabulary further includes a bidirectional long short-term memory (Bi-LSTM) training module having a text and image information-based convolutional neural network (TI-CNN) module.
5. The computer receiving the verbal request comprises: said computer decomposing said received verbal request; the computer using K-means clustering to identify content from the personalized vocabulary that is similar to the received verbal request; The method of claim 1 further comprising:
6. The computer utilizes a regional language corpus of integrated vocabulary, phrases, and voice personalization to find content similar to the received verbal request. The method of claim 5 further comprising:
7. The computer responding to the received verbal request using the trained voice response system comprises: The computer uses the identified content to construct a response to the received verbal request. The method of claim 5 further comprising:
8. 1. A computer system for personalized voice response, comprising: a computer system including one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more computer-readable tangible storage media to be executed by at least one of the one or more processors via at least one of the one or more computer-readable memories, the computer system being capable of performing a method, the method comprising: Collecting a plurality of user data from Internet of Things (IoT) connected sensors; identifying a personalized vocabulary based on the collected plurality of user data; training a voice response system based on the collected plurality of user data and the identified personalized vocabulary; receiving a verbal request; Responding to the received verbal request using the trained voice response system and personalized vocabulary of similar or nearby users, where the similar or nearby users are identified using a regional language corpus employing a lexical similarity ontology; 2. A computer system comprising:
9. 9. The computer system of claim 8, wherein the collected plurality of user data is selected from the group consisting of user movement patterns, user movement locations, user preferences, user requests, user responses, user identification information, and user activities.
10. identifying the personalized vocabulary based on the collected plurality of user data utilizing Latent Dirichlet Allocation (LDA) contextual analysis to identify topic information from the collected user data; determining a contextual significance of the collected plurality of user data based on social network contributions; The computer system of claim 8 further comprising:
11. 10. The computer system of claim 8, wherein training the voice response system based on the collected plurality of user data and the identified personalized vocabulary further comprises a bidirectional long short-term memory (Bi-LSTM) training module having a text and image information-based convolutional neural network (TI-CNN) module.
12. receiving the verbal request decomposing the received verbal request; using K-means clustering to identify content from the personalized vocabulary that is similar to the received verbal request; The computer system of claim 8 further comprising:
13. Leveraging an integrated regional language corpus of lexical, phrasal, and speech personalization to find content similar to the received verbal request. The computer system of claim 12 further comprising:
14. utilizing the trained voice response system to respond to the received verbal request, Utilizing the identified content to construct a response to the received verbal request. The computer system of claim 12 further comprising:
15. 1. A computer program for personalized voice response, comprising: Collecting a plurality of user data from Internet of Things (IoT) connected sensors; identifying a personalized vocabulary based on the collected plurality of user data; training a voice response system based on the collected plurality of user data and the identified personalized vocabulary; receiving a verbal request; Responding to the received verbal request using the trained voice response system and personalized vocabulary of similar or nearby users, where the similar or nearby users are identified using a regional language corpus employing a lexical similarity ontology; A computer program that causes a computer to execute the following.
16. 16. The computer program product of claim 15, wherein the collected plurality of user data is selected from the group consisting of user movement patterns, user movement locations, user preferences, user requests, user responses, user identification information, and user activities.
17. identifying the personalized vocabulary based on the collected plurality of user data utilizing Latent Dirichlet Allocation (LDA) contextual analysis to identify topic information from the collected user data; determining a contextual significance of the collected plurality of user data based on social network contributions; 16. The computer program of claim 15, further comprising:
18. 16. The computer program product of claim 15, wherein training the voice response system based on the collected plurality of user data and the identified personalized vocabulary further comprises a bidirectional long short-term memory (Bi-LSTM) training module having a text and image information-based convolutional neural network (TI-CNN) module.
19. receiving the verbal request decomposing the received verbal request; using K-means clustering to identify content from the personalized vocabulary that is similar to the received verbal request; 16. The computer program of claim 15, further comprising:
20. Leveraging an integrated regional language corpus of lexical, phrasal, and speech personalization to find content similar to the received verbal request.
20. The computer program of claim 19, further comprising:
Citation Information
Patent Citations
Recommendation system using social behavior analysis and lexical classification
JP2011513802A
Interaction processing program, interaction processing method, and information processing apparatus
JP2017203808A
Automatic response suggestions for received images in messages using language models
JP2019536135A
Virtual assistant for generating personalized responses within a communication session
US20190005021A1