Authentication of packetized audio signals
The data processing system authenticates packetized audio signals to manage network traffic, reducing bandwidth and processor load by disabling harmful transmissions and optimizing communication sessions, enhancing network efficiency and security.
Patent Information
- Application Number
- DE112017000177
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-12-30
- Filing Date
- 2017-08-31
- Publication Date
- 2026-01-22
- Estimated Expiration
- 2037-08-31
AI Technical Summary
Excessive packet-based network traffic data transmissions can lead to inefficient bandwidth usage, processor overload, and potential malicious activity, complicating data routing and degrading response quality in computer networks.
A data processing system authenticates packetized audio signals by identifying requests and trigger keywords, analyzing signal properties, and establishing alarm states to manage communication sessions, thereby reducing harmful transmissions and optimizing network bandwidth and processor utilization.
The system enhances network efficiency by preventing harmful audio signal processing, reducing bandwidth usage, and conserving power by disabling unauthorized communication sessions, thus improving computing performance and security.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Excessive packet-based or other network traffic data transmissions between computer devices can prevent a computer device from properly processing the network traffic data, completing an operation associated with the network traffic data, or responding to the network traffic data in a timely manner. Excessive network traffic data transmissions can also complicate data routing or degrade the quality of the response if the responding computer device reaches or exceeds its processing capacity, potentially resulting in inefficient bandwidth usage. Some of the excessive network transmissions may include malicious network transmissions.
[0002] US 2014 / 0249817A1 describes techniques for using speaker identification information and other features associated with received voice commands to determine how and whether to respond to those commands. A user can interact with a device by voice, by entering voice commands. After an interaction with the user begins, the device can detect subsequent speech, which may originate from the user, another user, or another source. The device can then use speaker identification information and other features associated with the speech to attempt to determine whether the user interacting with the device uttered the speech. The device can then interpret the speech as a valid voice command and, in response to the finding that the user did indeed utter the speech, perform an appropriate operation.However, if the device determines that the user did not make the speech, it may not take any action regarding the speech.
[0003] US 2005 / 0 185 779 A1 describes a system for the automatic detection of fraudulent activity in a transaction network, where each transaction on the network is assigned an identifier. In one embodiment of US 2005 / 0 185 779 A1, the system includes a voice comparison device for comparing an initial sampled voice of a user in a first transaction with a subsequent sampled voice of a user in a subsequent transaction that has the same identifier as the first transaction. A voice-based fraud detection engine is provided to derive a user usage profile from the comparison that is representative of the total number of different users of the associated identifier.
[0004] US 2016 / 0093304A1 describes systems and processes for generating a speaker profile for use in speaker identification for a virtual assistant. An example process might involve receiving audio input containing user speech and determining whether a speaker of that user speech is a designated user, based on a speaker profile for that designated user. If the user speech is found to be the designated user, the user speech can be added to the speaker profile, and the virtual assistant can be triggered. If the user speech is found to be not the designated user, the user speech can be added to an alternative speaker profile, and the virtual assistant might not be triggered.In some examples, contextual information can be used to verify the results generated by the speaker identification process.
[0005] US 2015 / 0371639A1 refers to methods, systems, and devices, including computer programs encoded on a computer storage medium, for a dynamic threshold for speaker verification.
[0006] US 2015 / 0 142 438 A1 describes a speech recognition method for use in an electronic device that includes a speech input module. The method includes: receiving speech data by the speech input module, performing an initial pattern speech recognition on the received speech data, including determining whether the speech data contains initial speech recognition information, performing a second pattern speech recognition on the speech data if the speech data contains the initial speech recognition information, and performing or rejecting an operation corresponding to the initial speech recognition information based on the result of the second pattern speech recognition. SUMMARY
[0007] The present disclosure relates generally to the authentication of packetized audio signals in a speech-enabled computer network environment to reduce the amount of excessive network transmissions. A natural language processor component executed by a data processing system can receive data packets. The data packets may contain an input audio signal detected by a sensor of a client computer device. The natural language processor component can parse the input audio signal to identify a request and a trigger keyword according to the request. A network security device can analyze one or more properties of the input audio signal. Based on these properties, the network security device can establish an alarm state. The network security device can then notify a content selection component of the data processing system of the alarm state.Based on the alarm state, the content selection component can select a content item via a real-time content selection process. An audio signal generator component, executed by the data processing system, can include an output signal that comprises the content item. An interface of the data processing system can transmit data packets containing the output signal generated by the audio signal generator component to cause an audio driver component, executed by the client computer, to drive a speaker on the client computer to generate an acoustic wave corresponding to the output signal. The data processing system can receive a response audio signal. The response audio signal is received in reaction to the output signal generated by the client computer. The response audio signal can contain characteristics that are analyzed by the network security device.Based on the characteristics of the response audio signal, the network security device can terminate or suspend a communication session between a service provider and a client computer device.
[0008] According to one aspect of the disclosure, a system for authenticating packetized audio signals in a voice-activated computer network environment is disclosed according to claims 1 to 15.
[0009] According to another aspect of the disclosure, a method for authenticating packetized audio signals in a voice-activated computer network environment is disclosed according to claims 16 to 18.
[0010] According to one aspect of the disclosure, a system for authenticating packetized audio signals in a voice-activated computer network environment is disclosed according to claims 19 to 20.
[0011] These and other aspects and implementations are explained in more detail below. The preceding information and the following detailed description include illustrative examples of various aspects and implementations and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The drawings provide an illustration and further understanding of the various aspects and implementations and are incorporated into and form part of this specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are not to scale. Identical reference symbols and labels in the various drawings refer to similar elements. For clarity, not every component may be labeled in every drawing. The drawings include: Fig. Figure 1 presents an exemplary system for executing packetized audio signals in a speech-activated packet-based (or other protocol-based) computer network environment; Fig. Figure 2 illustrates a flowchart that demonstrates an exemplary operation of a system for performing authentication of packetized audio signals; Fig. Figure 3 illustrates an exemplary procedure for authenticating packetized audio signals in a voice-activated packet-based (or other protocol-based) computer network environment using the method described in Fig. 1 illustrated system; and Fig. Figure 4 shows a block diagram illustrating a general architecture for a computer system that can be used to implement elements of the systems and procedures described and illustrated herein. DETAILED DESCRIPTION
[0013] More detailed descriptions of various concepts relating to processes, devices, and systems and their implementations follow. The various concepts introduced above and explained in more detail below can be implemented in any of numerous ways.
[0014] The present disclosure relates generally to a data processing system for authenticating packetized audio signals in a speech-activated computer network environment. The data processing system can improve the efficiency and effectiveness of transmitting auditory data packets over one or more computer networks by, for example, disabling harmful transmissions before they are transmitted across the network. The present solution can also improve computing performance by disabling remote computer processes that may be affected or caused by the harmful audio signal transmissions. By disabling the transmission of harmful audio signals, the system can reduce bandwidth usage by preventing the transmission of data packets carrying the harmful audio signal across networks. Processing the naturally spoken audio signal can be a computationally intensive task.By detecting potentially harmful audio signals, the system can reduce processing overhead by skipping or temporarily omitting the processing of such signals. The system can also reduce processing overhead by disabling communication sessions when harmful activity is detected.
[0015] The systems and procedures described herein may include a data processing system that receives an audio input query, also referred to as an audio input signal. From the audio input query, the data processing system can identify a request and a trigger keyword according to the request. The system can generate action data structures based on the audio input query. The system can also measure features of the audio input query. The system can determine whether the features of the audio input query match the predicted or expected properties of the audio input query. If the features do not match the expected properties, the system can select a content element to be transmitted back to the source of the audio input query. A communication session can then be initiated with the source.The content element can include an output signal that can be played back through a speaker associated with the source. The system can receive a response audio signal from the content element. This response audio signal can also contain characteristics that the system compares to expected properties. If the characteristics of the response audio signal do not match the expected properties, the system can disable communication sessions with the source and prevent the source from initiating communication sessions with third parties or content providers, thereby saving network bandwidth, reducing processor utilization, and conserving power.
[0016] This solution can prevent the transmission of unsafe audio-based user interactions by authenticating the interaction. Securing audio-based user interactions prevents malicious processes from running under the user account (or that of another user). Preventing the execution of malicious processes can also reduce network bandwidth and processor utilization or load. This solution can reduce network bandwidth usage by blocking the transmission of unauthorized audio-based user interactions.
[0017] Fig. Figure 1 presents an exemplary System 100 for executing packetized audio signals in a speech-activated packet-based (or other protocol-based) computer network environment. The System 100 can include at least one data processing system 105. The data processing system 105 can include at least one server, which has at least one processor. The data processing system 105 can, for example, include a plurality of servers located in at least one data center or server farm. The data processing system 105 can determine a request and a trigger keyword associated with the request from an audio input signal.Based on the request and the trigger keyword, the data processing system 105 can determine or select a thread containing a variety of sequence-dependent operations and initiate content elements (and other actions as described herein) in an order that does not correspond to the order of dependent operations, for example, as part of a speech-activated communication or scheduling system. The content elements can include one or more audio files that, when played back, provide an audio output or acoustic wave. In addition to audio content, the content elements can also include other content (for example, text, video, or image content).
[0018] The 105 data processing system can include multiple logically grouped servers and support distributed computing processes. The logical group of servers can be referred to as a data center, server farm, or computer farm. The servers can be geographically distributed. A data center or computer farm can be managed as a single entity, or the computer farm can comprise multiple computer farms. The servers in a computer farm can be heterogeneous—one or more of the servers or computers can run on one or more types of operating system platforms. The 105 data processing system can include servers in a data center housed in one or more high-density rack systems, as well as associated storage systems located, for example, in an enterprise data center.The Data Processing System 105 with consolidated servers can improve system administration, data security, physical system security, and system performance by searching for servers and high-performance storage systems across localized, high-performance networks. Centralizing all or some of the Data Processing System 105 components, including servers and storage systems, and connecting them to enhanced system management tools enables more efficient use of server resources, thereby saving power and processing requirements and reducing bandwidth usage.
[0019] The data processing system 105 can include at least one natural language processing (NLP) component 110, at least one interface 115, at least one network security device 123, at least one content selection element component 125, at least one audio signal generator component 130, at least one direct action application programming interface (API) 135, at least one session handling element component 140, at least one communication API 136, and at least one data container 145. The NLP component 110, interface 115, network security device 123, content selection element component 125, audio signal generator component 130, direct action API 135, and session handling element component 140 can each include at least one processing unit, server, virtual server, circuit, machine, agent, device, or other logic device, such as...The programmable arrays are configured to communicate with the data container 145 and with other computer devices (e.g., the client computer device 150, the content provider computer device 155, or the service provider computer device 160) via at least one computer network 165. The network 165 can include computer networks such as the Internet, local area networks, regional networks, and wide area networks or other area networks, intranets, satellite networks, or other computer networks such as voice or data-related mobile networks and combinations thereof.
[0020] Session handler component 140 can, for example, establish a communication session between data processing system 105 and client computer device 150. Session handler component 140 can initiate the communication session based on receiving an input audio signal from computer device 150. Session handler component 140 can set the initial duration of the session communication based on the time of day, the location of client computer device 150, the context of the input audio signal, or a voiceprint. Session handler component 140 can terminate the communication session after the session expires. Authentication may only be required once per communication session.For example, the data processing system 105 can determine that there was a previous successful authentication during the communication session and does not require any additional authentication before the communication session expires.
[0021] A subset of information sources that include or form elements linked to, or selectable from, a content ordering or search engine results system, such that these third-party content elements are included as part of a content ordering campaign. The network 165 can be used by the data processing system 105 to access information resources such as web pages, internet presences, domain names, or URLs that can be presented, output, reproduced, or displayed by the client computer device 150. For example, through the network 165, a user of the client computer device 150 can access information or data provided by the content provider computer device 155 or the service provider computer device 160.
[0022] Network 165 can, for example, include a point-to-point network, a broadcast network, a wide area network, a local area network, a telecommunications network, a data communications network, a computer network, an ATM (Asynchronous Transfer Mode) network, a SONET (Synchronous Optical Network), an SDH (Synchronous Digital Hierarchy) network, a wireless network, or a wired network, and combinations thereof. Network 165 can include a wireless connection, such as an infrared channel or a satellite frequency band. The topology of Network 165 can include a bus, star, or ring network topology.Network 165 can include cellular networks using any protocol or protocols suitable for communication with mobile devices, including Advanced Mobile Phone Protocol (AMPS), Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Global System for Mobile Communication (GSM), General Packet Radio Services (GPRS), and Universal Mobile Telecommunications System (UMTS). Different types of data can be transmitted over different protocols, or the same types of data can be transmitted over different protocols.
[0023] The client computer device 150, the content provider computer device 155, and the service provider computer device 160 may each include at least one logic device, such as a computer device with a processor, for communicating with each other or with the data processing system 105 over the network 165. The client computer device 150, the content provider computer device 155, and the service provider computer device 160 may each include at least one server, processor, or memory, or a plurality of computing resources or servers located in at least one data center. The client computer device 150, the content provider computer device 155, and the service provider computer device 160 may each include at least one computer device, such as a desktop computer, laptop, tablet, personal digital assistant, smartphone, portable computer, thin client computer, virtual server, or other computer device.
[0024] The client computer device 150 can include at least one sensor 151, at least one transducer 152, at least one audio driver 153, and at least one loudspeaker 154. The sensor 151 can include a microphone or an audio input sensor. The sensor 151 can also include at least one GPS sensor, proximity sensor, ambient light sensor, temperature sensor, motion sensor, accelerometer, or gyroscope. The transducer 152 can convert the audio input signal into an electronic signal. The audio driver 153 can include a script or program that is executed by one or more processors of the client computer 150 to control the sensor 151, the transducer 152, or the audio driver 153, among other components of the client computer 150, to process audio input or provide audio output. The loudspeaker 154 can transmit the audio output signal.
[0025] The client computer device 150 can be assigned to an end user who inputs voice queries as audio input into the client computer device 150 (via the sensor 151) and receives audio output in the form of a computer-generated voice, which can be provided to the client computer device 150 by the data processing system 105 (or the content provider computer device 155 or the service provider computer device 160) and is emitted by the speaker 154. The computer-generated voice can include recordings of a real person or computer-generated speech.
[0026] The content provider computer device 155 can provide audio-based content elements for display by the client computer device 150 as an audio output content element. The content element can include an offer for a good or service, such as a voice-based message like, "Would you like me to order a taxi for you?" For example, the content provider computer device 155 can include memory to store a series of audio content elements provided in response to a voice-based request. The content provider computer device 155 can also provide audio-based content elements (or other content elements) to the data processing system 105, where they can be stored in the data container 145.The data processing system 105 can select the audio content elements and deliver them to the client computer device 155 (or instruct the content provider computer device 150 to deliver them). The content may include security questions generated to authenticate the user of the client computer device 150. The audio-based content elements may consist solely of audio or be combined with text, image, or video data.
[0027] The service provider computer device 160 can include at least one service provider natural language processor (NLP) component 161 and at least one service provider interface 162. The service provider NLP component 161 (or other components, such as a direct action API of the service provider computer device 160) can control the client computer device 150 (via or bypassing the data processing system 105) to establish a back-and-forth, real-time speech- or audio-based conversation (e.g., a session) between the client computer device 150 and the service provider computer device 160. The service provider interface 162 can, for example, receive data messages to or send data messages to the direct action API 135 of the data processing system 105. The service provider computer device 160 and the content provider computer device 155 can be associated with the same entity.For example, the service provider computer device 155 can generate, store, or provide content for a ride-sharing service, and the service provider computer device 160 can establish a session with the client computer device 150 to initiate the provision of a taxi or car from the ride-sharing service to pick up the end user of client computer 150. The data processing system 105 can also establish the session with the client computer device, including or bypassing the service provider computer device 160, via the direct action API 135, the NLP component 110, or other components, to initiate, for example, the provision of a taxi or car from the ride-sharing service.
[0028] The service provider device 160, the content provider device 155, and the data processing system 105 can include a conversation API 136. The end user can interact with the content and the data processing system 105 via a voice conversation through a communication session. The voice conversation can take place between the client device 150 and the conversation API 136. The conversation API 136 can be executed by the data processing system 105, the service provider 160, or the content provider 155. The data processing system 105 can directly receive additional information about the end user's interaction with the content when the data processing system executes the conversation API 136.When the service provider 160 or the content provider 155 executes the conversation API 136, the communication session can either be routed through the data processing system 105, or the respective entities can forward data packets of the communication session to the data processing system 105. The network security application described herein can terminate the communication session when the conversation API 136 is executed by the data processing system 105. The network security device 105 can send instructions to the service provider 160 or content provider 155 to terminate (or otherwise disable) the communication session when the service provider 160 or content provider 155 executes the conversation API 136.
[0029] Data container 145 can contain one or more local or distributed databases and can include a database management system. Data container 145 can contain computer data storage or memory and can store one or more parameters 146, one or more policies 147, content data 148, and templates 149 along with other data. Parameters 146, policies 147, and templates 149 can contain information such as rules about a voice-based session between the client computer device 150 and the data processing system 105 (or the service provider computer device 160). Content data 148 can include content elements for audio output or associated metadata, as well as input audio messages that may be part of one or more communication sessions with the client computer device 150.
[0030] The data processing system 105 can include an application, script, or program installed on the client computer device 150, such as an application to communicate input audio signals to the interface 115 of the data processing system 105 and to control components of the client computer device to play back output audio signals. The data processing system 105 can receive data packets or other signals that contain or identify an audio input signal. For example, the data processing system 105 can execute or cause the NLP component 110 to execute in order to receive the audio input signal. The audio input signal can be detected by the sensor 151 (e.g., a microphone) on the client computer device.The NLP component 110 can convert the audio input signal into recognized text by comparing it to a stored representative set of audio waveforms and selecting the closest matches. The representative waveforms can be generated from a large group of input signals. The user can provide some of these input signals. Once the audio signal has been converted into recognized text, the NLP component 110 can match the text with words that are linked, for example, through a learning phase, to actions that the system 200 can perform. The client computer device 150 can provide the audio input signal to the data processing system 105 (e.g., via the network 165) via the converter 152, the audio driver 153, or other components. There, it can be received (e.g., through the interface 115) and provided to the NLP component 110, or stored as content data 148 in the data container 145.
[0031] NLP component 110 can receive the audio input signal. From the input audio signal, NLP component 110 can identify at least one request or at least one trigger keyword that corresponds to the request. The request can indicate the intention or subject of the input audio signal. The trigger keyword can indicate a type of action that is expected to be performed. For example, NLP component 110 can parse the input audio signal to identify at least one request to go out for dinner and to the cinema in the evening. The trigger keyword can contain at least one word, phrase, stem, part of a word, or derivative that indicates an action to be performed. For example, the trigger keyword "go" or "go to" from the input audio signal can indicate a need for transportation.In this example, the input audio signal (or the identified request) does not directly express an intent to transport, but the trigger keyword indicates that a transport is an additional action to at least one other action indicated by the request.
[0032] The content selection element component 125 can obtain this information from the data container 145, where it can be stored as part of the content data 148. The content selection element component 125 can query the data container 145 to select or otherwise identify the content element, for example, from the content data 148. The content selection element component 125 can also select the content element from the content provider computer device 155. For example, the content provider computer device 155, responding to a request from the data processing system 105, can provide a content element to the data processing system 105 (or a component thereof) for later output by the client computer device 150.
[0033] The audio signal generator component 130 can generate or otherwise receive an output signal containing the content element that responds to the third action. For example, the data processing system 105 can execute the audio signal generator component to generate or produce an output signal corresponding to the content element. The interface 115 of the data processing system 105 can provide or transmit one or more data packets containing the output signal to the client computer device 150 over the computer network 165. For example, the data processing system 105 can provide the output signal from the data container 145 or from the audio signal generator component 130 to the client computer device 150. The data processing system 105 can also instruct the content provider computer device 155 or the service provider computer device 160 to provide the output signal to the client computer device 150 via data packet transmissions.The output signal can be received, generated, converted, or transmitted to the client computer device 150 as one or more data packets (or another communication protocol) from the data processing system 105 (or another data processing device).
[0034] The content selection element component 125 can select the content element for the action of the input audio signal as part of a real-time content selection process. For example, the content element can be provided to the client computer device for transmission as plain text audio output as a direct response to the input audio signal. The real-time content selection process for identifying the content element and providing it to the client computer device 150 can be completed within one minute or less from the time of the input audio signal and can be considered real-time.
[0035] The output signal corresponding to the content element, e.g., an output signal received or generated by the audio signal generator component 130, which is transmitted to the client computer device 150 via the interface 115 and the computer network 165, can cause the client computer device 150 to execute the audio driver 153 to drive the loudspeaker 154 and generate an acoustic wave corresponding to the output signal. The acoustic wave can contain words corresponding to the content.
[0036] The direct action API 135 of the data processing system can generate action data structures based on the trigger keyword. The direct action API 135 can execute a specific action to fulfill the end user's intent as determined by the data processing system 105. Depending on the action specified in its inputs, the direct action API 135 can execute code or a dialog script that identifies the parameters required to fulfill a user request. The action data structure can be generated in response to the request. The action data structure can be included in messages transmitted to or received from the service provider computer device 160. Based on the request parsed by the NLP component 110, the direct action API 135 can determine which of the service provider computer devices 160 the message should be sent to.For example, if an input audio signal includes "Order a taxi," the NLP component 110 can identify the trigger word "Order" and the taxi request. The Direct Action API 135 can package the request into an action data structure and transmit it as a message to a service provider computer device 160 of a taxi service. The message can also be forwarded to the content picker component 125. The action data structure can include information to complete the request. In this example, the information can include a pickup location and a destination. The Direct Action API 135 can retrieve a template 149 from the data container 145 to determine which fields to include in the action data structure. The Direct Action API 135 can determine necessary parameters and package the information into an action data structure.The Direct Action API 135 can retrieve content from the Data Container 145 to obtain information for the fields of the data structure. The Direct Action API 135 can populate the template fields with this information to generate the data structure. The Direct Action API 135 can also populate the fields with data from the input audio signal. The templates 149 can be standardized for categories of service providers or standardized for specific service providers. For example, ride-sharing service providers can use the following standardized template 149 to generate the data structure: {client_device_identifier; authentication_credentials; pick_up_location; destination_location; no_passengers; service_level}. The action data structure can then be sent to another component, such as the Content Picker component 125, or to the service provider computer device 160 to be populated.
[0037] The direct action API 135 can communicate with the service provider computer device 160 (which can be associated with the content element, such as a ride-sharing company) to order a taxi or ride-sharing vehicle to the cinema's location at the time the film ends. The data processing system 105 can receive this location or time information as part of the data packet (or other protocol) based on data message communication with the client computer device 150, the data storage device 145, or from other sources, such as the service provider computer device 160 or the content provider computer device 155. Confirmation of this order (or other conversion) can occur as audio communication from the data processing system 105 to the client computer device 150 in the form of an output signal from the data processing system 105, which then drives the client computer device 150 to produce audio output, such as...“Great, you have a car waiting for you outside the cinema at 11 p.m.” The data processing system 105 can communicate with the service provider computer device 160 via the direct action API 135 to confirm the order for the car.
[0038] The data processing system 105 can receive the response (e.g., "Yes, please") to the content ("Would you like a ride home from the cinema?") and route a packet-based data message to the service provider NLP component 161 (or another component of the service provider computer device). This packet-based data message can cause the service provider computer device 160 to perform a conversion, such as making a reservation to pick up a car from outside the cinema. This conversion—or confirmed order—(or any other conversion of another action of the thread) can occur before the completion of one or more actions of the thread, such as before the end of the film, as well as after the completion of one or more actions of the thread, such as after dinner.
[0039] The Direct Action API 135 can receive content data 148 (or parameters 146 or policies 147) from the data container 145, as well as data received with the end user's consent from the client computer device 150, to determine location, time, user accounts, logistical, or other information for reserving a car from the ride-sharing service. The content data 148 (or parameters 146 or policies 147) can be included in the action data structure. If the content in the action data structure includes end-user data used for authentication, the data can be hashed before being stored in the data container 145.Using the Direct Action API 135, the data processing system 105 can also communicate with the service provider computer device 160 to complete the conversion by making the reservation for the carpool pickup in this example.
[0040] The data processing system 105 can cancel actions associated with content elements. This cancellation can occur in response to the network security device 123 generating an alarm condition. The network security device 123 can generate an alarm condition if it predicts that the input audio signal is harmful or not otherwise provided by an authorized end user of the client computer device 150.
[0041] The data processing system 105 can include a network security device 123, interface with it, or otherwise communicate with it. The network security application 123 can authenticate signal transmissions between the client computer device 150 and the content provider computer device 155. The signal transmissions can be the audio inputs from the client computer device 150 and the response audio signals from the client computer device 150. The response audio signals can be generated in response to content that the data processing system 105 transmits to the client computer device 150 during one or more communication sessions. The network security device 123 can authenticate the signal transmission by comparing the action data structure with one or more properties of the input audio signals and response audio signals.
[0042] The network security device 123 can determine features of the input audio signal. These features can include a voiceprint, a keyword, the number of detected voices, an audio source identification, and the location of an audio source. For example, the network security device 123 can measure the spectral components of the input audio signal to generate a voiceprint of the voice used for the input audio signal. The voiceprint generated in response to the input audio signal can be compared to a stored voiceprint held by the data processing system 105. The stored voiceprint can be an authenticated voiceprint—for example, a voiceprint generated by an authenticated user of the client computer device 150 during a system setup phase.
[0043] The network security device 123 can also determine non-audio properties of the input audio signal. The client computer device 150 can incorporate non-audio information into the input audio signal. The non-audio information can be a location as determined or specified by the client computer device 150. The non-audio information can include a client computer device 150 identifier. Non-audio properties or information can also include physical authentication devices, such as answering the security question with a one-time password device or a fingerprint reader.
[0044] The network security device 123 can set an alarm state if the properties of the input audio signal do not match the action data structure. For example, the network security device 123 can detect mismatches between the action data structure and the properties of the input audio signal. In one example, the input audio signal might include the location of the client computer device 150. The action data structure might include a predicted location of the end user, such as a location based on the general location of the end user's smartphone. If the network security device 123 determines that the location of the client computer device 150 is not within a previously defined range of the location contained in the action data structure, the network security device 123 can set an alarm state.In another example, the network security device 123 can compare the voiceprint of the input audio signal with a voiceprint of the end user stored in the data container 145 and contained in the action data structure. If the two voiceprints do not match, the network security device 123 can set an alarm condition.
[0045] The network security device 123 can determine which input audio signal properties the authentication will use based on the response to the request contained in the input audio signal. Authentication with different properties may have different computational requirements. For example, comparing voiceprints may be more computationally intensive than comparing two locations. Selecting computationally intensive authentication methods can be excessively computationally intensive if they are unsuitable. The network security device 123 can improve the efficiency of the data processing system 105 by selecting the properties used for authentication based on the request. For example, if the security risk of the input audio signal is low, the network security device 123 can select an authentication method with a non-computationally intensive property.The network security device 123 can select the property based on the cost required to complete the request. For example, a voiceprint property might be used if the input audio signal is "Order a new laptop computer," but a location property might be selected if the input audio signal is "Order a taxi." The property selection can also be based on the time or computational intensity required to complete the request. Properties that consume more computational resources can be used to authenticate input audio signals that generate requests requiring more resources. For example, the input audio signal "OK, I'd like to go to dinner and the movies" might involve multiple actions and requests, as well as multiple service providers 160.The input audio signal can generate requests to search for movies, check restaurant availability, make restaurant reservations, and buy movie tickets. Completing this input audio signal is both more computationally intensive and slower than completing the input audio signal "OK, what time is it?".
[0046] The network security device 123 can also set an alarm state based on the request contained in the input audio signal. The network security device 123 can automatically set an alarm state if the transmission of the action data structure to a service provider computer device 160 could result in a financial charge to the end user of the client computer device 150. For example, a first input audio signal, "OK, order a pizza," might generate a monetary charge, while a second input audio signal, "OK, what time is it?", would not. In this example, the network security device 123 can automatically set an alarm state when it receives an action data structure corresponding to the first input audio signal, and not set an alarm state when it receives an action data structure corresponding to the second input audio signal.
[0047] The network security device 123 can set an alert state based on the determination that the action data structure is intended for a specific service provider device 160. For example, the end user of the client computer device 150 can set restrictions on which service providers the data processing system 105 may interact with on behalf of the end user without further authorization. For example, if the end user has a child, in order to prevent the child from purchasing toys through a service provider that sells toys, the end user can set a restriction that action data structures cannot be transmitted to the toy seller without further authentication.When the network security device 123 receives an action data structure intended for a specific service provider device 160, the network security application 123 can look up a policy in the data container to determine whether to automatically set an alarm state.
[0048] The network security device 123 can send alerts about the alarm status to the content selection component 125. The content selection component 125 can select a content element to be transmitted to the client computer device 150. The content element can be an auditory request for a passphrase or additional information for authenticating the input audio signal. The content element can be transmitted to the client computer device 150, where the audio driver 153 converts the content element into sound waves via the transducer 152. The end user of the client computer device 150 can respond to the content element. The end user's response can be digitized by the sensor 151 and transmitted to the data processing system 105. The NLP component 110 can process the response audio signal and provide the response to the network security device 123.The network security device 123 can compare a property of the response audio signal with a feature of the input audio signal or the action data structure. For example, the content element might be a request for a passphrase. The NLP component 110 can recognize the text of the response audio signal and forward it to the network security device 123. The network security device 123 can perform a hash function on the text. After the end-user's authenticated passphrase has been hashed using the same hash function, it can be stored in the data container 145. The network security device 123 can compare the hashed text with the secure, hashed passphrase. If the hashed text and the hashed passphrase match, the network security device 123 can authenticate the input audio signal.If the hashed text and the hashed pass phase do not match, the network security device 123 can set a second alarm state.
[0049] The network security device 123 can terminate communication sessions. The network security device 123 can transmit instructions to a service provider computer device 160 to disable, interrupt, or otherwise terminate a communication session established with the client computer device 150. Termination of the communication session can occur in response to the network security device 123 setting a second alarm state. The network security device 123 can disable the computer device's ability to generate communication sessions with a service provider computer device 160 via the data processing system 105.For example, if the network security device 123 sets a second alarm state in response to the input audio signal "OK, order a taxi," the network security device 123 can disable the possibility of communication sessions established between the client computer device 150 and the taxi service provider device. An authorized user can reauthorize the taxi service provider device at a later time.
[0050] Fig. Figure 2 illustrates a flowchart demonstrating an exemplary operation of System 200 for performing audio signal authentication. System 200 may include one or more of the components or elements described above in relation to System 100. For example, System 200 may include a data processing system 105 that communicates with a client computer device 150 and a service provider computer device 160, for example, via the network 165.
[0051] The operation of system 200 can begin when the client computer device 150 transmits an input audio signal 201 to the data processing system 105. Once the data processing system 105 receives the input audio signal, the NLP component 110 of the data processing system 105 can parse the input audio signal into a request and a trigger keyword that corresponds to the request. A communication session can then be established between the client computer device 150 and the service provider computer device 160 via the data processing system 105.
[0052] The Direct Action API 135 can generate an action data structure based on the request. For example, the input audio signal might be "I want to go to the movies." In this example, the Direct Action API 135 can determine if the request is for a car service. The Direct Action API 135 can determine the current location of the client computer device 150 that generated the input audio signal and the location of the nearest movie theater. The Direct Action API 135 can generate an action data structure that includes the location of the client computer device 150 as the pickup location for the car service and the location of the nearest movie theater as the destination for the car service. The action data structure can also include one or more properties of the input audio signal. The data processing system 105 can forward the action data structure to the network security device to determine whether an alarm condition should be set.
[0053] If the network security device detects an alarm condition, the data processing system 105 can select a content element via the content selection component 125. The data processing system 105 can then provide the content element 202 to the client computer device 150. The content element 202 can be provided to the client computer device 150 as part of a communication session between the data processing system 105 and the client computer device 150. The communication session can have the flow and feel of a real-time, person-to-person conversation. For example, the content element can include audio signals that are played back on the client computer device 150. The end user can respond to the audio signal, which can be digitized by the sensor 151 and transmitted to the data processing system 105.The content element can be a security question, a content element, or another question transmitted to the client computer device 150. The question can be posed to the end user who generated the input audio signal via the converter 152. In some implementations, the security question can be based on previous interactions between the client computer device 150 and the data processing system 105. For example, if the user ordered a pizza via the system 200 before transmitting the input audio signal by providing the input audio signal of "OK, order a pizza," the security questions could include "What did you order for dinner last night?" The content element can also include a request to provide a password to the data processing system 105.The content element can include a push notification to a second computer device 150 that is linked to the first computer device 150. For example, a push notification requesting confirmation of the input audio signal can be sent to a smartphone linked to the client computer device 150. The user can select the push notification to confirm that the input audio signal is authentic.
[0054] During the communication session between client computer 150 and data processing system 105, the user can respond to the content element. The user can respond verbally. The response can be digitized by sensor 151 and transmitted to data processing system 105 as a response audio signal 203, carried by a variety of data packets. The audio signal can also include characteristics that can be analyzed by the network security device. If the network security device determines that an alarm condition persists based on the conditions of the response audio signal, the network security device can send a message 204 to the service provider computer 160. Message 204 can contain instructions for the service provider computer 160 to disable the communication session with client computer 150.
[0055] Fig. Figure 3 illustrates an exemplary Procedure 300 for authenticating packetized audio signals in a speech-activated data packet (or other protocol)-based computer network environment. Procedure 300 may involve receiving data packets containing an input audio signal (ACT 302). For example, the data processing system may execute, start, or call the NLP component to receive packet- or other protocol-based transmissions over the network from the client computer device. The data packets may contain or correspond to an input audio signal detected by the sensor, such as an end user speaking into a smartphone: “OK, I want to go out for dinner tonight and then watch a movie in the evening.”
[0056] Procedure 300 can involve identifying a request and a trigger keyword within the input audio signal (ACT 304). For example, the NLP component can analyze the input audio signal to identify requests (such as "dinner" or "movie" in the example above) as well as the keywords "go" and "to go" or "in order to go" that correspond to or relate to the request.
[0057] Procedure 300 involves generating an initial action based on the request (ACT 306). The Direct Action API can generate a data structure that can be transmitted and processed by the service provider computer or content provider computer to fulfill the request of the input audio signal. For example, continuing the example above, the Direct Action API can generate an initial action data structure that is transmitted to a restaurant reservation service. The initial action data structure can search for a restaurant that is near the current location of the client computer and meets other specifications associated with the client computer user (e.g., cuisine preferences of the client computer user). The Direct Action API can also determine a preferred time for the reservation.For example, the data processing system can determine that the restaurant selected in the search is 15 minutes away and that the current time is 6:30 PM. The data processing system can set the preferred reservation time after 6:45 PM. In this example, the first action data structure can include the restaurant name and the preferred reservation time. The data processing system can transmit the first action data structure to the service provider computer or the content provider computer. ACT 306 can include generating multiple action data structures. For the input audio signal above, a second action data structure containing a movie title and restaurant name can be generated, and a third action data structure containing pickup and drop-off locations can be generated.The data processing system can provide the second action data structure to a cinema ticket reservation service and the third action data structure to a car reservation service.
[0058] Procedure 300 may also include comparing the first action data structure with a property of the input audio signal (ACT 308). The network security device may compare the property of the input audio signal with the first action data structure to determine the authenticity of the input audio signal. Determining the authenticity of the input audio signal may involve determining whether the person who generated the input audio signal is authorized to generate input audio signals. Properties of the input audio signal may include a voiceprint, a keyword, a number of detected voices, an identification of an audio source (for example, an identification of the sensor or client computer device from which the input audio signal originates), a location of an audio source, or the location of another client computer device (and the distance between the other client computer device and the audio source).For example, during a setup phase, an authorized voiceprint can be generated by having a user speak passages. When these passages are spoken, the network security device can generate a voiceprint based on the frequency content, quality, duration, intensity, dynamics, and pitch of the signal. The network security device can generate an alert if it determines that the characteristics of the input audio signal do not match the initial action data structure or other expected data. For example, if an action data structure is generated for "OK, I want to go out for dinner tonight and then watch a movie in the evening," the data processing system can generate an action data structure for a car reservation service that includes a pickup location based on the user's smartphone location.The action data structure can include the location. The input audio signal can be generated via an interactive speaker system. The location of the interactive speaker system is transmitted to the data processing system along with the input audio signal. In this example, if the user's smartphone location does not match the location of the interactive speaker system (or is not within a predefined distance of the interactive speaker system), then the user is not near the interactive speaker system, and the network security device can determine that the user most likely did not generate the input audio signal. The network security device can then generate an alert condition. The distance between the client computer device 150 and a secondary client device (e.g., a smartphone) is also considered.The distance (from the end user's smartphone) can be calculated as the straight-line distance between the two devices, or as the driving distance between them. The distance can also be calculated based on the travel time between the locations of the two devices. Alternatively, the distance can be based on other location-specific properties, such as IP address and Wi-Fi network location.
[0059] Procedure 300 can include selecting a content element (ACT 310). The content element can be generated based on the trigger keyword and the alarm state and selected through a real-time content selection process. The content element can be selected to authenticate the input audio signal. The content element can be a notification, an online document, or a message displayed on a client computer device, such as a user's smartphone. The content element can be an audio signal transmitted to the client computer device and sent to the user via the transducer. The content element can be a security question. The security question can be a predefined security question, such as a password prompt. The security question can be dynamically generated.For example, security may be a question generated based on the user's history or the client computer device.
[0060] Procedure 300 can involve receiving data packets containing auditory signals (ACT 312). The data packets can transmit auditory signals that are exchanged between the client computer and the data processing system's conversational API. The conversational API can establish a communication session with the data processing system in response to the interaction with the content element. The auditory signals can include the user's response to the content element transmitted to the client computer during ACT 310. For example, the content element can cause the client computer to generate an audio signal asking, "What is your authorization code?" The auditory signals can also include the end user's response to the content element. The end user's response to the content element can be a feature of the response audio signal.
[0061] Procedure 300 can also include comparing a property of the response audio signal with a property of the input audio signal (ACT 314). The response audio signal can include a passphrase or other properties. The content element can include instructions for the client computer device to capture one or more specific properties of the response audio signal. For example, the property of the input audio signal can be the location of the client computer device. The property of the response audio signal can be different from the property of the input audio signal. For example, the property of the response audio signal can be a voiceprint. The content element can include instructions for capturing the voiceprint property. The instructions can include capturing the response audio signal at a higher sampling frequency so that additional frequency content for the voiceprint can be analyzed.If the system detects a discrepancy between the properties of the response audio signal and the input audio signal, it may trigger an alarm. For example, if the properties of the response audio signal include a passphrase that does not match a passphrase associated with the input audio signal, an alarm may be triggered.
[0062] If the properties of the response audio signal match those of the input audio signal (e.g., the passphrases (or hashes thereof) match), a pass state can be set. When a pass state is set, the system can send instructions to a third party to resume the communication session with the client device. These instructions can authenticate the communication session for a predetermined period, thus preventing the need for re-authentication until that period expires.
[0063] Procedure 300 may also include transmitting an instruction to a third-party device to disable the communication session (ACT 316). Disabling the communication session can prevent messages and action data structures from being transmitted to the service provider's device. This can improve network utilization by reducing unwanted network traffic. Disabling the communication session can also reduce computational overhead, as the service provider's devices will not process requests that are malicious or have been generated incorrectly.
[0064] Fig. Figure 4 shows a block diagram of an exemplary computer system 400. The computer system or computer device 400 can include the system 100 or its components, such as the data processing system 105, or it can be used to implement them. The computer system 400 includes a bus 405 or other communication component for transmitting information, as well as a processor 410 or processing circuitry coupled to the bus 405 for processing information. The computer system 400 can also include one or more processors 410 or processing circuitry coupled to the bus for processing information. The computer system 400 further includes main memory 415, such as random-access memory (RAM) or other dynamic storage device, coupled to the bus 405 for storing data, as well as instructions to be executed by the processor 410.The main memory 415 can be or include the data container 145. The main memory 415 can also be used to store position data, temporary variables, or other medium-term information when instructions are executed by the processor 410. The computer system 400 can also include a read-only memory (ROM) 420 or other static storage device coupled to the bus 405 to store static information and instructions for the processor 410. A storage device 425, such as a solid-state device, magnetic disk, or optical disk, can be coupled to the bus 405 to store information and instructions permanently. The storage device 425 can include or be part of the data container 145.
[0065] The computer system 400 can be connected via bus 405 to a display 435, such as a liquid crystal display (LCD) or active matrix display, to show information to a user. An input device 430, such as a keyboard with alphanumeric and other keys, can also be connected via bus 405 to transmit selected information and commands to the processor 410. The input device 430 can include a touchscreen display 435. The input device 430 can also include cursor control, such as a mouse, trackball, or arrow keys on the keyboard, allowing directional data and selected commands to be transmitted to the processor 410 and controlling the movement of the cursor on the display 435. The display 435 can, for example, be part of the data processing system 105, the client computer device 150, or other components of Fig. Be 1.
[0066] The processes, systems, and procedures described herein can be implemented by the computer system 400 in response to the processor 410 executing a set of instructions contained in main memory 415. These instructions can be read into main memory 415 from another computer-readable medium, such as storage device 425. The execution of the set of instructions contained in main memory 415 causes the computer system 400 to execute the processes described and illustrated herein. In a multi-processor arrangement, one or more processors can be used to execute the instructions contained in main memory 415. Hardwired circuitry can be used instead of, or in combination with, software instructions in conjunction with the systems and procedures described herein. The systems and procedures described herein are not limited to any specific combination of hardware circuitry and software.
[0067] Although an exemplary computer system in Fig. As described in section 4, the subject matter, including the processes described in this specification, can be implemented in other types of digital electronic circuits or in computer software, firmware or hardware, including in the structures disclosed in this specification and their structural equivalents or in combinations of one or more of the same.
[0068] In situations where the systems described herein collect or potentially use personal information about users, users may be given the option to configure whether programs or functions collect user information (e.g., information about a user's social network, social actions or activities, user preferences, or user location), or to configure whether and to what extent the same system can receive content from a content server or other data processing system that may be more relevant to the user. Additionally, certain data may be anonymized in one or more ways before being stored or used, so that personal data is removed when parameters are generated.For example, a user's identity can be anonymized so that no personally identifiable information can be determined for the user, or a user's geographic location can be generalized by extracting location information (such as city, postal code, or state) so that a specific location of a user cannot be determined. This allows the user to control how information about them is collected and used by a content server.
[0069] The subject matter and the processes described in this specification can be implemented in digital electronic circuit arrangements or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations thereof. The subject matter described in this description can be implemented as one or more computer programs, for example, as one or more circuits of computer program instructions encoded on one or more computer storage media for execution by or control of data processing devices.Alternatively or additionally, the program instructions may be encoded in an artificially generated propagating signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiving device for execution by a data processing device. A computer storage medium may be or include a computer-readable storage device, a computer-readable storage substrate, a freely addressable or serial access memory array or device, or a combination thereof. Although a computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in an artificially generated propagating signal. Furthermore, the computer storage medium may consist of one or more separate components or media (e.g.,(multiple CDs, data carriers, or other storage devices, or which may be contained therein). The operations described in this specification can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.
[0070] The terms "data processing system," "computer device," "component," or "data processing device" encompass various devices, apparatus, and machines for processing data, including, for example, a programmable processor, a computer, one or more systems on a chip, several of the same, or combinations thereof. The device may include specialized logic circuitry, such as an FPGA (field-programmable general-purpose circuit) or an ASIC (application-specific integrated circuit). In addition to hardware, the device may also include code that creates an execution environment for the corresponding computer program, such as code representing processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination thereof.The device and execution environment can implement various computer model infrastructures, such as web services, as well as distributed computing and geographically distributed computing infrastructures. For example, the Direct Action API 135, the Content Selection Component 125, the Network Security Device 123, or the NLP Component 110, and other data processing system 105 components can include or share one or more data processing devices, systems, computer equipment, or processors.
[0071] A computer program (also called a program, software, software application, app, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be provided in any form, such as a standalone executable program or module, component, subroutine, object, or any other unit suitable for use in a computer environment. A computer program can correspond to a file in a file system. A computer program can be stored in a portion of a file containing other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in several coordinated files (such as files containing one or more modules, subprograms, or code snippets).A computer program can be deployed and run on one computer or on multiple computers located at one or more sites and connected to each other via a communication network.
[0072] The processes and logical sequences described in this specification can be performed by one or more programmable processors executing one or more computer programs (e.g., components of Data Processing System 105) to perform operations by processing input data and generating outputs. The processes and logical sequences can also be executed by a specialized logic circuit, such as a field-programmable general-purpose circuit (FPGA) or an application-specific integrated circuit (ASIC), and devices in the form of such circuits can be implemented. Media suitable for storing computer program instructions and data include all types of solid-state storage, media, and storage devices, including semiconductor memory elements such as EPROM, EEPROM, and flash memory devices; magnetic hard disks, such as...Internal hard disks or removable disks; magneto-optical hard disks; and CD-ROM and DVD-ROM drives. The processor and memory can be supplemented by or integrated into a special logic circuit.
[0073] The object described herein can be implemented in a computer system that includes a back-end component, such as a data server, or a middleware component, such as an application server, or a front-end component, such as a client computer with a graphical user interface, or a combination of one or more of said back-end, middleware, or front-end components, or a web browser, through which a user can interact with an implementation of the object described in this specification. The system components can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0074] The computer system, such as System 100 or System 400, can include clients and servers. A client and a server are generally located remotely and typically interact via a communication network (e.g., Network 165). The client-server relationship arises from computer programs running on the respective computers, which establish a client-server relationship. In some implementations, a server sends data (e.g., data packets representing a content element) to a client device (e.g., for the purpose of displaying data and receiving user input from a user interacting with the client device). Data generated in the client device (e.g., a result of user interaction) can be received by the client device at the server (e.g.,received by the data processing system 105 from the computer device 150 or the content provider computer device 155 or the service provider computer device 160).
[0075] Although the processes in the drawings are depicted in a specific sequence, it is not necessary that these processes be carried out in the depicted order or in sequential order, nor is it necessary that all illustrated processes be carried out. The actions described herein can be performed in any order.
[0076] The separation of different system components does not require separation in all implementations; moreover, the described program components can be contained in a single hardware or software product. For example, the NLP component 110, the content selection component 125, or the network security device 123 can be a single component, an application or program, a logic device with one or more processing circuits, or part of one or more servers of the data processing system 105.
[0077] Having described several illustrative implementations, it is clear that the foregoing serves for illustration and not as a limitation, and has been presented merely as an example. In particular, although many of the examples presented herein involve specific combinations of process operations or system elements, these operations and elements can be combined in other ways to achieve the same goals. Operations, elements, and features explained in connection with one implementation are not intended to preclude a similar role in other implementations or embodiments.
[0078] The language and terminology used herein serve a descriptive purpose and should not be considered restrictive. The use of the words "including," "comprehensive," "exhibiting," "containing," "incorporating," "characterized by," "characterized by," and variations thereof, here means that the items listed thereafter, their equivalents, and additional items, as well as alternative implementations consisting solely of the items listed thereafter, are included. In an implementation, the systems and procedures described herein consist of one, any combination of more than one, or all of the elements, modes of operation, or components described herein.
[0079] Any references to implementations, elements, or modes of operation of the systems and methods mentioned herein in the singular may also include implementations including a plurality of such elements, and any reference to an implementation, element, or mode of operation of any kind mentioned herein in the plural may also include implementations including a single element. References to the singular or plural form are not intended to limit the systems and methods disclosed herein, their components, modes of operation, or elements to single or multiple configurations.References to a mode of operation or an element of any kind, based on information, modes of operation or elements of any kind, may include implementations whose mode of operation or element is based at least partially on information, modes of operation or elements of any kind.
[0080] Each of the implementations disclosed herein may be combined with any other implementation or embodiment, and references to "an implementation," "some implementations," "the implementation," or the like do not necessarily exclude one another but indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation or embodiment. Such terms, as used herein, do not necessarily refer to the same implementation. Each implementation may be combined with any other implementation, including or exclusively, and in any manner consistent with the aspects and implementations disclosed herein.
[0081] References to "or" can be interpreted as inclusive, meaning that any term described by "or" can refer to any single term, more than one, or all of the terms described. For example, a reference to "at least one of 'A' and 'B'" can include only 'A', only 'B', or both 'A' and 'B'. These references, when used in conjunction with "comprehensive" or other open terminology, can include additional elements.
[0082] If technical features in the drawings, detailed description, or any claim are followed by reference numerals, these reference numerals have been included to enhance the clarity of the drawings, detailed description, or claims. Accordingly, neither the presence nor the absence of such reference numerals restricts the scope of the claim elements.
[0083] The systems and methods described herein can also be implemented by other embodiments without deviating from their essential properties. The preceding implementations are considered illustrative rather than limiting for the systems and methods described herein. The scope of the systems and methods described herein is therefore defined more by the appended claims than by the preceding description, whereby modifications that fall within the meanings and scope of equivalence of the claims are thus included herein.
Claims
[1] System (100; 200) for authenticating packetized audio signals in a voice-activated computer network environment, comprising: a natural language processor component (110) executed by a data processing system (105) to receive data packets via an interface (115) of the data processing system, comprising an input audio signal (201) detected by a sensor (151) of a client device; the natural language processing component to parse the input audio signal in order to identify a request and a trigger keyword according to the request; a direct action application programming interface (135) of the data processing system to generate an initial action data structure in response to the request based on the trigger keyword; a network security device (123) for comparing the first action data structure with a first property of the input audio signal to detect an alarm condition; a content selection component executed by the data processing system to receive the trigger keyword identified by the natural language processor and the specification of the first alarm state, and based on the trigger keyword and the specification, select a content item (202); the network security device for: Receiving data packets carrying a response audio signal (203) transmitted between the client device and a conversational application programming interface (136) that has established a communication session with the client device; Comparing a second property of the response audio signal with the first property of the input audio signal to detect a second alarm state; and Transmitting an instruction based on the second alarm state to a third-party device to disable the communication session established with the client device. [2] System according to claim 1, comprising the network security device for: Determining the first property of the input audio signal; and Determining the second property of the response audio signal, wherein the first property and the second property include at least one of a voiceprint, a keyword, a number of detected voices, an identification of the client device, and a location of a source of the input audio signal. [3] System according to claim 1, wherein the first property differs from the second property. [4] System according to claim 1, comprising the network security device for: Receiving the location of a second client device; Determining the distance between the location of one client device and the location of the second client device; and Detecting the alarm state based on the distance between the location of one client device and the location of the second client device. [5] System according to claim 4, comprising the network security device for: Detecting the alarm state based on the distance between the location of one client device and the location of the second client device that is above a previously defined threshold. [6] System according to claim 1, wherein the content element comprises instructions for generating an auditory signal at the client device. [7] System according to claim 6, wherein the auditory signal comprises a security question. [8] System according to claim 1, comprising the network security device for: Disabling the first action data structure in response to the detection of the first alarm condition. [9] System according to claim 1, comprising the content selection element for: Generating instructions to capture the second property of the response audio signal in the content element. [10] System according to claim 1, comprising the network security device for: Closing the communication session established with the client device in response to the interaction with the content element. [11] System according to claim 1, comprising the network security device for: Determine the amount of computing resources required to complete the request. [12] System according to claim 11, comprising the network security device for setting the alarm state in response to the fact that the amount of computing resources exceeds a previously defined threshold. [13] System according to claim 1, comprising the natural language processor component for: Parsing the response audio signal to identify a passphrase. [14] System according to claim 13, comprising the network security device for: Setting the second alarm state based on a passphrase that does not match a saved passphrase. [15] System according to claim 13, wherein the passphrase is the second feature. [16] Method (300) for authenticating packetized audio signals in a speech-activated computer network environment, comprising: Receiving (302) data packets by a natural language processor component (110) executed by a data processing system (105), comprising an input audio signal (201) detected by a sensor (151) of a client device; Parsing the input audio signal by the natural language processing component to identify a request and a trigger keyword according to the request (304); Generating an initial action data structure through a direct action application programming interface (135) of the data processing system, based on the trigger keyword in response to the request; Comparing (308) the first action data structure with a first property of the input audio signal by a network security device (123) to detect an alarm condition; Selecting (310) a content element (202) based on the trigger keyword and the alarm state by a content selection component executed by the data processing system; Receiving (312) data packets by the network security device carrying a response audio signal (203) transmitted between the client device and a conversational application programming interface (136) that has established a communication session with the client device; (314) comparing a second property of the response audio signal with the first property of the input audio signal by the network security device to detect a second alarm condition; and Transmitting (316) an instruction to a third-party device by the network security device based on the second alarm state to disable the communication session established with the client device in response to the interaction with the content element. [17] The method of claim 16, comprising: Determining the first property of the input audio signal by the network security device; and Determining the second property of the response audio signal by the network security device, wherein the first property and the second property include at least one of a voiceprint, a keyword, a number of detected voices, an identification of the client device, and a location of a source of the input audio signal. [18] The method of claim 16, comprising: Receiving the location of a second client device by the network security device; Determining the distance between the location of one client device and the location of the second client device using the network security device; and Detecting the alarm state based on the distance between the location of one client device and the location of the second client device by the network security device. [19] System (100; 200) for authenticating packetized audio signals in a voice-activated computer network environment, comprising: a natural language processor component (110) executed by a data processing system (105) to receive data packets via an interface (115) of the data processing system, comprising an input audio signal (201) detected by a sensor (151) of a client device; the natural language processing component to parse the input audio signal in order to identify a request and a trigger keyword according to the request; a direct action application programming interface (135) of the data processing system to generate an initial action data structure in response to the request based on the trigger keyword; a network security device (132) for comparing the first action data structure with a first property of the input audio signal to detect an alarm condition; a content selection component executed by the data processing system to receive the trigger keyword identified by the natural language processor and the specification of the first alarm state, and based on the trigger keyword and the specification, select a content item (202); the network security device for: Receiving data packets carrying a response audio signal (203) transmitted between the client device and a conversational application programming interface (136) that has established a communication session with the client device; Comparing a second property of the response audio signal with the first property of the input audio signal to detect a match state; and Transmitting an instruction based on the pass state to a third-party device to continue the communication session established with the client device. [20] System according to claim 19, comprising the network security device for: Determining the first property of the input audio signal; and Determining the second property of the response audio signal, wherein the first property and the second property include at least one of a voiceprint, a keyword, a number of detected voices, an identification of the client device and a location of a source of the input audio signal, and wherein the second property includes a security question.
Citation Information
Patent Citations
System and method for the detection and termination of fraudulent services
US20050185779A1
Identification using Audio Signatures and Additional Characteristics
US20140249817A1
Voice recognition method, voice controlling method, information processing method, and electronic apparatus
US20150142438A1
Dynamic threshold for speaker verification
US20150371639A1
Speaker identification and unsupervised speaker adaptation techniques
US20160093304A1