SYSTEM AND METHOD FOR TRANSMITTING USER-SPECIFIC DATA TO A DEVICE - Patent application
The system addresses the lack of integrated user data capture and display flexibility by using a cross-modal voice interface engine to process user requests and transmit user-specific data across devices, ensuring flexible and user-controlled display options.
Patent Information
- Application Number
- JP2022535133
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-02-18
- Publication Date
- 2025-08-06
- Estimated Expiration
- 2041-02-18
AI Technical Summary
Existing systems lack an integrated approach to intuitively capture and store user-specific data for future use, provide flexible display options, and allow users to select display preferences.
A system and method for connecting a communication device to a display device, utilizing a cross-modal voice interface engine to receive user requests, process them through an IVR system and natural language understanding, and plan data transmission based on user parameters, enabling user-specific data display on any device via a telecommunications network.
Enables independent data transmission across various devices, providing users with flexible display options and allowing them to select how their data is presented, regardless of the device type or network capabilities.
Smart Images

Figure 0007719508000001 
Figure 0007719508000002 
Figure 0007719508000003
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to systems and methods for transmitting user-specific data, and more particularly, the present invention relates to a method for connecting / interfacing a communication device to a display device and transmitting user-specific data via the display. [Background technology]
[0002] The field of data transmission has undergone remarkable changes in recent years. Marketing teams around the world face the constant challenge of reaching target customers and successfully selling their products or services (hereinafter referred to as the "Products"). Data transmission serves an important purpose, as its content, instance, and duration aid in decision-making at various stages of data processing. At the same time, different users may need different types of products, and they often face the challenge of locating data according to their needs. Various modes of data transmission are constantly undergoing innovative changes to enable data to reach target populations. These forms of data transmission have their own limitations regarding the privacy of personal data and do not guarantee the availability of outputs / products related to the corresponding data in all locations.
[0003] Personal technology (Personal Tech) includes a variety of electronic devices, which include various services that cater to individual needs and preferences for the type and amount of data required by the user. The increasing number and variety of Personal Tech devices have made it possible to individually identify requests and deliver output data / products that meet those requests. This has opened the door to a new era of demand, where individual request patterns are identified and intuitive request suggestions are provided based on user data such as age, gender, location, purchase history, menstrual cycle, internet search history, and medical history. Personal technology has made it possible to capture search history, understand user requests and purchasing trends, store data for future reference, and constantly study user behavior to establish patterns useful for precise advertising.
[0004] A voice user interface (VUI) is a module that allows users to input voice commands to enter their queries and retrieve search results. For example, if a user inquires about a product / request over the phone, the VUI displays the product / request and directly explains the product / request's features to the user. If the user is in a remote location, the VUI suggests the nearest possible location for the product / request. However, it can be difficult for VUI users to understand the capabilities of the VUI and maximize the benefits of using it. However, similar benefits are unavailable if users only have access to a regular landline phone or a less powerful device. In current devices, VUIs are integrated with displays, requiring users to focus their attention on such displays. Regardless of the input method, a robust system is still needed where user data can be collected, stored for future use, and played back when requested by the user. Summary of the Invention [Problem to be solved by the invention]
[0005] Available systems and methods have limitations and there is a need for an integrated approach that intuitively captures and stores user-specific data for future use, provides flexible display options, and thereby allows the user to select / decide on display options. [Means for solving the problem]
[0006] The present invention discloses a system and method for transmitting user-specific data. More particularly, the present invention relates to a method for connecting / interfacing a communication device to a display device and transmitting user-specific data via the display. The system includes a service request means for inputting a request by a user and a cross-modal voice interface engine configured to receive the user request via a telecommunications network. The cross-modal voice interface engine includes an IVR system configured to receive the user request and a natural language understanding component configured to receive the user request from the IVR system and convert the user request into a machine-readable transcript. The system further includes a dialog and media planner. The dialog and media planner is connected to the IVR system and configured to plan and schedule the transmission of data to the user based on user request parameters calculated by the dialog and media planner.
[0007] In another embodiment, the present invention provides a method for connecting / interfacing a communication device to a display device and transmitting user-specific data via the display, the method including receiving a user request over a network, the method further including processing the user request with a cross-modal voice interface engine, planning data to the user based on user request parameters calculated by the cross-modal voice interface engine, and transmitting the data to the user.
[0008] Therefore, by using the above-described data transmission system and method, the system is made independent of the type of personal technology device using the telecommunications network, and user-specific data is displayed on the display device according to the user's request.
[0009] The present invention will now be described in detail with reference to the drawings. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 illustrates an overall system setup according to one embodiment of the present disclosure. [Figure 2] FIG. 1 illustrates an overview of a cross-modal speech interface engine, according to one embodiment of the present disclosure. [Figure 3] FIG. 1 illustrates a traditional phone+smart device hardware setup according to one embodiment of the present disclosure. [Figure 4] FIG. 1 illustrates a smartphone hardware configuration according to one embodiment of the present disclosure. [Figure 5A] FIG. 1 illustrates a method for delivering services to a user according to one embodiment of the present disclosure. [Figure 5B] FIG. 1 illustrates a method for delivering services to a user according to one embodiment of the present disclosure. [Figure 5C] FIG. 1 illustrates a method for delivering services to a user according to one embodiment of the present disclosure. [Figure 6] FIG. 1 illustrates a configuration for data transmission using a URL transmitted via SMS using a feature phone. DETAILED DESCRIPTION OF THE INVENTION
[0011]
[0006] Described herein are systems and methods for delivering user-specific data over a telecommunications network and providing users with an efficient means for displaying such data. The systems and methods are described with reference to figures, which are intended to be illustrative and not limiting, to facilitate description of exemplary systems and methods according to embodiments of the present invention.
[0012] The foregoing description of specific embodiments sufficiently reveals the general nature of the embodiments herein, so that others, by applying current knowledge, can easily modify and / or adapt such specific embodiments to various uses without departing from the general concept; therefore, such adaptations and modifications should, and are intended to, be understood within the meaning and range of equivalents of the disclosed embodiments. It should be understood that the phraseology or terminology used herein is for purposes of description and not of limitation. Thus, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modifications within the spirit and scope of the appended claims.
[0013] As used herein, the term "network" refers to any form of communication network used to carry data and connect communication devices (e.g., phones, smartphones, computers, servers) to one another. According to one embodiment of the present invention, the data includes at least one of processed data and unprocessed data. Such data includes data obtained from automated data processing, manual data processing, or unprocessed data. Data according to the present invention is in the form of text, images, audio files, or video files.
[0014] As used herein, the term "artificial intelligence" refers to a set of executable instructions stored on a server and generated using machine learning techniques.
[0015] As used herein, the term "hub device" refers to an electronic device that includes a processor and that is capable of connecting to other electronic / computing devices, displays, and sensors over a network.
[0016] As used herein, the term "VUI" refers to a voice user interface module that allows a user to enter a request / query by inputting voice commands and obtain a set of results corresponding to the request / query.
[0017] As used herein, the term "IVR" refers to an interactive voice response system that provides responses to incoming user queries.
[0018] FIG. 1 illustrates a system (100) according to one embodiment of the present disclosure. A user or set of users (101) may make a call using a device (102) via a carrier's network (105). The device (102) may be a telecommunications device, a smartphone, or a regular landline phone, and the user may input their service request through the device (102). The device (102) functions as a service request means for inputting the user request. The user's input voice is processed using a system and method for speech-to-text conversion of any incoming voice of any user. Artificial intelligence techniques are used to collect data from several users or individual users and extract key phrases indicative of user preferences. The call is further connected to a cross-modal voice interface engine (106). The cross-modal voice interface engine (106), which is composed of different components and configured to process the call, is described in detail in FIG. 2. The hub device (103) is communicatively connected to the cross-modal voice interface engine (106). According to one embodiment of the present invention, the hub device (103) may be a computing device or microcontroller capable of connecting to one or more networks and multiple devices (104) or sensors (107), and the hub device (103) displays a list of the devices (104) or sensors (107) to which it is connected. In one embodiment of the present invention, the hub device (103) may be connected to the devices (104) and the cross-modal voice interface engine (106) using the Internet, Bluetooth, or any other connection / interface medium for transmitting data from the hub device (103) to the devices (104) and / or sensors (107). A data transmission system for transmitting user-specific data includes a smart device connected to the hub device (103) and having sensors (107) capable of providing ambient information to the cross-modal voice interface engine (106).The cross-modal voice interface engine (106) is configured to use the hub device (103) to detect the presence of at least one of the devices (104) and sensors (107) in the user's (101) vicinity and transmit a signal to the user indicating the presence of the device (104). By nearby devices, it is understood that the devices are communicating over the same network or different networks. Thus, nearby devices include devices communicating over the same network or different networks. It is further noted that transmission of data to a user based on a user request is the same as communication of data to a user, and ultimately, the data is transmitted only to an electronic device, and the user receives such data through access to only such electronic device.
[0019] As the call progresses, the engine (106) provides the user with the option to select one of the devices (104) for displaying output or for communicating data to the devices (104). The hub device (103) receives commands and data for displaying data from the engine (106) via one or more available devices (104, 107) selected by the user.
[0020] Figure 2 details the components of the cross-modal voice interface engine (106). An interactive voice response (IVR) (201) system receives a call from a user (101). According to the present invention, the IVR (201) is configured to handle multiple calls from different users simultaneously. Furthermore, the IVR (201) can be connected to different subsystems to process incoming calls. A natural language understanding component (202) is connected to the IVR (201) to receive user input and convert the natural language input into machine-readable information. The natural language understanding component (202) generates metadata for transcribed content, intent, entities, and callers.
[0021] The machine-readable information generated by the natural language understanding component (202) defines the user's request, and the system logic generates a plan for answering the user's query and providing possible display options. A dialog and media planner (203) is provided, serving as central logic for determining further steps based on available input and output devices. A user-specific knowledge database (208) is connected to the media planner (203) to store user-specific data for future reference. The data in the database (208) includes the user's current and past interactions with the system (100). A global knowledge database (207) is connected to the user-specific knowledge database (208) to provide additional relevant information that may be useful to the user (101). A service repository (210) is connected to the planner (203) to search for and provide all possible services to the user (101) based on the user's (101) intent. In yet another embodiment of the present invention, a data transmission system for transmitting user-specific data includes a third-party service repository connected to the service repository (210) to provide related third-party services to the user. According to one embodiment of the present invention, the third-party service module (212) is integrated with the service repository (210) to search for and provide additional relevant services based on the user's (101) intent. Relevant data indicating available service options is also obtained from the registered smart device's sensors (107) and provided to the planner (203). The planner (203) can communicate with the user-specific knowledge database (208) to search various available parameters, such as the current user intent, the global knowledge database (208), the service repository (210), the third-party service module (212), and the registered smart device's sensors (107), to prepare a presentation plan for the user (101). According to one embodiment of the present invention, the presentation plan includes audio and / or visual outputs and the sequence in which the plan is presented to the user.The natural language generation component (211) creates speech output for presentation to the user (101). The visual output is sent to an available display device using the hub device (103). The system not only allows the user (101) to see and hear all available options specific to the request, but also allows the user (101) to see other relevant information that will improve the quality of the answer they seek and make a more comprehensive decision.
[0022] FIG. 3 illustrates the integration of a traditional telephone and a smart device hardware configuration according to one embodiment of the present disclosure. In a situation where a user (101) connects to an IVR system (201) via a regular landline phone (102) without network capabilities, both audio and visual output may still be provided to the user according to the current embodiment. An Internet of Things (IoT) device (103) containing hub software connects to a cross-modal voice interface engine (106) via an Internet connection, and the device (103) connects to a display, such as a local smart TV (104), using any known connection means, such as Bluetooth. Once the display device (104) is registered with the dialog and media planner (203), the dialog and media planner generates visual display commands to display the results. Thus, the embodiment of FIG. 3 generates audio and visual results in a manner similar to the embodiment of FIG. 2, even when the input device is a regular landline phone without network capabilities.
[0023] Example 1: When a user (101) wants to travel to the southern coast of Bangladesh and wants to book a holiday home, the user (101) invokes the cross-modal voice interface engine (106). The system (100) then explains the options and displays several options on the nearby screen (104). The user (100) can refer to the second home by saying "Tell me more about the second one" or by touching it on the screen (104), and the system (100) can understand this and respond by showing details about the option the user selected.
[0024] Figure 4 illustrates a smartphone hardware configuration according to one embodiment of the present disclosure, in which the smartphone (102) hosts a hub application (103) and includes multiple sensors (107) and a small display screen. An external display (104) may be connected to the smartphone (102) to provide additional display options. The smartphone (102) is registered with the planner (203), and when the user (101) calls the cross-modal voice interface engine (106), the IVR (201) processes the call in the manner described with reference to Figure 2 and provides dynamic output for further user selection. The user (101) may further select multiple options from the results provided by the cross-modal voice interface engine (106) and request the cross-modal voice interface engine (106) to provide the selected option.
[0025] Example 2: A user (101) calls the cross-modal voice interface engine (106) and says he or she wants to find a new shopping mall. The system (100) responds by explaining what stores the user will find there and any special offers currently available. At the same time, the system (100) displays images and videos of the location and the stores on the large display (104). When the user (101) announces that he or she wants to go there, the system (100) responds by explaining how to get there, displaying an overview map on the large display (104), and offering to guide the user to the location via the user's smartphone (102). If the user verbally agrees, the display (104) is turned off, and the system (100) dynamically guides the user to the location by displaying navigation instructions on the smartphone (102) while retrieving GPS data from there.
[0026] The system (100) of the above embodiment may be used to facilitate data transmission to a user (101). When a user invokes the cross-modal voice interface engine (106), the cross-modal voice interface engine (106) may determine whether appropriate data related to the user's request is available for the given user. If appropriate data related to the user (101) is already available in the system (100), the system (100) may begin transmitting the data during the call and interact with the user to explain any additional requests for data. If the relevant data is not available in the system (100), the system may locate the appropriate data from the network and initiate a new call with the user (101) to transmit the data. The user (101) has the option to continue or disconnect from the provided data. The user's (101) selection is analyzed to identify the relevance of the data to the user (101), which is stored for future reference.
[0027] 5A-5C illustrate a method (500) for transmitting data to a user based on a user request, according to one embodiment of the present invention. As shown in FIG. 5A, the method (500) includes receiving a user request over a network (510). The method further includes processing the user request (520) with the cross-modal voice interface engine 106. The method (500) further includes planning data to the user based on user request parameters calculated by the cross-modal voice interface engine and transmitting the data to the user (530).
[0028] The step of receiving a user request (510) includes receiving the user request via a smartphone, a regular telephone, or an Internet of Things (IoT) device. The method (500) further includes storing user request data in a user-specific knowledge database, the user-specific data being obtained using a transcribed speech-based interaction in a machine-readable format, an analysis of the acoustic characteristics of the speech, interaction metadata, and user persona data.
[0029] As shown in FIG. 5B, processing a user request by a cross-modal voice interface engine (520) includes receiving a user request via an IVR system (521), converting the request into a machine-readable transcript by a natural language understanding component (522), and generating a data transmission plan by a dialog and media planner (523).
[0030] As shown in FIG. 5C , generating a data transmission plan (523) includes obtaining inputs from a user-specific knowledge database, a global knowledge database, a service repository, a third-party service, and sensors (524), generating audio output using a natural language generation component (525), and generating visual output based on available visuals that are displayed on available display devices (526).
[0031] Transmitting the data to the user (530) includes delivering the relevant data to the user in audio or visual form, and displaying the visual information includes using a registered display device to visualize the data.
[0032] According to yet another embodiment of the present invention, FIG. 6 depicts a user (101) using a feature phone (102) to engage in a call over a telecommunications network (105) to interact with a system (106). The feature phone (102) hosts a set of standardized software for common tasks, including a web browser (103) for rendering web pages. When interacting with one or more services via the phone, the cross-modal speech interface engine (106) dynamically generates a URL pointing to additional media, e.g., a picture, that is published on the World Wide Web by the system on the fly. According to one embodiment of the present invention, this URL is transmitted via the SMS protocol so that the user can access this URL via SMS via an SMS application on the feature phone (102) during or after the call. Clicking on the URL automatically opens the web browser (103) and directs it to the dynamically generated URL, displaying the image that the cross-modal speech interface engine (106) intended to show to the caller.
[0033] The embodiments and examples discussed herein are intended to be illustrative of the present invention only. It is clear that many other forms of the present invention can be envisioned by those skilled in the art without departing from the scope of the present invention. The claims and embodiments are intended to encompass all such forms and modifications that fall within the scope of the claims, and the description of the embodiments should not be construed as limiting the scope of the description.
Claims
1. A data transmission system for transmitting user (101) specific data, comprising: service request means (102) for inputting a user request as a request by a user via a telecommunications network (105); a cross-modal speech interface engine (106) configured to receive the user request via a telecommunications network (105), an IVR system (201) configured to receive the user request; a natural language understanding component (202) configured to receive the user request from the IVR system (201) and convert the user request into a machine-readable transcript; a dialog and media planner (203) connected to the IVR system (201) and configured to plan and schedule transmission of the data to the user based on user request parameters calculated by the dialog and media planner, and to generate output corresponding to the data based on available output devices; a cross-modal speech interface engine (106) comprising: A data transmission system comprising:
2. 2. The data transmission system for transmitting user-specific data according to claim 1, wherein the data transmission system for transmitting user-specific data comprises a user-specific knowledge database connected to the IVR system and configured to store and update the user requests.
3. 3. The data transmission system for transmitting user-specific data according to claim 2, further comprising a global knowledge database (207) connected to the user-specific knowledge database (208) for providing information based on an associated global trend analysis of similar requests.
4. 2. The data transmission system for transmitting user-specific data according to claim 1, comprising a service repository (210) connected to the dialog and media planner (203) and configured to provide services related to the user's intentions.
5. 5. The data transmission system for transmitting user-specific data according to claim 4, wherein the data transmission system for transmitting user-specific data comprises a third-party service repository connected to the service repository for providing related third-party services to the user.
6. 2. The data transmission system for transmitting user-specific data according to claim 1, comprising a hub device (103) connected to the cross-modal speech interface engine (106) and capable of transmitting the data via a device (104).
7. 7. The data transmission system for transmitting user-specific data according to claim 6, wherein the hub device is an Internet of Things (IoT) device equipped with hub software.
8. 10. The data transmission system for transmitting user-specific data according to claim 1, comprising a smart device connected to a hub device capable of transmitting said data and having sensors capable of providing ambient information to said cross-modal voice interface engine.
9. 6. The data transmission system for transmitting user-specific data according to claim 1, 3 or 5, wherein the user request parameters calculated by the dialog and media planner (203) are based on inputs from at least one of the following: the user's intention, a user-specific knowledge database (208), a global knowledge database (207), a service repository (210), a third-party service module, and a sensor (107).
10. 2. The data transmission system for transmitting user-specific data according to claim 1, wherein the dialog and media planner is configured to prepare a presentation plan for a service to be provided to the user.
11. 2. The data transmission system for transmitting user-specific data according to claim 1, wherein the data transmission system comprises a natural language generation component connected to the dialog and media planner and capable of generating and transmitting data to the user.
12. 2. The data transmission system for transmitting user-specific data according to claim 1, wherein the service request means comprises a smartphone, a regular telephone, or any other input device that may be connected via the telecommunications network.
13. 12. The data transmission system for transmitting user (101) specific data according to claim 11, wherein the data transmitted to the user comprises data in audio format.
14. 12. The data transmission system for transmitting user-specific data of claim 11, wherein the data transmitted to the user includes data corresponding to at least one of the user request and system-generated data generated via the natural language generation component.
15. 12. A data transmission system for transmitting specific data of a user (101) as described in claim 11, wherein the data transmitted to the user is further transmitted from the IVR system (201) to a device (104) connected to a hub device (103) capable of transmitting the data.
16. A method for transmitting user (101) specific data, comprising: receiving a user request as a request by a user (101) via a telecommunications network (105); processing the user request by a cross-modal speech interface engine (106); planning data for the user based on user request parameters calculated by the cross-modal speech interface engine and transmitting the data to the user via output generated based on available output devices; A method for transmitting user (101) specific data, comprising:
17. The step of processing the user request by a cross-modal speech interface engine (106) includes: receiving said user request via an IVR system (201); converting the request into a machine-readable transcript by a natural language understanding component (202); generating a data transmission plan by a dialogue and media planner (203); 17. The method for transmitting user (101) specific data according to claim 16, comprising:
18. The step of generating a data transmission plan includes: obtaining inputs from a user-specific knowledge database, a global knowledge database, a service repository, third-party services and sensors; generating an audio output using a natural language generation component; generating a visual output based on the available visuals to be displayed on the available display devices; 20. The method for transmitting user (101) specific data according to claim 17, comprising:
19. 17. The method for transmitting user-specific data of claim 16, wherein receiving the user request includes receiving the user request via a smartphone, a regular telephone, or an Internet of Things (IoT) device.
20. 17. A method for transmitting user (101) specific data according to claim 16, comprising storing user request data in a user specific knowledge database.
21. 21. The method for transmitting user (101) specific data of claim 20, wherein the user specific data is obtained using a transcribed speech-based dialogue in a machine-readable format, an analysis of the acoustic characteristics of the speech, metadata of the dialogue, and persona data of the user.
22. 17. The method for transmitting user-specific data of claim 16, wherein transmitting user-specific data to the user comprises transmitting the data to the user in an audio or visual format.
23. 17. The method for transmitting user (101) specific data of claim 16, comprising displaying visual information via a registered display device (104, 107) using a hub device (103).
Citation Information
Patent Citations
Information transfer method, information transfer system and information transfer device, information reception terminal equipment used therefor
JP1996009053A
Output control device and program
JP2019087014A
SYSTEM AND METHOD FOR SECURE INTERNET OF THINGS (IoT) DEVICE PROVISIONING
JP2019502206A
Service Interfacing for Telephony
US20110064208A1