Method and system for triggering intelligent dialogue through real-world audio and video
By leveraging real-world audio and video inputs, the system addresses the limitations of current chatbots by providing contextually relevant and personalized dialogue through environmental integration and machine learning.
Patent Information
- Application Number
- JP2024229232
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current natural language chatbots lack the ability to provide relevant and real-time responses due to their lack of integration with the user's environment, limiting their effectiveness in providing situational answers.
A method and system that utilizes real-world audio and video to trigger intelligent conversations by activating a user's camera and microphone to capture environmental data, integrating location information, and employing a chatbot that uses natural language processing and machine learning to generate contextually relevant dialogue.
Enables chatbots to provide personalized and timely responses by incorporating environmental objects and sounds, enhancing the relevance and effectiveness of conversations.
Smart Images

Figure 2025110882000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method of initiating an intelligent conversation in a browsing interface, and more particularly to a method and system for triggering an intelligent conversation via real-world audio and video that provides a service for initiating an intelligent conversation when a user browses the surrounding environment using real-world video.
Background Art
[0002] ChatGPT can interact with users in the form of natural language by learning and training a large amount of information on the Internet. However, the general content for answering users is a standard answer obtained through learning, and it is not possible to provide users with an answer that responds to the current situation in real time. Although it is called a natural language chatbot, it lacks content corresponding to the relevance with users and real-time situations.
Summary of the Invention
Problems to be Solved by the Invention
[0003] In addition to the above drawbacks, the services provided by current natural language chatbots are limited to general conversations and cannot meet all needs. For example, since they are not linked to the real environment in applications, they cannot provide effective answers according to the user's situation at that time.
Means for Solving the Problems
[0004] To provide a new type of intelligent conversation suitable for the user's situation at that time, the present disclosure provides a method and system for triggering an intelligent conversation via real-world audio and video. The introduced chatbot can interact with the user based on the user's preferences and current location information, not only form relevant recommended information, but also acquire the video and audio of the place where the user is, and generate conversation content that is more suitable for the user's situation at that time in the intelligent conversation.
[0005] According to an embodiment, a system for triggering an intelligent conversation through real audio and video includes a cloud server and a database, and the cloud server executes a method for triggering an intelligent conversation through real audio and video. In this method, first, a real video interface is opened on the user device, the camera is activated to obtain environmental video, and further the microphone is activated to obtain environmental audio information.
[0006] The cloud server receives location information and a real video request from the user device. Based on the user's location information and the requested real video request, the cloud server provides corresponding location-based data, and it is also possible to display a link image displayed at a spatial position on the real video interface started on the user device.
[0007] On the user device, environmental objects included therein may be obtained by recognizing the captured environmental video. Further, environmental sound may be obtained by recognizing environmental audio information. The cloud server receives an intelligent conversation request generated by triggering an intelligent conversation link point displayed on the real video interface from the user device, and activates an intelligent conversation process between the cloud server and the user device. As a result, after the intelligent conversation interface is started and a chatbot is introduced, the chatbot executes a natural language model based on location information, environmental objects, and / or environmental sound to generate conversation content.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0009] The present disclosure provides a method and system for triggering an intelligent dialogue via real audio and video. Here, by using a user device to activate the sound collection functions of a camera and a microphone, and triggering the activation of a real video interface, the current environmental video is captured and displayed, and the current environmental sound and human voices can be received using the microphone. Based on the location information transmitted to the server, corresponding location-based data is obtained from the server, and link images of content can be displayed at spatial positions related to the location-based data in the real video interface. One specific embodiment is to display the real video within the visible range captured by the camera (for example, the surrounding environmental video) in the real video interface, and combine one or more link images displayed at one or more spatial positions to form an augmented reality image. Thereby, while the user operates the user device to view the environmental video from the display, the user can also view the link images and character information displayed at the spatial coordinates by augmented reality (AR) technology. In particular, a link point that can trigger the intelligent dialogue process by the user tapping is additionally provided, achieving the purpose of the method for triggering an intelligent dialogue via real audio and video.
[0010] According to the embodiment, the system that triggers intelligent conversation through the real voice and video can provide social media services through a network by a cloud server, and users can share text content, image content, and voice and video content after subscribing. Thereby, the location-based data obtainable from the real video can include text content, image content, voice and video content, etc. shared by the user. Further, the cloud server provides chatbots in each field, and users can have conversations using the services provided by the cloud server. Here, the utilization of artificial intelligence technology includes chatbots that use machine learning algorithms and natural language processing (NLP) technologies to learn and train data in each field to provide conversation services. By learning the activities of users on social media, the preferences of users are obtained, and the chatbots are based on the semantics of the users' conversations and the preferences of the users, and combine the environmental information obtained in real time to provide conversation content more suitable for the individual needs of the users and the current environmental characteristics.
[0011] Figure 1 is a diagram of an embodiment showing the structure of a system that triggers intelligent conversation through real voice and video. The structure of the system shown in the figure is mainly composed of a system that triggers intelligent conversation through real voice and video implemented by a server, a database, and related software and hardware on the server side, and an application program executed by a device on the user side.
[0012] The cloud server 100 shown in the figure is constructed from a computer system, a database, and a network, and as various functional modules realized by the cooperation of software and hardware, as shown in the figure, it has a natural language processing module 101, a machine learning module 103, an external system interface module 105, and a user interface module 107. The natural language processing module 101 that processes natural language information is realized by a chatbot having natural language processing capabilities. The machine learning module 103 not only executes machine learning algorithms and trains natural language models, but also learns the user's activities on the network using deep learning methods to obtain the user's preference information, so that the chatbot can provide dialogue content according to the user's preferences. The cloud server 100 includes an external system interface module 105. The external system interface module 105 executes a circuit and related application software for connecting to an external system (for example, the first external system 111, the second external system 112) (for example, via the network 10) and obtaining data via an application programming interface (API). The cloud server 100 includes a user interface module 107. By providing a connection between the user device 150 and the cloud server 100 through the network connection function in the user interface module 107, executing a server (web server) for providing web services, and executing an application program for obtaining services correspondingly on the user device 150, the services provided by the cloud server 100 can be obtained.
[0013] In order to realize the function of triggering an intelligent dialogue through real-time voice and video, the cloud server 100 is provided with an audio-visual processing module 108 capable of processing the audio-visual data provided by the user device 150. The audio-visual processing module 108 may be realized by the cooperation of a processor and a related software program. Mainly, by executing the video processing process and the audio processing process, the audio-visual data provided by the user device 150 is analyzed, and the features of the video and audio are extracted, so that the objects and voices in the video can be recognized. These features of the audio-visual may be provided to the natural language processing module 101 and utilized as materials when the chatbot provides the dialogue service.
[0014] According to the embodiment, an augmented reality module 109 is implemented in the cloud server 100 using augmented reality (AR) technology. When the user activates the augmented reality process in response to a real-time video request generated from the user device 150, the augmented reality module 109 can display the video of the surrounding environment through the display and simultaneously view the objects in the augmented reality integrated with the augmented reality video provided by the cloud server 100. For example, when the real-time video interface 115 starts on the user device 150, the operation of the augmented reality module 109 of the cloud server 100 obtains the link images of the location-based data displayed at one or more spatial positions within the visible range and obtains the corresponding information. For example, when the user activates the camera of the user device 150 to obtain and display the surrounding environmental video, the current location information of the user device 150 is also transmitted to the cloud server 100. As a result, the software program in the cloud server 100 can identify the user's current location and the surrounding environmental objects (such as buildings, roads, landscapes, etc.) that can be seen in that space, and can obtain the location-based data related to these environmental objects by querying the database. Subsequently, these data are transmitted to the user device 150, integrated into the augmented reality video through the real-time video interface 115, and the link images corresponding to the location-based data are displayed at the corresponding spatial coordinates.
[0015] According to the structure shown in the figure, the cloud server 100 is equipped with a built-in or externally connected database. Thereby, the cloud server 100 provides data services. The audio-visual database 110 shown in the figure provides audio-visual content that can be uploaded and shared by each end-user stored and accessible by the user device 150 via the network 10. What the user uploads and shares may include text content and image content. The user database 120 stores user data including the user's personal information, uploaded text content, image content, and audio-visual content, and acquires the user's activity data (e.g., network activities such as content viewed, followed, liked button, shared, subscribed, etc.) in the network services provided by the cloud server 100. Based on these, a user profile is created. Further, when interactive content is continuously generated over time, the user database 120 may store and update user data including the user's past interaction history on the time axis and provide it to the natural language model as an interaction history for learning by the machine learning algorithm. The vector database 130 stores structured data obtained by computing various text content, image content, and audio-visual content through vectorization, and may be used for matching various information that conforms to the user's personalization.
[0016] The database may further include a map database 140. The map database 140 is for a user to search for location-based data associated with a specific geographical location or spatial coordinates, and is associated with, for example, a country, an administrative region, a tourist destination, a landmark, and a place tagged by the user at a specific geographical location. In particular, not only data on geographical locations on a plane but also location-based data including spatial coordinates may be stored. For example, a restaurant installed on a certain floor of a building is given geographical location (e.g., longitude, latitude) and spatial coordinates including altitude information (e.g., spherical coordinate system described by radial distance (γ), polar angle (θ), and azimuth angle (φ), or rectangular coordinate system described by the X-axis, Y-axis, and Z-axis). Furthermore, since the map database 140 can be used to search for environmental objects with altitude information, the system can identify environmental objects that a user can view through the user device 150 based on the location information (e.g., elevation) of the user device 150 and the shooting direction (combination of azimuth angle and polar angle) of the user device 150, and can form the user's visible range.
[0017] According to the schematic diagram of the structure of the system shown in the figure, the cloud server 100 may obtain data of an external system via the network 10 or a connection of a specific protocol. The external systems are the first external system 111 and the second external system 112 schematically shown in the figure, and are, for example, servers installed by a government or a company and are used to provide open data. Thereby, the cloud server 100 can obtain real-time information that meets the needs, such as real-time weather information, real-time road conditions, real-time news, and network information related to the real-time location, via the application programming interfaces provided by the external systems respectively using the external system interface module 105.
[0018] The user device 150 runs an application capable of obtaining the services provided by the cloud server 100. For example, for the social media service provided by the cloud server 100, the user device 150 runs the corresponding social media application to obtain the social media service via the user interface module 107. In particular, the cloud server 100 provides a natural language chatbot via the natural language processing module 101, and the user can interact with the chatbot using the intelligent dialogue interface started at the augmented reality interface 115. On the other hand, the cloud server 100 uses the machine learning module 103 to learn the activity data of the user using various services of the cloud server 100. The activity data includes the activity data obtained by the user using the social media application and the augmented reality interface 115. In this way, the machine learning module 103 can learn the user's interests and characteristics and build user data.
[0019] According to the embodiment, the system that triggers intelligent interaction through the actual voice and video provides the corresponding application program to the user device 150, and the application program cooperates with the hardware in the user device 150 to execute an augmented reality process 151, a location identification process 152, an audio processing process 153, a video processing process 154, an intelligent model 155, etc. Here, the augmented reality process 151 activates the actual video interface to realize the application of augmented reality. The location identification process 152 generates location information, the audio processing process 153 processes the audio received by the microphone, the video processing process 154 processes the image captured by the camera, and the intelligent model 155 recognizes the environmental objects where the user device 150 is located based on the characteristics of the video and audio. In this way, when the cloud server 100 obtains the information provided by the user device 150, it provides an intelligent interaction service that can reflect the situation at that time based on the location information of the user device 150 and the currently obtained video and audio through the software program to be executed.
[0020] Here, it should be mentioned that the cloud server 100 obtains various text contents, image contents, and voice / video contents uploaded and shared by a large number of users as unstructured data, converts them into vectorized data using an encoding method, makes it easier to obtain the meaning of the content, and is convenient for data search. Furthermore, the vectorized data may be used when collating with the user's search keywords. Using a specific distance function, the distance between the search keyword and the vectorized data in the database is calculated, and the closer the distance, the more relevant the data is. In this way, the user can perform data search using the vector database 130.
[0021] According to an embodiment, the vector database 130 in the cloud server 100 can support multimode search services such as text content and video content. Here, the vector database 130 provides structured data. For example, after various text contents, image contents, and audio / video contents are texturized, vectorized data is calculated using a vector algorithm. The vectorized data is not only utilized in the search service but also applied to the natural language processing process. The natural language processing process associates the vectorized data with a vector space using a natural language model. Taking the words input by the user as an example, a word vector is calculated using a vector algorithm.
[0022] According to an embodiment, in a method for triggering an intelligent conversation through actual audio / video according to the present disclosure, the function of conducting an intelligent conversation is implemented as a chatbot operating on the cloud server 100. The chatbot can communicate with the user using natural language including text and audio, not only respond to the user's input message, but also obtain user data from the cloud server 100 in advance before the conversation to understand the user's personality and habits. Furthermore, the chatbot can also obtain video by utilizing a video capture device such as the camera of the user device 150, and obtain the audio received on-site by utilizing a sound collection device such as a microphone. In addition, real-time status (for example, local weather and news based on the user's current location) is obtained from an external system (the first external system 111, the second external system 112), so that the content to which the chatbot responds not only matches the user's preferences but also can reflect the actual situation.
[0023] Furthermore, in a cloud server that executes a method for triggering an intelligent conversation via real voice and video, chatbots trained in various fields may be installed. When a user needs additional information during a conversation, the chatbot can introduce relevant field chatbots (for example, those specialized in restaurants, food courts, businesses related to night markets, products, or fields). Thereby, the relevant field chatbots can continue the conversation with the user using natural language and provide more professional and accurate conversation content.
[0024] A system for triggering an intelligent conversation via real video provides an intelligent conversation service via a real video interface, uses a machine learning method to learn data related to the user's preferences from the conversation history and the user's activities on social media, and forms structured data in the system. As shown in FIG. 2, in an example diagram showing a data structure for executing a natural language model in a system for triggering an intelligent conversation via real voice and video, the data is divided into social media platform data 21, user data (for example, user profile) 23, and user activity data 25.
[0025] The social media platform data 21 is non-public data within the system and includes viewer data 211 of users who access various contents provided by the cloud server using a system for triggering an intelligent conversation via real voice and video, content creator data 212 of users who provide various contents within the system, business data 213 in which enterprises create enterprise information and facilitate advertising in the system, and various location data 214 related to the acquisition of geographical locations for the system to provide location-based services.
[0026] User data 23 is public data within the system and includes information that can be edited by the user himself / herself. Here, the data that can be edited by the user himself / herself includes viewer data 231 obtained by the system from various user activity data. The viewer data 231 includes viewer preference data (interest data) obtained by machine learning with the user as the viewer, such as recent preference data, past preference data (history interest), and location-related preference data (location interest).
[0027] The content producer data 232 in the user data 23 is the relevant information of the user as the content producer and includes data related to the content producer's preferences and the location of the content producer obtained by the system through machine learning. For example, it includes the user's data as the content producer, the content producer's preferred type obtained by learning, geographical location, or location information including the specific location of a place.
[0028] The business data 233 in the user data 23 includes the business type of the enterprise and the characteristics of its products obtained by the system through machine learning when the user is an enterprise.
[0029] The user activity data 25 is non-public data within the system and includes statistical data on the user's activities in various services provided by the cloud server and data obtained by machine learning. Specifically, it mainly includes viewer data 251, content producer data 252, and business data 253.
[0030] The viewer data 251 is the viewing rate, viewing time, and activity data such as following, preferences, comments, subscriptions, etc. of the user using the services provided by the cloud server. The content producer data 252 is statistical data when the user acts as a content producer, and includes, for example, followers of channels and accounts, the number of views of produced content, and the viewing rate of the account. The business data 253 is followers obtained when the user is a company, the number of views of content, and overall impression data, etc.
[0031] The above-mentioned social media platform data 21, user data 23, and user activity data 25 collected and learned by the cloud server 100 serve as the basis for the dialogue service realized by utilizing the natural language processing and generation artificial intelligence technologies according to the present disclosure. Inside the cloud server 100, by calculating various data mentioned above by the processing circuit, it becomes possible to provide a chatbot according to the personalization and real-time needs of the user.
[0032] According to the embodiment, the natural language model executed in the cloud server 100 first executes a vector algorithm on the content input by the user through the dialogue interface, the user's preferences, and real-time environment information, tags the obtained text, calculates the vector of each word, searches the database based on the vector distance between words to obtain relevant content, and thereby generates dialogue content suitable for the user's preferences and real-time environment information. In the process of online dialogue, a transformer model may be used to execute processes such as machine translation, document summarization, and document generation on the texturized data. Then, the semantics of the user during the dialogue may be obtained so that the chatbot can generate dialogue content.
[0033] FIG. 3 is a flowchart of an embodiment showing a method of triggering an intelligent dialogue via real audio and video. The method of triggering an intelligent dialogue via the real audio and video is executed on a server such as the cloud server shown in FIG. 1, and can provide an augmented reality and an intelligent dialogue service via a network. Thereby, a user device installed with a corresponding application program can trigger an intelligent dialogue when executing an augmented reality video.
[0034] When an application program executed on a user device starts a real video interface, a real video request to the server is generated. After the server receives the location information and the real video request sent from the user device (step S301), it calculates a plurality of visible ranges corresponding to a plurality of viewing angles based on the location information using the location information (step S303). When the user device transmits the current location information, the software program on the server calculates a plurality of visible ranges that can be viewed from different viewing angles of the user's location, and then searches a database (the map database shown in FIG. 1) based on the real video request and the calculated plurality of visible ranges, and obtains environmental objects (for example, buildings, landmarks, tourist attractions) that can be seen in the video captured by the camera of the user device. By searching for location-based data within the visible range, the spatial position of a link image capable of displaying the location-based data and related text descriptions, etc. are specified (step S305).
[0035] Thereafter, search results such as one or more location-based data within the visible ranges obtained from different viewing angles are transmitted to the user device (step S307). In the real video interface started on the user device, one or more link images corresponding to each of the one or more location-based data within the visible range are displayed at the corresponding spatial position, and an intelligent dialogue link point is further provided.
[0036] Here, it should be noted that when the user device starts the real - world video interface, the displayed video is the surrounding real - world video captured by the camera of the user device. At this time, the software program executed on the server or the user device calculates the visible range based on the current position, that is, the position of the user can reflect the altitude (for example, on top of a building, on top of a mountain, etc.), and the shooting direction of the user device can reflect the visual angle. At this time, the visible range is displayed on the screen of the user device, and at the same time, position information and a real - world video request are generated and sent to the server. The server identifies the environmental objects that can be seen within the visible range, and combines the real - world video within the visible range captured by the camera of the user device with one or more linked images displayed at one or more spatial positions, thereby forming an augmented reality video.
[0037] Furthermore, according to the embodiment, after opening the real - world video interface on the user device, a video processing process for recognizing features of the acquired video is executed, and an audio processing process for receiving audio information and acquiring audio features is executed. In the augmented reality process, video and audio are continuously received (step S309), and an intelligent model is activated to recognize the objects in the video and various sounds in the audio information (step S311).
[0038] When the user taps on the intelligent interaction link point displayed on the augmented reality interface, the intelligent interaction interface is launched. At the same time, when the server receives an intelligent interaction request generated from the user device (step S313), it initiates an intelligent interaction process between the user device and the server (step S315). In this way, in the intelligent interaction process, a chatbot is introduced, and at the same time, the location information of the user device and one or more location-based data within the visible range are obtained. Furthermore, the objects and / or sounds at the scene obtained by recognizing video and audio information may also be included. The chatbot generates dialogue content suitable for the current situation for the user information using a natural language model based on the various information received. In this way, the user can receive a dialogue service closer to the current situation.
[0039] According to an embodiment of the flowchart shown, the user operates a software program executed on the user device to start a graphical user interface. FIG. 6 shows a schematic diagram of an embodiment of browsing location-based data in a map interface 60. A user interface with an electronic map as the background is shown, through which location-based data within different geographical ranges can be browsed. In the figure, voice / video link points 601, 602, 603 displayed at different positions are shown. At the lower part of the interface, some functions provided by the software program, such as play 611, dialogue 612, assistant 613, search 614, and return to user homepage 615, are further displayed. Furthermore, an augmented reality link point 605 is provided, and the user can start the augmented reality process by tapping on it.
[0040] When the user taps on one interesting point such as any of the audio / video link points 601, 602, 603 shown in FIG. 6, a browsing page is started to play text data, image data, or audio / video data related to the interesting point, and location-based data may be displayed on this browsing page. When tapping on the augmented reality link point 605(AR) among them, the user device activates the camera to capture the surrounding video, the real video mode is activated, and the real video interface is started. In the real video mode, based on the search results obtained by searching the database from the server, a link image of the location-based data related to the position of the interesting point may be displayed at one or more spatial positions in the real video interface.
[0041] When the user device activates the real video mode, reference may continue to be made to the schematic diagram of the embodiment showing the augmented reality interface in the method of triggering an intelligent dialogue through real audio / video shown in FIG. 7.
[0042] As an example, in the augmented reality interface 70, it is schematically shown that the real video 701 is displayed, and the user can activate the sound collection function and the shooting function through images such as the microphone 703 and the camera 705. Thereby, the application program can obtain the currently acquired video and the received audio information, and further provide them to the cloud server, and can realize the method of triggering an intelligent dialogue through real audio / video.
[0043] According to the embodiment, when the user activates the intelligent dialogue process in the real video interface, a chatbot is introduced, and the flowchart of the embodiment for processing natural language information in the method of triggering an intelligent dialogue through real audio / video shown in FIG. 4 is executed.
[0044] When the user taps on the interaction link point in the real - time video interface, the cloud server receives the selection of the intelligent interaction (step S401) and activates the said intelligent interaction process (step S403). Next, the intelligent interaction interface is started, and the user can input text content, image content, or specific audio - video content (for example, input a link where the audio - video content is shared) via the intelligent interaction interface. The cloud server receives the content input by the user via the user interface module (step S405). According to an embodiment, the said intelligent interaction process is implemented as a chatbot using a natural language model, interacts with the user via the intelligent interaction interface, and executes a natural language information processing method for each content input by the user. An input field for the user to input content is provided in the intelligent interaction interface, and an interaction display area for the interaction content output by the chatbot and the content input by the user is displayed.
[0045] At this time, the cloud server obtains the content input by the user using the user interface module. The content received via the interaction interface may be text content, audio content, or video content. When audio content or video content is received, after being converted into text using a text - conversion process, semantic analysis is performed to obtain semantic features (step S407). While the above - mentioned process is being executed, the cloud server obtains user data from the user database and obtains real - time environmental information from an external system (using the external system interface module 105 shown in FIG. 1) (step S409).
[0046] Furthermore, when the user activates the function to simultaneously acquire current video and audio information, by using an intelligent model to recognize the objects in the video and various sounds in the audio information, the cloud server can obtain more information from the user device and can further provide intelligent conversations that match the current situation and content more suitable for the user's current needs. For example, based on the characteristics of the video, it is possible to identify the people, vehicles, and events occurring around the user, and from the characteristics of the audio, it is possible to identify the music being played near the user, people's conversations, noises, and the sounds of events occurring. All of these can be used as materials in the intelligent conversation process.
[0047] In this way, the software program operating on the cloud server can determine (or search and select from the database) the semantic features that match the content input by the user, the user's preferences obtained from the user data, the real-time environment information, and the content of the current audio and video situation (step S411), perform processing by the natural language model executed in the intelligent conversation process, and generate conversation content (step S413). Then, introduce the conversation content into the intelligent conversation process and output the conversation content to the conversation interface (step S415). The flow may repeat the above steps S405 to S413.
[0048] Furthermore, when the natural language model on the cloud server operates, multi-dimensional information is recorded using the database or the system memory. The multi-dimensional information may include the past conversation history in the same intelligent conversation process. Thereby, before generating a conversation, in addition to considering the user's semantic features, the user's preferences, and the real-time environment information in the conversation as in step S411, the chatbot can also consider the user's past conversation history in the current intelligent conversation process (step S417). In this way, the conversation content generated by the natural language model (step S415) will be conversation content that matches the current situation.
[0049] Next, refer to the flowchart of another embodiment of natural language information processing shown in FIG. 5.
[0050] In the flow shown in FIG. 5, the user activates an intelligent dialogue process using an application program (step S501). Through the dialogue with the chatbot, the system receives the dialogue content input by the user (step S503), and further obtains the semantic features of the user. According to the embodiment, by utilizing the natural language processing module in the cloud server, transformation (transformer) operations and vector operations can be executed to obtain semantic features (step S505).
[0051] It should be mentioned here that in natural language information processing, it is possible to utilize artificial intelligence technology to learn natural language. After performing natural language understanding, text classification and grammar analysis are executed. Here, when processing the dialogue content input by the user, the deep learning method of the transformation model (the transformer model, proposed by the Brain team of Google (registered trademark) in 2017) can be used to process the natural language content with a temporal order input by the user. If the input content is not text content, it is necessary to convert it into text to obtain the text. In this way, in the online dialogue process, this transformation model can be utilized to execute machine translation, document summarization, document generation, etc.
[0052] After obtaining the semantic features of the user's dialogue content, the system analyzes the user's preferences and the user's current location obtained, or analyzes the location information that the user is interested in from the dialogue content, and obtains real-time environmental information from an external system in real time based on that location. Further, as can be seen from the above-described embodiments, not only does the user activate the camera to shoot a video in the real video mode, but also when the user activates a sound collection device such as a microphone to receive the current environmental sound of the user device, information such as the user's location information, video, recognized objects, and voice can be simultaneously transmitted to the cloud server (step S507). Here, the real-time environmental information may include one or an arbitrary combination of real-time weather, real-time road conditions, real-time news, and network information related to the real-time location (for example, POI on the map, evaluation of POI, etc.) obtained from one or more external systems in real time. The user device processes the video and audio information obtained using the intelligent model 155 by its software program (the audio processing process 153 and the video processing process 154 shown in FIG. 1), recognizes the environmental objects and environmental sounds around the environment in the video, and provides them to the cloud server, so that they can be used as materials for natural language processing in intelligent dialogue, and the dialogue content approaches the current situation.
[0053] After that, the system utilizes the vector database to calculate the closest answer based on the user's semantic features, the user's preferences, the real-time environmental information, the currently recognized environmental objects and voice, or in combination with the past dialogue history (step S509). It should be mentioned here that the data in the vector database is structured information obtained by a vector algorithm. Thereby, the system can extract words with similar semantics from the content obtained based on the vector distance.
[0054] In this embodiment, for the content input by the user, the content in which the user is interested, and the real-time environment information, the past conversation history is added as necessary, the vector algorithm is executed, the obtained text is tagged, and the vector of each word is calculated. Then, relevant content is obtained based on the vector distance between words, and based on this, the user's preferences, the objects and voices in the environment at the current position of the user, and conversation content suitable for the real-time environment information are generated. Further, according to the embodiment, when the vector algorithm is executed on the past conversation history recorded in the cloud server 100, conversation content suitable for the current emotion of the user can also be generated. For example, the same topic handled in the past conversation history can be continued, and terms that match the analyzed emotion can be used.
[0055] Furthermore, based on the above information, the system searches the audio-visual database to obtain appropriate audio-visual content, and adds the user's current location, the environmental objects and location-based content displayed by the real-time video interface (step S511). Then, the chatbot generates conversation content using natural language processing and generation artificial intelligence technology (step S513), and the generated conversation content is output to the conversation interface (step S515). Also, in an embodiment, by continuously executing the above steps during the chat, the chatbot can interact with the user using natural language (text or voice) and provide real-time content (audio-visual content, text content) in which the user is interested.
[0056] When entering the intelligent conversation process, as relevant embodiments, reference may be made to the conversation interface 80 shown in FIG. 8, the conversation interface 90 shown in FIG. 9, the conversation interface 1000 shown in FIG. 10, etc. In the conversation interfaces shown in each example, an input field for the user to input content is provided, and further, a conversation display area for the conversation content output by the chatbot and the user input content is provided.
[0057] As a related example, reference may be made to FIG. 8. In the dialogue interface 80 shown in FIG. 8, several pieces of dialogue content 801, 802, 803 between the user and the chatbot are displayed. Further, the chatbot may search a database based on the semantic features of the user obtained from the dialogue content 802 and provide recommended audio-visual content 804. At the lower part of the dialogue interface 80, an input field 805 is provided for the user to further input dialogue content.
[0058] As another mode, there is the dialogue interface 90 shown in FIG. 9. In this example, at the start of the online dialogue process, the system directly provides natural language dialogue content 901, 902, 904 based on the user's preferences and real-time information, and directly presents recommended audio-visual content 903. In this way, the user can continue to use the input field 905 in the dialogue interface 90 to respond to the above-mentioned dialogue content.
[0059] In the online dialogue process, the dialogue interface 1000 shown in FIG. 10 includes dialogue content 1001 generated by the chatbot based on the semantic features of the user. In this example, through semantics, the chatbot determines that the user is trying to select a specific item, and accordingly presents several recommended options 1002. Of particular note is that the chatbot provides recommended options to the user based on the real-time environmental information obtained by the system from an external system and the objects and sounds in the environment where the user is currently located.
[0060] For example, the chatbot may provide recommended option 1002 based on real-time weather information, road conditions, time, and the user's current location. Here, if it is exactly meal time, considering the user's eating habits, meal options can be provided based on restaurants that are open in the vicinity of the user's current location. In another embodiment, the cloud server acquires the user's current location information from the user device, and environmental objects and sounds recognized from environmental video and audio information. If it is determined from the environmental objects that the user is in a commercial area and it is determined from the environmental sound that the user is talking about a product they want to buy with a friend, the chatbot can provide information on related products and stores during the process of intelligent conversation, and can also recommend related audio-visual content such as the recommended option 1002 shown in the figure. Accordingly, when the user indicates that they want to view audio-visual content, the recommended option 1002 that responds during the process of intelligent conversation may be multiple recommended audio-visual contents. Also, when the user indicates that they want to find friends with common interests, the recommended option 1002 may be multiple friend recommendation links.
[0061] Furthermore, the user uses the input field 1006 to input dialogue content 1003 in response to these recommended options 1002. Thereafter, the chatbot responds with dialogue content 1004 based on the semantics of the dialogue content 1003, and proposes multiple recommended contents 1005 based on the semantics of the above-mentioned dialogue content.
Claims
1. Executed on a cloud server, opening a reality video interface on a user device, activating a camera to obtain environmental video, and / or activating a microphone to obtain environmental audio information; receiving location information and a reality video request from the user device, obtaining environmental objects by recognizing the environmental video, and / or obtaining environmental sounds by recognizing the environmental audio information; receiving a knowledge dialogue request generated by triggering a knowledge dialogue link point displayed on the reality video interface from the user device, and starting a knowledge dialogue process; in the knowledge dialogue process, starting a knowledge dialogue interface, introducing a chatbot, and executing a natural language model based on the location information, the environmental objects, and / or the environmental sounds to generate dialogue content; including A method for triggering a knowledge dialogue via real audio and video, characterized by the above.
2. The cloud server further obtains the user's preferences based on user data received from the user device, obtains real-time environmental information from one or more external systems, and causes the chatbot to generate the dialogue content based on the user's preferences and the real-time environmental information. The method for triggering a knowledge dialogue via real audio and video according to Claim 1.
3. Receiving the content input by the user via the knowledge dialogue interface, obtaining the semantic features of the content input by the user, so that the chatbot executes the natural language model based on the semantic features, the user's preferences, and / or the real-time environmental information to generate the dialogue content, and the dialogue content includes providing a plurality of recommended options, a plurality of recommended audio / video contents, and / or a plurality of friend recommendation links. The method for triggering a knowledge dialogue via real audio and video according to Claim 2.
4. Using the position information, calculate a plurality of visible ranges corresponding to a plurality of viewing angles based on the position information, search a database based on the real video request and the calculated plurality of visible ranges, and obtain one or more location-based data. Thus, the cloud server transmits link information of the one or more location-based data within each visible range to the user device. In the real video interface, based on the spatial positions of the one or more location-based data within the visible range, one or more link images linked to the one or more location-based data are displayed on the real video interface. The real video of the visible range captured by the camera of the user device is displayed on the real video interface and combined with the one or more link images displayed at one or more spatial positions, thereby forming an augmented reality video. The method for triggering an intelligent dialogue via real audio and video according to claim 1.
5. In the user device, using an intelligent model to process the acquired environmental video and environmental audio information, recognize the environmental objects and environmental sounds around the user device, and provide them to the cloud server, so that the natural language model serves as the basis for generating the dialogue content. The method for triggering an intelligent dialogue via real audio and video according to any one of claims 1 to 4.
6. Comprising a cloud server provided with a database. The cloud server Open a real video interface on the user device, activate the camera to acquire environmental video, and / or activate the microphone to acquire environmental audio information. Receive position information and a real video request from the user device, obtain environmental objects by recognizing the environmental video, and / or obtain environmental sounds by recognizing the environmental audio information. Receive a knowledge dialogue request generated by triggering a knowledge dialogue link point displayed on the real video interface from the user device, and activate a knowledge dialogue process. In the intelligent dialogue process, an intelligent dialogue interface is started, a chatbot is introduced, and a natural language model is executed based on the position information, the environmental object, and / or the environmental sound to generate dialogue content. Executing a method for triggering intelligent dialogue via real audio and video including A system for triggering intelligent dialogue via real audio and video, characterized in that. **Claim 7** The cloud server further obtains the user's preferences based on the user data received from the user device, obtains real-time environmental information from one or more external systems, and causes the chatbot to generate the dialogue content based on the user's preferences and the real-time environmental information. The system for triggering intelligent dialogue via real audio and video according to claim 6. **Claim 8** Receiving the content input by the user via the intelligent dialogue interface and obtaining the semantic features of the content input by the user, whereby the chatbot executes the natural language model based on the semantic features, the user's preferences, and / or the real-time environmental information to generate the dialogue content, and the dialogue content includes providing a plurality of recommended options, a plurality of recommended audio and video contents, and / or a plurality of friend recommendation links. The system for triggering intelligent dialogue via real audio and video according to claim 7. **Claim 9** The database includes an audio and video database that provides the user device with audio and video content uploaded and shared by each stored end user accessible via a network, and a user database that stores and updates a user profile including the past dialogue history of each user on a time axis and provides it to the natural language model for learning by a machine learning algorithm as the dialogue history. The system for triggering intelligent dialogue via real audio and video according to claim 6. **Claim 10** The database includes a vector database that stores structured data obtained by subjecting various text contents, image contents, and audio / video contents to an algorithm for vectorization and that is used for collating various information conforming to user personalization, and a map database that is used to provide each user with a function of searching for location-based data associated with a specific geographical location or spatial coordinates. The system for triggering an intelligent dialogue via real audio / video according to claim 9.
11. The cloud server calculates a plurality of visible ranges corresponding to a plurality of viewpoints based on the position information by using the position information, searches the database based on the real video request and the calculated plurality of visible ranges, and acquires one or more location-based data, and thereby transmits link information of the one or more location-based data within each visible range to the user device. In the real video interface, one or more link images linked to the one or more location-based data are displayed on the real video interface based on the spatial positions of the one or more location-based data within the visible range. The system for triggering an intelligent dialogue via real audio / video according to claim 6.
12. The user device causes the real video of the visible range captured by the camera to be displayed on the real video interface and combines it with the one or more link images displayed at one or more spatial positions, thereby forming an augmented reality video. The system for triggering an intelligent dialogue via real audio / video according to claim 11.
13. In the user device, an intelligent model is used to process the acquired environmental video and environmental audio information, recognize environmental objects and environmental sounds around the user device, and provide them to the cloud server, so that the natural language model serves as a basis for generating the dialogue content. The system for triggering an intelligent dialogue via real audio / video according to claim 6.
14. The natural language model executed in the cloud server uses a conversion model to execute processes of machine translation, document summarization, and document generation to generate the dialogue content. A system for triggering an intelligent dialogue via real audio and video according to any one of claims 6 to 13.
15. In the cloud server, for the content input by the user, the user's preferences, real-time environment information, and one or more location-based data within each visible range, further execute a vector algorithm, tag the obtained text, calculate the vector of each word, and obtain relevant content based on the vector distance between words, thereby generating the dialogue content suitable for the user's preferences and the real-time environment information. A system for triggering an intelligent dialogue via real audio and video according to claim 14.
Citation Information
Patent Citations
Server, client terminal, control method, and storage medium
WO2018066191A1
Conversation control program, conversation control method, and information processing device
WO2021205543A1