Method and system for triggering intelligent dialogue through reality video

The method and system address the limitations of current chatbots by integrating real-world video and augmented reality to provide context-aware dialogue, ensuring relevant and real-time responses through location-based data and chatbot interactions.

JP2025110881AActive Publication Date: 2025-07-29PLACY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024228051
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2024-12-24
Publication Date
2025-07-29
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Current natural language chatbots lack the ability to provide relevant and real-time responses due to their lack of linkage to the user's environment, limiting their effectiveness in providing situational answers.

Method used

A method and system that utilizes a real-world video interface to trigger intelligent conversations by calculating visible ranges, displaying location-based data, and integrating augmented reality to provide context-aware dialogue through a chatbot.

Benefits of technology

Enables contextually relevant and real-time responses by linking chatbot interactions to the user's surroundings, enhancing the relevance and accuracy of the conversation content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025110881000001_ABST
    Figure 2025110881000001_ABST
Patent Text Reader

Abstract

To provide a method and a system for triggering an intelligent dialogue through reality video.SOLUTION: A method comprises: receiving location information and a reality video request sent from a user device; calculating a plurality if visual ranges corresponding to a plurality of viewing angles from the location information; searching a database based on the plurality of calculated visual ranges to obtain location-based data; sending link information on the location-based data to the user device; making a reality video interface display one visual range; displaying one or more link images linked to the location-based data; and providing an intelligent dialogue link point. When an intelligent dialogue request is generated by the user triggering the intelligent dialogue link point, an intelligent dialogue process is activated, a chatbot is introduced, and a dialogue content is generated to enable a response to a content input by the user.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method of initiating an intelligent conversation in a browsing interface, and particularly to a method and system for triggering an intelligent conversation via a real-world video that provides a service for initiating an intelligent conversation when a user browses the surrounding environment using the real-world video.

Background Art

[0002] ChatGPT can interact with users in the form of natural language by learning and training on a large amount of information on the Internet. However, the general content for answering users is the standard answers obtained through learning, and it is not possible to provide users with answers that respond to the current situation in real time. Although it is called a natural language chatbot, it lacks content corresponding to the relevance with users and real-time situations.

Summary of the Invention

Problems to be Solved by the Invention

[0003] In addition to the above-mentioned drawbacks, the services provided by current natural language chatbots are limited to general conversations and cannot meet all needs. For example, since they are not linked to the real environment in applications, they cannot provide effective answers according to the user's situation at that time.

Means for Solving the Problems

[0004] To provide a new form of method for triggering an intelligent conversation, the present disclosure provides a method and system for triggering an intelligent conversation via a real-world video. This system includes a server and a database, and uses the server to provide that a user device connects and is displayed in a real-world video, obtains a link image with location-based data and related information, and provides a link for triggering an intelligent conversation in the real-world video.

[0005] In a method for triggering an intelligent dialogue via a real-world video, after receiving the location information and the real-world video request transmitted from a user device, based on the location information, a plurality of visible ranges corresponding to a plurality of viewpoints are calculated from the received location information, and a database is searched based on the real-world video request and the calculated plurality of visible ranges, and one or more location-based data are obtained.

[0006] After that, after transmitting the link information of one or more location-based data within each visible range to the user device, a real-world video interface is activated on the user device, and based on the spatial position of each of the one or more location-based data within the obtained visible range, one or more link images linked to the one or more location-based data are displayed on the real-world video interface, and an intelligent dialogue link point is provided.

[0007] When the user taps the intelligent dialogue link point, an intelligent dialogue request is generated, an intelligent dialogue process is activated between the server and the user device, a chatbot is introduced, and at the same time, the location information and one or more location-based data within the visible range are obtained, and an intelligent dialogue interface is started on the user device to conduct a dialogue.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Modes for Carrying Out the Invention

[0009] The present disclosure provides a method and system for triggering an intelligent dialogue via a real-world video. Here, by using a user device to activate the camera function and activate the real-world video interface, the current environmental video can be captured and displayed, and based on the location information sent to the server, corresponding location-based data can be obtained from the server, and link images of content can be displayed at spatial positions related to the location-based data in the real-world video interface. One specific embodiment is to display the real-world video within the visible range captured by the camera (e.g., the surrounding environmental video) in the real-world video interface, and combine one or more link images displayed at one or more spatial positions to form an augmented reality image. Thereby, while the user operates the user device to view the environmental video from the display, the user can also view the link images and character information displayed at the spatial coordinates by the augmented reality (AR) technology. In particular, a link point that can trigger the intelligent dialogue process by the user tapping is additionally provided, achieving the purpose of the method for triggering the intelligent dialogue via the real-world video.

[0010] According to the embodiment, the system that triggers intelligent interaction through the real video can provide social media services through a network by a cloud server, and after joining, the user can share text content, image content, and audio / video content. Thereby, the location-based data obtainable from the real video can include text content, image content, audio / video content, etc. shared by the user. Furthermore, the cloud server provides chatbots in each field, and the user can have conversations using the services provided by the cloud server. Here, the utilization of artificial intelligence technology includes chatbots that provide conversation services by learning and training data in each field using machine learning algorithms and natural language processing (NLP) technology, and by learning the user's activities on social media to obtain the user's preferences. The chatbot is based on the semantics of the user's conversation and the user's preferences, and combines the environment information obtained in real time to provide conversation content more suitable for the user's individual needs and current environmental characteristics.

[0011] FIG. 1 is a diagram of an embodiment showing the structure of a system that triggers intelligent interaction through real video. The system structure shown in the figure mainly includes a system that triggers intelligent interaction through real video implemented by a server, a database, and related software and hardware on the server side, and an application program executed by a device on the user side.

[0012] The cloud server 100 shown in the figure is constructed from a computer system, a database, and a network, and as various functional modules realized by the cooperation of software and hardware, as shown in the figure, it has a natural language processing module 101, a machine learning module 103, an external system interface module 105, and a user interface module 107. The natural language processing module 101 that processes natural language information is realized by a chatbot having natural language processing capabilities. The machine learning module 103 not only executes machine learning algorithms and trains natural language models, but also learns the user's online activities using deep learning methods to obtain user preference information, so that the chatbot can provide dialogue content according to the user's preferences. The cloud server 100 includes an external system interface module 105. The external system interface module 105 executes a circuit and related application software for connecting to an external system (for example, the first external system 111, the second external system 112) (for example, via the network 10) and obtaining data via an application programming interface (API). The cloud server 100 includes a user interface module 107. By providing a connection between the user device 150 and the cloud server 100 through the network connection function in the user interface module 107, executing a server (web server) for providing web services, and executing an application program for correspondingly obtaining services on the user device 150, the services provided by the cloud server 100 can be obtained.

[0013] According to the embodiment, an augmented reality module 109 is implemented in the cloud server 100 using augmented reality (AR) technology. The augmented reality module 109 can, in response to a real-time video request generated from the user device 150, display the video of the surrounding environment via a display when the user activates the augmented reality process, and at the same time, can also view the objects in the augmented reality integrated with the augmented reality video provided by the cloud server 100. For example, when the real-time video interface 115 is started on the user device 150, the operation of the augmented reality module 109 of the cloud server 100 acquires the link images of the location-based data displayed at one or more spatial positions within the visible range and obtains the corresponding information. For example, when the user activates the camera of the user device 150 to acquire and display the surrounding environmental video, the current location information of the user device 150 is also transmitted to the cloud server 100. As a result, the software program in the cloud server 100 can identify the user's current location and the surrounding environmental objects (such as buildings, roads, scenery, etc.) that can be seen in that space, and can obtain the location-based data related to these environmental objects by querying the database. Subsequently, this data is transmitted to the user device 150, integrated into the augmented reality video via the real-time video interface 115, and the link images corresponding to the location-based data are displayed at the corresponding spatial coordinates.

[0014] According to the structure shown in the figure, the cloud server 100 is equipped with a built-in or externally connected database. Thereby, the cloud server 100 provides data services. The audio-visual database 110 shown in the figure provides audio-visual content stored and shared by each end-user accessible by the user device 150 via the network 10. What the user uploads and shares may include text content and image content. The user database 120 stores user data including the user's personal information, uploaded text content, image content, and audio-visual content, and acquires the user's activity data (e.g., network activities such as content viewed, followed, liked, shared, subscribed) in the network service provided by the cloud server 100. Based on these, a user profile is created. Furthermore, when interactive content is continuously generated over time, the user database 120 may store and update user data including the user's past interaction history on the timeline and provide it to the natural language model as an interaction history for learning by the machine learning algorithm. The vector database 130 stores structured data obtained by computing various text content, image content, and audio-visual content through vectorization, and may be used for matching various information that matches the user's personalization.

[0015] The database may further include a map database 140. The map database 140 is for a user to search for location-based data associated with a specific geographical location or spatial coordinates. For example, it is associated with a country, administrative region, tourist destination, landmark, and places tagged by the user at a specific geographical location. In particular, it may store not only data on geographical locations on a plane but also location-based data including spatial coordinates. For example, a restaurant installed on a certain floor of a building is given geographical location (e.g., longitude, latitude) and spatial coordinates including altitude information (e.g., spherical coordinate system described by radius distance (γ), polar angle (θ), and azimuth angle (φ), or rectangular coordinate system described by X-axis, Y-axis, and Z-axis). Furthermore, since the map database 140 can be used to search for environmental objects with altitude information, the system can identify environmental objects that a user can see through the user device 150 based on the location information (e.g., elevation) of the user device 150 and the shooting direction (combination of azimuth angle and polar angle) of the user device 150, and form the user's visible range.

[0016] According to the schematic diagram of the structure of the system shown in the figure, the cloud server 100 may obtain data from an external system via the network 10 or a connection of a specific protocol. The external systems are the first external system 111 and the second external system 112 schematically shown in the figure, which are, for example, servers installed by a government or a company and used to provide open data. Thereby, the cloud server 100 can obtain real-time information that meets the needs, such as real-time weather information, real-time road conditions, real-time news, and network information related to the real-time location, via the application programming interfaces provided by the external systems respectively using the external system interface module 105.

[0017] The user device 150 runs an application capable of obtaining the services provided by the cloud server 100. For example, for the social media service provided by the cloud server 100, the user device 150 runs the corresponding social media application to obtain the social media service via the user interface module 107. In particular, the cloud server 100 provides a natural language chatbot via the natural language processing module 101, enabling the user to interact with the chatbot using the intelligent dialogue interface initiated in the augmented reality interface 115. On the other hand, the cloud server 100 uses the machine learning module 103 to learn the activity data of the user using various services of the cloud server 100. The activity data includes the activity data obtained by the user using the social media application and the augmented reality interface 115. In this way, the machine learning module 103 can learn the user's interests and characteristics and construct user data.

[0018] It should be noted here that the cloud server 100 obtains various text contents, image contents, and audio / video contents as unstructured data, converts them into vectorized data using an encoding method, making it easier to obtain the meaning of the content and convenient for data search. Furthermore, the vectorized data may be used when collating with the user's search keywords. Using a specific distance function, the distance between the search keywords and the vectorized data in the database is calculated, indicating that the closer the distance, the more relevant the data. In this way, the user can perform data search using the vector database 130.

[0019] According to an embodiment, the vector database 130 in the cloud server 100 can support multi-mode search services such as text content and video content. Here, the vector database 130 provides structured data. For example, after various text contents, image contents, and audio / video contents are texturized, vectorized data is calculated using a vector algorithm. The vectorized data is not only utilized in the search service but also applied to the natural language processing process. The natural language processing process associates the vectorized data with a vector space using a natural language model. Taking the words input by the user as an example, a word vector is calculated using a vector algorithm.

[0020] According to an embodiment, in a method for triggering an intelligent conversation through a real video according to the present disclosure, the function of performing the intelligent conversation is implemented as a chatbot operating on the cloud server 100. The chatbot can interact with the user using natural language including text and voice. It can not only respond to the user's input message but also obtain the user's data from the cloud server 100 in advance before the conversation to understand the user's personality and habits. In addition, it can obtain real-time statuses (e.g., local weather and news based on the user's current location) from external systems (the first external system 111 and the second external system 112), so that the content to which the chatbot responds not only matches the user's preferences but also reflects the real situation.

[0021] Furthermore, in a cloud server that executes a method for triggering an intelligent conversation through a real video, chatbots trained in various fields may be installed. When the user needs further information during the conversation, the chatbot can introduce chatbots related to the relevant field (e.g., those specialized in restaurants, food courts, night market-related operations, products, or fields). Thereby, the chatbots related to the relevant field continue the conversation with the user using natural language and provide more professional and accurate conversation content.

[0022] A system that triggers an intelligent dialogue through real-world video provides an intelligent dialogue service via a real-world video interface, uses machine learning methods to learn data on user preferences from the dialogue history and the user's activities on social media, and forms structured data in the system. As shown in FIG. 2, in a system that triggers an intelligent dialogue through real-world video, in an example diagram showing a data structure for executing a natural language model, the data is divided into social media platform data 21, user data (e.g., user profile) 23, and user activity data 25.

[0023] The social media platform data 21 is non-public data within the system and includes viewer data 211 of users who access various contents provided by the cloud server using the system that triggers an intelligent dialogue through real-world video, content creator data 212 of users who provide various contents within the system, business data 213 in which enterprises create enterprise information and facilitate advertising in the system, and various location data 214 related to the acquisition of geographical locations for the system to provide location-based services.

[0024] The user data 23 is public data within the system and includes information that can be edited by the user himself / herself. Here, the data that can be edited by the user himself / herself includes viewer data 231 obtained by the system from various user activity data. The viewer data 231 includes viewer preference data (interest data) obtained by machine learning, such as recent preference data, past preference data (history interest), and location-related preference data (location interest), with the user as the viewer.

[0025] The content producer data 232 in the user data 23 is the relevant information of the user as the content producer, and the system includes data related to the preferred type of the content producer and the location of the content producer through machine learning. For example, it includes the data of the user as the content producer, the preferred type of the content producer obtained through learning, geographical location or location information including the specific location of a place.

[0026] The business data 233 in the user data 23 includes the business type of the enterprise and the features of its products obtained by the system through machine learning when the user is an enterprise.

[0027] The user activity data 25 is non-public data within the system, and includes statistical data on the activities of the user in various services provided by the cloud server and data obtained through machine learning. Specifically, it mainly includes viewer data 251, content producer data 252, and business data 253.

[0028] The viewer data 251 is the viewing rate, viewing time of the user using the services provided by the cloud server, and activity data such as following, liking, commenting, subscribing, etc. The content producer data 252 is the statistical data when the user acts as a content producer, and includes, for example, the followers of the channel or account, the number of views of the produced content, and the viewing rate of the account. The business data 253 is the followers obtained when the user is an enterprise, the number of views of the content, and the overall impression data, etc.

[0029] The above-mentioned social media platform data 21, user data 23, and user activity data 25 collected and learned by the cloud server 100 serve as the basis for the dialogue service realized by utilizing the natural language processing and generation artificial intelligence technologies according to the present disclosure. Inside the cloud server 100, various data described above are calculated by a processing circuit, making it possible to provide a chatbot according to the personalization and real-time needs of the user.

[0030] According to an embodiment, the natural language model executed in the cloud server 100 first executes a vector algorithm on the content input by the user via the dialogue interface, the user preference, and the real-time environment information, tags the obtained text, calculates the vector of each word, searches the database based on the vector distance between words to obtain relevant content, thereby generating dialogue content suitable for the user preference and the real-time environment information. In the process of online dialogue, a process such as machine translation, document summarization, and document generation may be executed on the texturized data using a transformer model. Then, the semantics of the user during the dialogue may be obtained so that the chatbot can generate dialogue content.

[0031] FIG. 3 is a flowchart of an embodiment showing a method of triggering an intelligent dialogue via a real-world video. The method of triggering an intelligent dialogue via the real-world video is executed on a server such as the cloud server shown in FIG. 1, and can provide an augmented reality and intelligent dialogue service via a network. Thereby, a user device installed with a corresponding application program can trigger an intelligent dialogue when executing an augmented reality video.

[0032] When an application program running on a user device starts a real-world video interface, it generates a real-world video request to the server. After the server receives the location information and the real-world video request sent from the user device (step S301), it calculates a plurality of visible ranges corresponding to a plurality of viewpoints based on the location information using the location information (step S303). When the user device transmits the current location information, the software program on the server calculates a plurality of visible ranges that can be seen from different viewpoints of the user's location. Next, based on the real-world video request and the calculated plurality of visible ranges, it searches a database (the map database shown in FIG. 1), and obtains environmental objects (for example, buildings, landmarks, tourist attractions) that can be seen in the video captured by the camera of the user device. By searching for location-based data within the visible range, it identifies the spatial positions of link images capable of displaying the location-based data and related text descriptions, etc. (step S305).

[0033] After that, it transmits search results such as one or more location-based data within the visible ranges obtained from different viewpoints to the user device (step S307). In the real-world video interface started on the user device, one or more link images corresponding to each of the one or more location-based data within the visible range are displayed at the corresponding spatial positions, and furthermore, intelligent interaction link points are provided.

[0034] Here, it should be noted that when the user device starts the real - world video interface, the displayed video is the surrounding real - world video captured by the camera of the user device. At this time, the software program executed on the server or the user device calculates the visible range based on the current position, that is, the user's position can reflect the altitude (for example, on top of a building, on top of a mountain, etc.), and the shooting direction of the user device can reflect the visual angle. At this time, the visible range is displayed on the screen of the user device, and at the same time, the location information and the real - world video request are generated and sent to the server. The server identifies the environmental objects that can be seen within the visible range, and combines the real - world video within the visible range captured by the camera of the user device with one or more linked images displayed at one or more spatial positions, thereby forming an augmented reality video.

[0035] When the user taps on the intelligent interaction link point displayed on the real - world video interface, the intelligent interaction interface is started. At the same time, when the server receives the intelligent interaction request generated from the user device (step S309), it activates the intelligent interaction process between the user device and the server (step S311). In this way, in the intelligent interaction process, a chatbot is introduced, and at the same time, the location information and one or more location - based data within the visible range are obtained, and the user interacts with the chatbot in the intelligent interaction interface.

[0036] Figure 4 is a flowchart showing an embodiment of another display mode of the method for triggering intelligent interaction via real - world video.

[0037] According to the embodiment of the flowchart shown, the user operates a software program executed on the user device to start a graphical user interface. FIG. 7 shows a schematic diagram of an embodiment of browsing location-based data in the map interface 70, where a user interface with an electronic map as the background is shown, through which location-based data within different geographical ranges can be browsed. In the figure, voice / video link points 701, 702, 703 displayed at different positions are shown, and at the lower part of the interface, several functions provided by the software program, such as play 711, dialogue 712, assistant 713, search 714, and return to user homepage 715, are further displayed.

[0038] In the flow shown in FIG. 4, the server receives a signal of one point of interest (POI) tapped by the user from the user device (step S401). The server searches the database based on the POI tapped by the user and can obtain information of the point of interest including, for example, voice / video content, map data, dialogue groups, etc. (step S403). At the same time, since the user taps one point of interest such as the voice / video link points 701, 702, 703 shown in FIG. 7, a browse page for playing text data, image data, or voice / video data related to the point of interest is started, and location-based data may be displayed on this browse page. When tapping the augmented reality link point 705 (AR) among them, the user device activates the camera to capture the surrounding video, activates the real video mode, and starts the real video interface. In the real video mode, based on the search results obtained by searching the database from the server, link images of location-based data related to the position of the point of interest may be displayed at one or more spatial positions in the real video interface (step S405).

[0039] When the user device activates the real - world video mode, an augmented reality interface 80 is displayed, where the user device captures and displays the surrounding video with the camera as shown in FIG. 8. In this augmented reality interface 80, for example, the shown video link points 801, 802 are corresponding link images displayed based on the spatial positions of the respective location - based data, and may include text descriptions. Through this link image, there is a relevance between the location - based data and one of the environmental objects within the visible range. In this example, three link points are displayed at the lower part of the augmented reality interface 80, all with thumbnails shown. For example, the audio - video link points 801, 802 and the interaction link point 803 are illustrated.

[0040] When the user taps the interaction link point 803 in the augmented reality interface 80, an intelligent interaction process is activated. At this time, the server receives an intelligent interaction request from the user device (step S407) and activates an intelligent interaction process between the user device and the server (step S409). At this time, in the intelligent interaction interface activated on the user device, a chatbot is introduced, and at the same time, the location information of the user device and one or more location - based data within the visible range are acquired. The intelligent interaction interface is started on the user device, enabling interaction.

[0041] Continuing, refer to the augmented reality interface 90 shown in FIG. 9. In this example, within the visible range captured by the camera of the user device, there is a building that occupies most of the screen. As a result of calculations by a software program on the server, it can be determined that this environmental object is included within the visible range formed from a certain viewing angle at the user's current location. Therefore, the link image restricted by the environmental object and displayed by the augmented reality interface 90 has an audio - video link point 901 displayed on this building. Furthermore, an interaction link point 902 is also provided, and prompt text 903 for providing an intelligent interaction service is also included.

[0042] According to an embodiment, in the intelligent dialogue process, the example flowchart of natural language information processing shown in FIG. 5 is executed by a chatbot.

[0043] When the user taps on a dialogue link point (dialogue link point 803 shown in FIG. 8) in the real-time video interface, the cloud server receives the selection of the intelligent dialogue (step S501) and activates the intelligent dialogue process (step S503). Next, the intelligent dialogue interface is started, and the user can input text content, image content, or specific audio / video content (for example, input a link where the audio / video content is shared) through the intelligent dialogue interface. The cloud server receives the content input by the user through the user interface module (step S505). According to an embodiment, the intelligent dialogue process is implemented as a chatbot using a natural language model, interacts with the user through the intelligent dialogue interface, and executes the natural language information processing method for each piece of content input by the user. The intelligent dialogue interface provides an input field for the user to input content, and a dialogue display area for the dialogue content output by the chatbot and the content input by the user is displayed.

[0044] At this time, the cloud server acquires the content input by the user using the user interface module. The content received via the dialogue interface may be text content, voice content, or video content. When voice content or video content is received, after being converted into text using a text conversion process, semantic analysis is performed to obtain semantic features (step S507). While the above process is being executed, the cloud server acquires user data from the user database and acquires real-time environmental information from an external system (using the external system interface module 105 shown in FIG. 1) (step S509).

[0045] Thereafter, it is possible to determine (or search and select from the database) the semantic features matching the content input by the user, the user preferences obtained from the user data, and the content of the real-time environmental information (step S511), perform processing by a natural language model executed in the intelligent dialogue process, and generate dialogue content (step S513). Thereafter, the dialogue content is introduced into the intelligent dialogue process, and the dialogue content is output to the dialogue interface (step S515). The flow may repeat steps S505 to S513 described above.

[0046] Furthermore, when the natural language model in the cloud server operates, multi-dimensional information is recorded using the database or the system memory. The multi-dimensional information may include the past dialogue history in the same intelligent dialogue process. Thereby, before generating a dialogue, in addition to considering the user's semantic features, user preferences, and real-time environmental information in the dialogue as in step S511, the chatbot can also consider the user's past dialogue history in the current intelligent dialogue process (step S517). In this way, the dialogue content generated by the natural language model (step S515) becomes dialogue content that matches the current situation.

[0047] Next, refer to the flowchart of another embodiment of natural language information processing shown in FIG. 6.

[0048] In the flow shown in FIG. 6, the user starts an intelligent dialogue process using an application program (step S601). Through the dialogue with the chatbot, the system receives the dialogue content input by the user (step S603), and further obtains the semantic features of the user. According to the embodiment, by utilizing the natural language processing module in the cloud server, the transformer operation and the vector operation can be executed to obtain the semantic features (step S605).

[0049] It should be mentioned here that in natural language information processing, it is possible to utilize artificial intelligence technology to learn natural language. After performing natural language understanding, text classification and grammar analysis are executed. Here, when processing the dialogue content input by the user, the deep learning method of the transformer model (proposed by Google's (registered trademark) Brain team in 2017) can be used to process the natural language content with the time sequence input by the user. If the input content is not text content, it is necessary to convert it into text to obtain the text. In this way, in the online dialogue process, this transformer model can be utilized to execute machine translation, document summarization, document generation, etc.

[0050] After obtaining the semantic features of the user's dialogue content, based on the user preferences and the current location of the user obtained by the system or by analyzing the location information that the user is interested in from the dialogue content, the system obtains real-time environmental information from an external system in real time based on that location (step S607). Here, the real-time environmental information may include one or an arbitrary combination of real-time weather, real-time road conditions, real-time news, and network information related to the real-time location (for example, POIs on a map, evaluations of POIs, etc.) obtained from one or more external systems in real time.

[0051] After that, the system utilizes a vector database to calculate the closest answer based on the user's semantic features, user preferences, and real-time environmental information, or in combination with past dialogue history (step S609). It should be noted here that the data in the vector database is structured information obtained by a vector algorithm. Thereby, the system can extract words with similar semantics from the content obtained based on the vector distance. As a specific example, the vector distance between the word "computer" that appears in the dialogue content and the word "calculation" in the database is relatively close, but the vector distance between the word "computer" and the word "running" is relatively far.

[0052] In this embodiment, for the content input by the user, the content in which the user is interested, and the real-time environment information, the past conversation history is added as necessary to execute a vector algorithm, the obtained text is tagged, and the vectors of each word are calculated. Then, relevant content is obtained based on the vector distance between words, and based on this, conversation content suitable for the user's preferences and real-time environment information is generated. Furthermore, according to the embodiment, when a vector algorithm is executed on the past conversation history recorded in the cloud server 100, conversation content suitable for the user's current emotion may be generated. For example, the same topic handled in the past conversation history may be continued, and terms that match the analyzed emotion may be used.

[0053] Furthermore, based on the above information, the system searches the audio-visual database to obtain appropriate audio-visual content, and adds the user's current location, environmental objects, and location-based content displayed by the real-time video interface (step S611). Then, the chatbot generates conversation content using natural language processing and generation artificial intelligence technology (step S613), and the generated conversation content is output to the conversation interface (step S615). Also, in an embodiment, by continuously executing the above steps during the chat, the chatbot may interact with the user using natural language (text or voice) and provide real-time content (audio-visual content, text content) in which the user is interested.

[0054] When entering the intelligent conversation process, as relevant embodiments, reference may be made to the conversation interface 1000 shown in FIG. 10, the conversation interface 1100 shown in FIG. 11, the conversation interface 1200 shown in FIG. 12, and the like. In the conversation interface shown in each example, an input field for the user to input content is provided, and further, a conversation display area for the conversation content output by the chatbot and the user input content is provided.

[0055] As a related example, reference may be made to FIG. 10. In the dialogue interface 1000 shown in FIG. 10, several pieces of dialogue content 1001, 1002, 1003 between the user and the chatbot are displayed. Also, the chatbot may search a database based on the semantic features of the user obtained from the dialogue content 1002 and provide recommended audio-visual content 1004. At the bottom of the dialogue interface 1000, an input field 1005 is provided for the user to input further dialogue content.

[0056] As another mode, there is the dialogue interface 1100 shown in FIG. 11. In this example, at the start of the online dialogue process, the system directly provides natural language dialogue content 1101, 1102, 1104 and directly presents recommended audio-visual content 1103 based on user preferences and real-time information. In this way, the user can continue to use the input field 1105 in the dialogue interface 1100 to respond to the above-mentioned dialogue content.

[0057] According to another embodiment, the dialogue content generated by the natural language model operating in the chatbot based on the semantic features of the content input by the user, user preferences, and real-time environmental information may include multiple recommended options, multiple recommended audio-visual contents, and / or multiple friend recommendation links. Reference may be made to the example shown in FIG. 12.

[0058] In an online interaction process, the interaction interface 1200 shown in FIG. 12 includes interaction content 1201 generated by the chatbot based on the semantic features of the user. In this example, through semantics, the chatbot determines that the user is trying to select a specific item, and accordingly presents several recommended options 1202. Of particular note is that the chatbot provides recommended options to the user based on real-time environmental information obtained by the system from an external system. For example, the chatbot may provide recommended options 1202 based on real-time weather information, road conditions, time, and the user's current location. Here, if it is exactly meal time, considering the user's eating habits, meal options can be provided based on restaurants that are open near the user's current location.

[0059] On the other hand, when the user indicates that they want to view audio-visual content, the recommended options 1202 may be multiple recommended audio-visual contents. Also, when the user indicates that they want to find friends with common interests, the recommended options 1202 may be multiple friend recommendation links.

[0060] Furthermore, the user uses the input field 1206 to input interaction content 1203 in response to these recommended options 1202. Subsequently, the chatbot responds with interaction content 1204 based on the semantics of the interaction content 1203, and proposes multiple recommended contents 1205 based on the semantics of the above-mentioned interaction content. Continuing with the above example, when the user answers that they want to eat one of the dishes, the chatbot proposes options for restaurants corresponding to the dish the user wants to eat based on real-time weather information, road conditions, and the user's current location obtained by the system from an external system. If the weather is bad and the road is congested, it recommends restaurants that the user can easily reach accordingly.

Claims

1. Executed on a server, Receiving location information and a real-time video request sent from a user device; Calculating a plurality of visible ranges corresponding to a plurality of viewing angles based on the location information using the location information; Searching a database based on the real-time video request and the calculated plurality of visible ranges to obtain one or more location-based data; Sending link information of the one or more location-based data within each visible range to the user device, starting a real-time video interface on the user device, and based on the spatial positions of the one or more location-based data within the visible range, one or more link images linked to the one or more location-based data are displayed on the real-time video interface, providing intelligent interaction link points; Receiving an intelligent interaction request generated by triggering the intelligent interaction link point displayed on the real-time video interface from the user device and starting an intelligent interaction process; In the intelligent interaction process, introducing a chatbot, simultaneously obtaining the location information and the one or more location-based data within the visible range, and starting an intelligent interaction interface on the user device to conduct an interaction; Including A method for triggering intelligent interaction via real-time video, characterized by the above.

2. In the real-time video interface, a real-time video of the visible range captured by the camera of the user device is displayed, and by combining with the one or more link images displayed at one or more spatial positions, an augmented reality video is formed. The method for triggering intelligent interaction via real-time video according to Claim 1.

3. In the real-time video interface, the link image and text description corresponding to the spatial position of each location-based data are displayed, and there is a relevance between the location-based data linked to the link image and one of the environmental objects within the visible range. The method for triggering intelligent interaction via real-time video according to Claim 2.

4. In the intelligent dialogue process, the chatbot generates location-based dialogue content related to the position of the user device based on the one or more location-based data within the visible range and the one or more recognized environmental objects. The method for triggering intelligent dialogue via a real-world video according to claim 3.

5. In the intelligent dialogue process, by the chatbot, Receiving the content input by the user via the intelligent dialogue interface; Obtaining the semantic features of the content input by the user; Obtaining user data and real-time environmental information; Based on the semantic features of the content input by the user, the user preferences obtained from the user data, and the content of the real-time environmental information, further referring to the one or more location-based data within each visible range, executing a natural language model, and generating dialogue content; Introducing the dialogue content into the intelligent dialogue process and outputting the dialogue content to the intelligent dialogue interface; is executed, The content received via the intelligent dialogue interface is text content, voice content, or video content. When voice content or video content is received, it is converted into text by a text conversion process, and the semantic features of the text are further obtained by semantic analysis. The natural language model executed on the server uses a conversion model to execute processes of machine translation, document summarization, and document generation to generate the dialogue content. The method for triggering intelligent dialogue via a real-world video according to any one of claims 1 to 4.

6. Comprising a server provided with a database, The server, Receiving position information and a real-world video request transmitted from a user device; Calculating a plurality of visible ranges corresponding to a plurality of viewing angles based on the position information using the position information; Searching the database based on the real-world video request and the calculated plurality of visible ranges to obtain one or more location-based data. Transmit the link information of the one or more location-based data within each visible range to the user device, start a reality video interface on the user device, and based on the spatial positions of the one or more location-based data within the visible range, one or more link images linked to the one or more location-based data are displayed on the reality video interface, and an intelligent dialogue link point is provided, Receive an intelligent dialogue request generated by triggering the intelligent dialogue link point displayed on the reality video interface from the user device, and start an intelligent dialogue process, In the intelligent dialogue process, a chatbot is introduced, and at the same time, the position information and the one or more location-based data within the visible range are acquired, and an intelligent dialogue interface is started on the user device to conduct a dialogue, Execute a method for triggering intelligent dialogue through a reality video including the above, A system for triggering intelligent dialogue through a reality video, characterized in that.

7. In the reality video interface, the reality video of the visible range captured by the camera of the user device is displayed, and by combining with the one or more link images displayed at one or more spatial positions, an extended reality video is formed, The link image and text description corresponding to the spatial position of each location-based data are displayed, and there is a relevance between the location-based data linked to the link image and one of the environmental objects within the visible range, The system for triggering intelligent dialogue through a reality video according to claim 6.

8. Search the map data and audio / video data in the database to obtain one or more environmental objects within the visible range and the location-based data corresponding to each environmental object, so that the link image and explanatory text are accurately displayed at the spatial position corresponding to each location-based data, The system for triggering intelligent dialogue through a reality video according to claim 7.

9. In the intelligent dialogue process, the chatbot generates location-based dialogue content related to the location of the user device based on the one or more location-based data within the visible range and the one or more recognized environmental objects. A system for triggering intelligent dialogue via a real-world video according to claim 8.

10. The server obtains real-time environmental information by providing an external system interface for connecting to one or more external systems, and the chatbot further generates the location-based dialogue content based on the real-time environmental information. A system for triggering intelligent dialogue via a real-world video according to claim 9.

11. In the intelligent dialogue process, by the chatbot Receiving the content input by the user via the intelligent dialogue interface; Obtaining the semantic features of the content input by the user; Obtaining user data and real-time environmental information; Based on the semantic features of the content input by the user, the user preferences obtained from the user data, and the content of the real-time environmental information, further referring to the one or more location-based data within each visible range, executing a natural language model, and generating dialogue content; Introducing the dialogue content into the intelligent dialogue process and outputting the dialogue content in the intelligent dialogue interface; are executed. A system for triggering intelligent dialogue via a real-world video according to any one of claims 6 to 10.

12. The content received via the intelligent dialogue interface is text content, voice content or video content. When voice content or video content is received, it is converted into text by a textification process, and the semantic features of the text are further obtained by semantic analysis. A system for triggering intelligent dialogue via a real-world video according to claim 11.

13. The natural language model executed on the server uses a conversion model to execute processes of machine translation, document summarization and document generation to generate the dialogue content. A system for triggering an intelligent dialogue via a real-world video according to claim 11.

14. In the server, for the content input by the user, the user preferences, the real-time environment information, and the one or more location-based data within each visible range, further execute a vector algorithm, tag the obtained text, calculate the vector of each word, and obtain relevant content based on the vector distance between words, thereby generating the dialogue content suitable for the user preferences and the real-time environment information. A system for triggering an intelligent dialogue via a real-world video according to claim 13.

15. Based on the semantic features of the content input by the user, the user preferences, the real-time environment information, and the one or more location-based data within each visible range, the dialogue content generated by the natural language model executed in the chatbot includes a plurality of recommended options, a plurality of recommended audio / video contents, and / or a plurality of friend recommendation links. A system for triggering an intelligent dialogue via a real-world video according to claim 11.

Citation Information

Patent Citations

  • Augmented reality (AR) imprinting method and system

    JP2022507502A

  • Information processing device and information processing method

    JP7082219B1

  • Shopping system using virtual reality technology capable of improving the real feeling of consumer shopping and increasing consumption pleasure

    TW202026996A

  • Method and system of an augmented / virtual reality platform

    US20210209676A1

  • System and method for location relevant augmented reality based communication

    WO2021064747A1