Method and system for triggering intelligent dialogue through real-world video

The method and system leverage real-world images to integrate location-based data and augmented reality for personalized chatbot interactions, addressing the limitations of current chatbots by providing relevant and context-aware responses.

JP7744706B2Active Publication Date: 2025-09-26PLACY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024228051
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2024-12-24
Publication Date
2025-09-26
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Current natural language chatbots lack the ability to provide relevant answers tailored to a user's real-time situation due to their inability to integrate with the real world, limiting their effectiveness in providing personalized interactions.

Method used

A method and system that utilizes real-world images to trigger intelligent dialogue by linking location-based data with augmented reality, enabling a chatbot to provide personalized interactions through a user's device, integrating real-time environmental information and user preferences.

Benefits of technology

Enables personalized and context-aware interactions by combining real-time environmental data and user preferences, enhancing the relevance and effectiveness of chatbot responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007744706000001
    Figure 0007744706000001
  • Figure 0007744706000002
    Figure 0007744706000002
  • Figure 0007744706000003
    Figure 0007744706000003
Patent Text Reader

Abstract

To provide a method and a system for triggering an intelligent dialogue through reality video.SOLUTION: A method comprises: receiving location information and a reality video request sent from a user device; calculating a plurality if visual ranges corresponding to a plurality of viewing angles from the location information; searching a database based on the plurality of calculated visual ranges to obtain location-based data; sending link information on the location-based data to the user device; making a reality video interface display one visual range; displaying one or more link images linked to the location-based data; and providing an intelligent dialogue link point. When an intelligent dialogue request is generated by the user triggering the intelligent dialogue link point, an intelligent dialogue process is activated, a chatbot is introduced, and a dialogue content is generated to enable a response to a content input by the user.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a method for initiating intelligent dialogue in a browsing interface, and more particularly to a method and system for triggering intelligent dialogue via real-world images that provides a service for initiating intelligent dialogue when a user uses real-world images to browse their surrounding environment. [Background technology]

[0002] ChatGPT is trained by learning from a vast amount of information on the Internet, allowing it to converse with users in natural language. However, the general content it responds to is a standard answer learned through learning, and it cannot provide users with answers that are relevant to their current situation in real time. Although it is called a natural language chatbot, it lacks content that is relevant to users and responds to real-time situations. Summary of the Invention [Problem to be solved by the invention]

[0003] In addition to the above drawbacks, the services provided by current natural language chatbots are limited to general conversations and cannot meet all needs. For example, because the application does not integrate with the real world, it cannot provide effective answers that are tailored to the user's current situation. [Means for solving the problem]

[0004] To provide a new type of method for triggering intelligent interactions, the present disclosure provides a method and system for triggering intelligent interactions through a real image, the system including a server and a database, using the server to allow a user device to connect and display in the real image, obtain link images with location-based data and related information, and provide links in the real image to trigger intelligent interactions.

[0005] In a method for triggering intelligent interaction through real-world images, after receiving location information and a real-world image request sent from a user device, based on the location information, calculate multiple visibility ranges corresponding to multiple viewing angles from the received location information, and search a database based on the real-world image request and the calculated multiple visibility ranges to obtain one or more location-based data.

[0006] Then, after sending link information of one or more location-based data within each visible range to the user device, a real-image interface is launched on the user device, and one or more link images linked to one or more location-based data based on the respective spatial positions of the one or more location-based data within the acquired visible range are displayed on the real-image interface, thereby providing intelligent interaction link points.

[0007] When a user taps on the intelligent interaction link point, an intelligent interaction request is generated, an intelligent interaction process is initiated between the server and the user device, a chatbot is introduced, and at the same time, the location information and one or more location-based data within the visible range are obtained, and an intelligent interaction interface is started on the user device to interact. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram of an embodiment showing the structure of a system for triggering intelligent dialogue through real-world video. [Figure 2] FIG. 1 is a diagram of an example of a data structure for implementing a natural language model in a system for triggering intelligent dialogue through real-world video. [Figure 3] 1 is a flowchart illustrating an example method for triggering intelligent interaction via real-world video. [Figure 4] 10 is a flow chart of another embodiment illustrating a method for triggering intelligent interaction via real-world video. [Figure 5]1 is a flowchart illustrating an example of natural language information processing. [Figure 6] 10 is a flowchart illustrating another embodiment of natural language information processing. [Figure 7] FIG. 1 is a schematic diagram showing a real-image interface. [Figure 8] FIG. 1 is a schematic diagram of an embodiment showing an augmented reality interface. [Figure 9] FIG. 1 is a schematic diagram of an embodiment showing an augmented reality interface. [Figure 10] FIG. 1 illustrates an example of a graphical user interface for implementing an intelligent dialogue process. [Figure 11] FIG. 1 illustrates an example of a graphical user interface for implementing an intelligent dialogue process. [Figure 12] FIG. 1 illustrates an example of a graphical user interface for implementing an intelligent dialogue process. DETAILED DESCRIPTION OF THE INVENTION

[0009] The present disclosure provides a method and system for triggering intelligent dialogue through real-world images. By activating a camera function on a user device and activating a real-world image interface, a current environmental image can be captured and displayed. Corresponding location-based data can be obtained from the server based on location information transmitted to the server, and linked images of content can be displayed in the real-world image interface according to the spatial location associated with the location-based data. In one specific embodiment, an augmented reality image is formed in the real-world image interface by displaying a real-world image (e.g., an image of the surrounding environment) captured by a camera within a visible range and combining it with one or more linked images displayed at one or more spatial locations. This allows a user to operate the user device to view the environmental image on the display, while simultaneously viewing the linked image and text information displayed at spatial coordinates using augmented reality (AR) technology. In particular, a link point is additionally provided that the user can tap to launch an intelligent dialogue process, thereby achieving the objective of the method for triggering intelligent dialogue through real-world images.

[0010] According to an embodiment, the system for triggering intelligent dialogue through real-world video can provide social media services via a network using a cloud server, allowing users to share text content, image content, and audio-video content after signing up. Thus, location-based data obtainable from real-world video can include the text content, image content, and audio-video content shared by users. Furthermore, the cloud server provides chatbots for various fields, allowing users to engage in dialogue using services provided by the cloud server. Here, the utilization of artificial intelligence technology includes a chatbot that learns and trains data for various fields using machine learning algorithms and natural language processing (NLP) technology to provide dialogue services. The chatbot learns the user's social media activity to acquire user preferences. The chatbot provides dialogue content more suited to the user's individual needs and current environmental characteristics by combining the semantics of the user's dialogue, user preferences, and environmental information acquired in real time based on the user's dialogue semantics and preferences.

[0011] 1 is a diagram showing an embodiment of a system for triggering intelligent interaction through real images. The system mainly includes a server-side server, a database, and related software and hardware, and an application program executed by a user-side device.

[0012] The cloud server 100 shown in the figure is constructed from a computer system, a database, and a network, and includes various functional modules realized through the cooperation of software and hardware, such as a natural language processing module 101, a machine learning module 103, an external system interface module 105, and a user interface module 107. The natural language processing module 101, which processes natural language information, is realized by a chatbot with natural language processing capabilities. The machine learning module 103 not only executes machine learning algorithms to train natural language models, but also uses deep learning methods to learn users' online activities and obtain user preference information, allowing the chatbot to provide conversational content tailored to the user's preferences. The cloud server 100 also includes an external system interface module 105. The external system interface module 105 executes circuits and related application software for connecting (e.g., via the network 10) with external systems (e.g., a first external system 111 and a second external system 112) and obtaining data via an application programming interface (API). The cloud server 100 also includes a user interface module 107. The user device 150 can obtain services provided by the cloud server 100 by providing a connection between the user device 150 and the cloud server 100 through the network connection function in the user interface module 107, running a server (web server) for providing web services, and running an application program in the user device 150 for correspondingly obtaining the services.

[0013] According to an embodiment, the cloud server 100 implements an augmented reality module 109 using augmented reality (AR) technology. The augmented reality module 109 allows a user to view an image of the surrounding environment through a display when the user activates an augmented reality process in response to a real-world image request generated by the user device 150, and simultaneously view an augmented reality object integrated into the augmented reality image provided by the cloud server 100. For example, the real-world image interface 115 is started on the user device 150, and the augmented reality module 109 of the cloud server 100 operates to obtain linked images of location-based data displayed at one or more spatial locations within the visible range and acquire corresponding information. For example, when a user activates the camera of the user device 150 to obtain and display an image of the surrounding environment, the current location information of the user device 150 is also transmitted to the cloud server 100. This allows a software program on the cloud server 100 to identify the user's current location and the surrounding environmental objects (e.g., buildings, roads, landscapes, etc.) visible in the space, and retrieve location-based data related to these environmental objects by searching a database. These data are then transmitted to the user device 150 and integrated into the augmented reality image via the real image interface 115, displaying a linked image corresponding to the location-based data at the corresponding spatial coordinates.

[0014] According to the illustrated structure, the cloud server 100 includes an internal or externally connected database. This allows the cloud server 100 to provide data services. The illustrated audio / video database 110 provides audio / video content uploaded and shared by each end user, which is accessible to the user device 150 via the network 10. The content uploaded and shared by users may include text content and image content. The user database 120 stores user data, including personal information, uploaded text content, image content, and audio / video content, and acquires user activity data (e.g., network activities such as content viewed, following, liking, sharing, and subscribing) on ​​the online services provided by the cloud server 100. A user profile is created based on this data. Furthermore, as dialogue content is continuously generated over time, the user database 120 may store and update user data, including the user's past dialogue history, over time and provide it to a natural language model as a dialogue history for machine learning algorithms to learn from. The vector database 130 stores structured data in which various text contents, image contents, audio and video contents are calculated by vectorization, and may be used to collate various information that matches the user's personalization.

[0015] The database may further include a map database 140. The map database 140 allows a user to search for location-based data associated with a specific geographic location or spatial coordinates, such as a country, administrative district, tourist destination, landmark, or location tagged by a user. In particular, location-based data including spatial coordinates may be stored in addition to data on a planar geographic location. For example, a restaurant located on a certain floor of a building may be assigned a geographic location (e.g., longitude and latitude) and spatial coordinates including altitude information (e.g., a spherical coordinate system described by a radial distance (γ), a polar angle (θ), and an azimuth angle (φ), or a Cartesian coordinate system described by an X-axis, a Y-axis, and a Z-axis). Furthermore, the map database 140 can be used to search for environmental objects with altitude information. Therefore, the system can identify environmental objects that the user can see through the user device 150 and form the user's viewfinder based on the location information (e.g., altitude) of the user device 150 and the shooting direction of the user device 150 (a combination of an azimuth angle and a polar angle).

[0016] According to the schematic diagram of the system structure shown in the figure, the cloud server 100 may obtain data from external systems via a network 10 or a connection using a specific protocol. The external systems, such as a first external system 111 and a second external system 112 shown in the figure, are servers installed by, for example, a government or a company, and are used to provide open data. As a result, the cloud server 100 can obtain real-time information tailored to its needs, such as real-time weather information, real-time road conditions, real-time news, and real-time location-related network information, through the application programming interfaces provided by the external systems using the external system interface module 105.

[0017] The user device 150 executes an application that can obtain services provided by the cloud server 100. For example, for a social media service provided by the cloud server 100, the user device 150 executes a corresponding social media application to obtain the social media service via the user interface module 107. In particular, the cloud server 100 provides a natural language chatbot via the natural language processing module 101, allowing the user to interact with the chatbot using an intelligent dialogue interface initiated in the real-world video interface 115. Meanwhile, the cloud server 100 uses the machine learning module 103 to learn activity data of the user's use of various services of the cloud server 100. The activity data includes activity data obtained by the user's use of the social media application and the real-world video interface 115. In this way, the machine learning module 103 can learn the user's interests and characteristics and construct user data.

[0018] It should be noted that the cloud server 100 acquires various text content, image content, and audio / video content as unstructured data and converts them into vectorized data using an encoding method, making it easier to understand the meaning of the content and facilitating data searches. Furthermore, the vectorized data may be used to match a user's search keywords. A specific distance function is used to calculate the distance between the search keyword and the vectorized data in the database, with the closer the distance, the more relevant the data. In this way, users can use the vector database 130 to perform data searches.

[0019] According to an embodiment, the vector database 130 in the cloud server 100 can support a multi-mode search service for text content, video content, etc. Here, the vector database 130 provides structured data, for example, converts various text content, image content, and audio / video content into text, and then calculates vectorized data using a vector algorithm. The vectorized data is used in search services and also in natural language processing processes. The natural language processing processes associate the vectorized data with a vector space using a natural language model. For example, for a word entered by a user, a word vector is calculated using a vector algorithm.

[0020] According to an embodiment, in the method for triggering intelligent dialogue through real-world video according to the present disclosure, the function of conducting intelligent dialogue is implemented as a chatbot running on the cloud server 100. The chatbot can converse with a user using natural language including text and voice, and not only respond to messages input by the user but also obtain user data from the cloud server 100 in advance of the dialogue to understand the user's personality and habits. In addition, the chatbot can obtain real-time conditions (e.g., local weather and news based on the user's current location) from external systems (first external system 111, second external system 112), so that the content to which the chatbot responds not only matches the user's preferences but also reflects the real-world situation.

[0021] Furthermore, a cloud server that executes the method for triggering intelligent dialogue via real-world video may be configured to install trained chatbots for various fields. If the user needs more information during the dialogue, the chatbot can introduce a chatbot for a related field (e.g., a chatbot specialized in a business, product, or field related to a restaurant, food court, or night market). The chatbot for the related field can then continue the dialogue with the user using natural language and provide more specialized and accurate dialogue content.

[0022] The system for triggering intelligent dialogue through real-world video provides intelligent dialogue services through a real-world video interface, and uses machine learning methods to learn data about user preferences from dialogue history and the user's social media activities to form structured data in the system. As shown in Figure 2, in an example diagram showing a data structure for implementing a natural language model in the system for triggering intelligent dialogue through real-world video, the data is divided into social media platform data 21, user data (e.g., user profile) 23, and user activity data 25.

[0023] The social media platform data 21 is private data within the system, and includes viewer data (viewers) 211 of users who access various contents provided by the cloud server using the system that triggers intelligent dialogue through real-world video, content creator data (creators) 212 of users who provide various contents within the system, business data (business) 213 that allows companies to create corporate information in the system and facilitate advertising, and various location data (locations) 214 related to obtaining geographic locations so that the system can provide location-based services.

[0024] The user data 23 is public data within the system and includes information that can be edited by the user himself / herself. Here, the data that can be edited by the user himself / herself includes viewer data 231 that the system acquires from various user activity data. The viewer data 231 includes viewer interest data acquired by machine learning, such as recent interest data, history interest data, and location interest data, for example.

[0025] The content creator data 232 in the user data 23 is information related to the user as a content creator, and includes data related to the content creator's preference type and the content creator's location obtained through machine learning by the system. For example, data on the user as a content creator, the content creator's preference type obtained through learning, and location information including a geographic location or a specific location of a place may be included.

[0026] The business data 233 in the user data 23 includes the business type of the company and the characteristics of its products, which the system obtains through machine learning if the user is a company.

[0027] The user activity data 25 is private data within the system, and includes statistical data on user activities in various services provided by the cloud server and data obtained by machine learning. Specifically, it mainly includes viewer data 251, content creator data 252, and business data 253.

[0028] The viewer data 251 is activity data such as the view rate, view time, and following, liking, commenting, and subscribing when a user uses services provided by the cloud server. The content creator data 252 is statistical data when a user acts as a content creator, such as the number of followers of a channel or account, the number of views of created content, and the view rate of the account. The business data 253 is data such as the number of followers, the number of views of content, and overall impression data when the user is a business.

[0029] The above-mentioned social media platform data 21, user data 23, and user activity data 25 collected and learned by the cloud server 100 form the basis for the dialogue service realized by utilizing the natural language processing and generative artificial intelligence technology of the present disclosure. Within the cloud server 100, the above-mentioned various data are calculated by a processing circuit, thereby enabling the provision of a chatbot that is personalized and meets the real-time needs of users.

[0030] According to an embodiment, the natural language model executed in the cloud server 100 first executes a vector algorithm on the content, user preferences, and real-time environmental information input by the user through the dialogue interface, tags the acquired text, calculates the vector of each word, and searches a database to obtain relevant content based on the vector distance between words, thereby generating dialogue content suited to the user preferences and real-time environmental information. During the online dialogue process, a transformer model may be used to perform processes such as machine translation, document summarization, and document generation on the text data. Then, the semantics of the user during the dialogue may be acquired so that the chatbot can generate dialogue content.

[0031] 3 is a flowchart illustrating an embodiment of a method for triggering intelligent interaction through a real image. The method for triggering intelligent interaction through a real image can be executed in a server such as the cloud server illustrated in FIG. 1 and can provide augmented reality and intelligent interaction services over a network. This allows a user device having a corresponding application program installed thereon to trigger intelligent interaction when executing an augmented reality image.

[0032] When an application program running on a user device starts a real-world image interface, it generates a real-world image request to a server. The server receives location information and the real-world image request from the user device (step S301), and then uses the location information to calculate multiple viewable areas corresponding to multiple viewing angles based on the location information (step S303). When the user device transmits its current location information, the software program on the server calculates multiple viewable areas that can be seen from different viewing angles of the user's location, and then searches a database (the map database shown in FIG. 1) based on the real-world image request and the calculated multiple viewable areas to obtain environmental objects (e.g., buildings, landmarks, tourist attractions) that can be seen in the image captured by the user device's camera. By searching for location-based data within the viewable areas, the software program identifies the spatial location of linked images that can display the location-based data, as well as associated text descriptions (step S305).

[0033] Then, search results such as one or more location-based data within the visible range obtained from different viewing angles are sent to the user device (step S307), and the real-life image interface launched on the user device displays one or more link images corresponding to each of the one or more location-based data within the visible range at corresponding spatial positions, and further provides intelligent interactive link points.

[0034] It should be noted that when a user device starts the real-image interface, the image displayed is a real-image of the surroundings captured by the user device's camera. At this time, a software program running on the server or the user device calculates the visible range based on the current location, i.e., the user's location can reflect altitude (e.g., on top of a building, a mountain, etc.), and the user device's shooting direction can reflect the visual angle. At this time, the visible range is displayed on the screen of the user device, and at the same time, location information and a real-image request are generated and sent to the server. The server identifies environmental objects visible within the visible range, and combines the real-image captured by the user device's camera within the visible range with one or more linked images displayed at one or more spatial locations to form an augmented reality image.

[0035] When the user taps the intelligent interaction link point displayed on the real-life image interface, the intelligent interaction interface is started, and at the same time, the server receives an intelligent interaction request generated by the user device (step S309), and starts an intelligent interaction process between the user device and the server (step S311). In this way, in the intelligent interaction process, a chatbot is introduced, and at the same time, location information and one or more location-based data within the visible range are obtained, and the user interacts with the chatbot in the intelligent interaction interface.

[0036] FIG. 4 is a flow chart illustrating another display mode embodiment of the method for triggering intelligent interaction via real-world video.

[0037] According to the illustrated flowchart embodiment, a user initiates a graphical user interface by operating a software program running on a user device. Figure 7 shows a schematic diagram of an embodiment for viewing location-based data in a map interface 70, which shows a user interface against the background of an electronic map, through which location-based data within different geographical areas can be viewed. The figure shows audio-visual link points 701, 702, and 703 displayed at different locations, and the bottom of the interface further displays several functions provided by the software program, such as playback 711, interaction 712, assistant 713, search 714, and return to user homepage 715.

[0038] In the flow shown in FIG. 4, the server receives from the user device a signal indicating a point of interest (POI) tapped by the user (step S401). The server then searches a database based on the POI tapped by the user to obtain information about the point of interest, including, for example, audio-video content, map data, and conversation groups (step S403). At the same time, when the user taps a point of interest such as audio-video link points 701, 702, or 703 shown in FIG. 7, a browse page is launched that plays text data, image data, or audio-video data related to the point of interest, and location-based data may be displayed on the browse page. When the user taps augmented reality link point 705 (AR), the user device activates the camera to capture an image of the surrounding area, activating the real-image mode and launching the real-image interface. In the real-image mode, based on the search results obtained from the database search by the server, a link image of location-based data related to the position of the point of interest may be displayed at one or more spatial locations in the real-image interface (step S405).

[0039] When the user device activates the real-image mode, an augmented reality interface 80, shown in FIG. 8, is displayed, displaying an image of the surroundings captured by the user device's camera. In this augmented reality interface 80, for example, the image link points 801 and 802 shown are corresponding link images displayed based on the spatial position of each piece of location-based data, and may include a text description. This link image establishes an association between the location-based data and one of the environmental objects within the visible range. In this example, three link points are displayed at the bottom of the augmented reality interface 80, each displayed as a thumbnail. For example, audio-visual link points 801 and 802 and an interactive link point 803 are shown.

[0040] When the user taps the interaction link point 803 in the augmented reality interface 80, an intelligent interaction process is initiated. At this time, the server receives an intelligent interaction request from the user device (step S407) and initiates an intelligent interaction process between the user device and the server (step S409). At this time, a chatbot is introduced into the intelligent interaction interface initiated by the user device, and simultaneously obtains the location information of the user device and one or more location-based data within the visible range, and starts the intelligent interaction interface on the user device to enable interaction.

[0041] Continuing with reference to the augmented reality interface 90 shown in FIG. 9, in this example, the visible area captured by the camera of the user device includes a building that occupies a large portion of the screen. As a result of calculations performed by a software program on the server, it is determined that this environmental object is included in the visible area formed from a certain viewing angle at the user's current location. Therefore, the link image displayed by the augmented reality interface 90, limited by the environmental object, includes an audio-visual link point 901 displayed for this building. In addition, an interaction link point 902 is also provided, including prompt text 903 for providing intelligent interaction services.

[0042] According to an embodiment, in an intelligent dialogue process, the example flowchart of natural language information processing shown in FIG. 5 is executed by a chatbot.

[0043] When a user taps an interaction link point (interaction link point 803 shown in FIG. 8 ) in the real-world image interface, the cloud server receives the selection of the intelligent interaction (step S501) and launches the intelligent interaction process (step S503). Next, the intelligent interaction interface is launched, and the user can input text content, image content, or specific audio-video content (e.g., input a link to share the audio-video content) through the intelligent interaction interface. The cloud server receives the content input by the user through the user interface module (step S505). According to an embodiment, the intelligent interaction process is implemented as a chatbot using a natural language model, interacts with the user through the intelligent interaction interface, and executes a natural language information processing method for each content input by the user. The intelligent interaction interface provides an input field for the user to input content, and displays an interaction display area showing the interaction content output by the chatbot and the content input by the user.

[0044] At this time, the cloud server acquires content input by the user using the user interface module. The content received through the dialogue interface may be text content, audio content, or video content. If audio content or video content is received, it is converted into text using a text conversion process, and then semantic analysis is performed to acquire semantic features (step S507). During the above process, the cloud server acquires user data from a user database and acquires real-time environmental information from an external system (using the external system interface module 105 shown in FIG. 1) (step S509).

[0045] Then, the semantic features matching the content input by the user, the user preferences obtained from the user data, and the content of real-time environmental information can be determined (or searched and selected from a database) (step S511), and the content is processed by a natural language model executed in the intelligent dialogue process to generate dialogue content (step S513).The dialogue content is then introduced into the intelligent dialogue process, and the dialogue content is output to the dialogue interface (step S515).The flow may repeat the above steps S505 to S513.

[0046] Furthermore, when the natural language model in the cloud server operates, multidimensional information is recorded using a database or system memory. The multidimensional information may include past dialogue history in the same intellectual dialogue process. This allows the chatbot to consider the user's past dialogue history in the current intellectual dialogue process (step S517) in addition to considering the user's semantic features, user preferences, and real-time environmental information in the dialogue before generating a dialogue (step S511). In this way, the dialogue content generated by the natural language model (step S515) is dialogue content that matches the current situation.

[0047] Reference is now made to a flowchart of another embodiment of natural language information processing shown in FIG.

[0048] 6, a user initiates an intelligent dialogue process using an application program (step S601). Through dialogue with the chatbot, the system receives the dialogue content input by the user (step S603) and further obtains the semantic features of the user. According to an embodiment, a natural language processing module is utilized in the cloud server to perform transformer operations and vector operations to obtain the semantic features (step S605).

[0049] It should be noted that natural language information processing can utilize artificial intelligence technology to learn natural language, and after natural language understanding, perform text classification and grammar analysis. When processing user-entered dialogue content, a deep learning method based on a transformer model (proposed by the Google® Brain team in 2017) can be used to process the natural language content entered by the user in chronological order. If the input content is not text, it must be converted into text to obtain the text. In this way, during the online dialogue process, this transformer model can be used to perform machine translation, document summarization, document generation, and other tasks.

[0050] After obtaining the semantic features of the user's interactive content, the system obtains real-time environmental information from an external system in real time based on the location by analyzing the obtained user preferences and the user's current location or the location information of the user's interest from the interactive content (step S607). Here, the real-time environmental information may include one or any combination of real-time weather, real-time road conditions, real-time news, and real-time location-related network information (e.g., POIs on a map, POI ratings, etc.) obtained from one or more external systems in real time.

[0051] The system then utilizes the vector database to calculate the closest answer based on the user's semantic features, user preferences, and real-time environmental information, or by combining past dialogue history (step S609). It should be noted that the data in the vector database is structured information obtained by a vector algorithm. This allows the system to extract semantically similar words from the retrieved content based on vector distance. For example, the vector distance between the word "computer" appearing in the dialogue content and the word "calculating" in the database is relatively close, while the vector distance between the word "computer" and the word "running" is relatively far.

[0052] In this embodiment, a vector algorithm is run on content entered by a user, content of the user's interest, and real-time environmental information, optionally adding past dialogue history, to tag the acquired text and calculate vectors for each word. Relevant content is then acquired based on the vector distance between words, and dialogue content suited to the user's preferences and real-time environmental information is generated based on this. Furthermore, according to this embodiment, running the vector algorithm on the past dialogue history recorded in the cloud server 100 may generate dialogue content suited to the user's current emotions, for example, by continuing the same topic covered in the past dialogue history and using terms that match the analyzed emotions.

[0053] Furthermore, based on the above information, the system searches an audio-visual database to obtain appropriate audio-visual content, and adds the user's current location and environmental objects and location-based content displayed by the real-world video interface (step S611).The chatbot then generates dialogue content using natural language processing and generative artificial intelligence techniques (step S613), and the generated dialogue content is output to the dialogue interface (step S615).In some embodiments, the system continuously performs the above steps during a chat, allowing the chatbot to interact with the user using natural language (text or voice) and provide real-time content (audio-visual content, text content) that the user is interested in.

[0054] When entering into the intelligent dialogue process, you may refer to the dialogue interface 1000 shown in Figure 10, the dialogue interface 1100 shown in Figure 11, and the dialogue interface 1200 shown in Figure 12 as related examples. In the dialogue interfaces shown in each example, an input field is provided for the user to input content, and an dialogue display area is also provided for the dialogue content output by the chatbot and the user input content.

[0055] For a related example, see Fig. 10. The dialogue interface 1000 shown in Fig. 10 displays several dialogue contents 1001, 1002, and 1003 between a user and a chatbot. The chatbot may also search a database based on the semantic features of the user obtained from the dialogue content 1002, and provide recommended audio-visual content 1004. An input field 1005 is provided at the bottom of the dialogue interface 1000, allowing the user to further input dialogue content.

[0056] Another mode is the dialogue interface 1100 shown in Fig. 11. In this example, at the beginning of the online dialogue process, the system directly provides natural language dialogue content 1101, 1102, 1104 based on user preferences and real-time information, and directly suggests audiovisual content 1103. In this way, the user can continue to use the input field 1105 in the dialogue interface 1100 to respond to the dialogue content.

[0057] According to another embodiment, the dialogue content generated by the natural language model running in the chatbot based on the semantic features of the content input by the user, the user preferences and real-time environmental information may include multiple recommended options, multiple recommended audio-visual contents and / or multiple friend recommendation links, see the example shown in FIG. 12.

[0058] In the online interaction process, the interaction interface 1200 shown in FIG. 12 includes interaction content 1201 generated by the chatbot based on the semantic features of the user. In this example, the semantics allows the chatbot to determine that the user is trying to select a specific item, and accordingly provides several recommended options 1202. It is particularly noteworthy that the chatbot provides the recommended options to the user based on real-time environmental information acquired by the system from external systems. For example, the chatbot may provide the recommended options 1202 based on real-time weather information, road conditions, time, and the user's current location. Here, if it is time for a meal, the chatbot may provide dining options based on the user's eating habits and restaurants open near the user's current location.

[0059] On the other hand, if the user indicates that he / she wants to watch audio / video content, the recommendation options 1202 may be a plurality of recommended audio / video content, or if the user indicates that he / she wants to find friends with common interests, the recommendation options 1202 may be a plurality of friend recommendation links.

[0060] Furthermore, the user uses the input field 1206 to input dialogue content 1203 according to these recommended options 1202. The chatbot then responds with dialogue content 1204 based on the semantics of the dialogue content 1203, and suggests multiple recommended contents 1205 based on the semantics of the above-mentioned dialogue content. Continuing with the above example, if the user answers that they would like to eat one of the dishes, the chatbot will suggest restaurant options corresponding to the dish the user wants to eat based on real-time weather information, road conditions, and the user's current location obtained by the system from an external system, and if the weather is bad and the roads are congested, the chatbot will accordingly recommend restaurants that the user can reach easily.

Claims

1. It runs on the server, receiving location information and a reality image request transmitted from a user device; calculating a plurality of visible ranges corresponding to a plurality of viewing angles based on the position information using the position information; searching a database based on the real-world image request and the calculated plurality of visibility ranges to obtain one or more location-based data; Sending link information of the one or more location-based data within each visible range to the user device, a real-image interface is started on the user device, and one or more link images linked to the one or more location-based data are displayed on the real-image interface based on the spatial positions of the one or more location-based data within the visible range, thereby providing intelligent interaction link points; receiving an intelligent interaction request from the user device generated by triggering the intelligent interaction link point displayed on the real-world video interface, and initiating an intelligent interaction process; In the intelligent interaction process, a chatbot is introduced, and simultaneously the position information and the one or more location-based data within the visible range are obtained, and an intelligent interaction interface is started on the user device to interact; Including, A method for triggering intelligent dialogue via a real video, comprising:

2. In the real image interface, a real image of the visible range captured by a camera of the user device is displayed, and an augmented reality image is formed by combining the real image of the visible range captured by a camera of the user device with the one or more linked images displayed at one or more spatial positions. The method for triggering intelligent interaction through real-world video according to claim 1.

3. In the real-image interface, the link image and a text description corresponding to a spatial position of each location-based data are displayed, and there is an association between the location-based data linked to the link image and one of the environmental objects within the visible range. The method for triggering intelligent interaction through real-world video according to claim 2.

4. In the intelligent interaction process, the chatbot generates location-based interaction content related to the location of the user device based on the one or more location-based data and the recognized one or more environmental objects within the visible range. The method for triggering intelligent interaction through real-world video according to claim 3.

5. In the intelligent dialogue process, the chatbot: receiving content input by a user through the intelligent dialogue interface; obtaining semantic features of the user-input content; acquiring user data and real-time environmental information; Executing a natural language model to generate dialogue content based on the semantic features of the content input by the user, user preferences obtained from the user data, and content of the real-time environmental information, and with reference to the one or more location-based data within each of the visibility ranges; introducing the dialogue content into the intelligent dialogue process and outputting the dialogue content to the intelligent dialogue interface; is executed, The content received through the intelligent dialogue interface is text content, audio content or video content, and when the audio content or video content is received, it is converted into text through a text conversion process, and further semantic features of the text are obtained through semantic analysis; the natural language model executed on the server uses a conversion model to perform machine translation, document summarization, and document generation processes to generate the dialogue content; A method for triggering intelligent dialogue via real images according to any one of claims 1 to 4.

6. a server on which a database is provided; The server receiving location information and a reality image request transmitted from a user device; calculating a plurality of visible ranges corresponding to a plurality of viewing angles based on the position information using the position information; searching the database based on the real-world image request and the calculated plurality of visibility ranges to obtain one or more location-based data; Sending link information of the one or more location-based data within each visible range to the user device, a real-image interface is started on the user device, and one or more link images linked to the one or more location-based data are displayed on the real-image interface based on the spatial positions of the one or more location-based data within the visible range, thereby providing intelligent interaction link points; receiving an intelligent interaction request from the user device generated by triggering the intelligent interaction link point displayed on the real-world video interface, and initiating an intelligent interaction process; In the intelligent interaction process, a chatbot is introduced, and simultaneously the position information and the one or more location-based data within the visible range are obtained, and an intelligent interaction interface is started on the user device to interact; Executing a method for triggering intelligent dialogue through a real-life image including: A system that triggers intelligent dialogue through real-world video.

7. In the real image interface, the real image of the visible range captured by the camera of the user device is displayed, and an augmented reality image is formed by combining the real image of the visible range captured by the camera of the user device with the one or more linked images displayed at one or more spatial positions; the link image and a text description corresponding to the spatial location of each location-based data are displayed, and there is an association between the location-based data linked to the link image and one of the environmental objects within the visible range; The system for triggering intelligent dialogue through real images according to claim 6.

8. searching the map data and audio / video data in the database to obtain one or more environmental objects within the visible range and the location-based data corresponding to each environmental object, thereby accurately displaying the link image and explanatory text at the spatial position corresponding to each location-based data; The system for triggering intelligent dialogue through real images according to claim 7.

9. In the intelligent interaction process, the chatbot generates location-based interaction content related to the location of the user device based on the one or more location-based data within the visible range and the recognized one or more environmental objects. The system for triggering intelligent dialogue through real images according to claim 8.

10. The server provides an external system interface for connecting with one or more external systems to obtain real-time environmental information, and the chatbot further generates the location-based interactive content based on the real-time environmental information. The system for triggering intelligent dialogue through real images according to claim 9.

11. In the intelligent dialogue process, the chatbot: receiving content input by a user through the intelligent dialogue interface; obtaining semantic features of the user-input content; acquiring user data and real-time environmental information; Executing a natural language model to generate dialogue content based on the semantic features of the content input by the user, user preferences obtained from the user data, and content of the real-time environmental information, and with reference to the one or more location-based data within each of the visibility ranges; introducing the dialogue content into the intelligent dialogue process and outputting the dialogue content in the intelligent dialogue interface; is executed, A system for triggering intelligent dialogue via a real image according to any one of claims 6 to 10.

12. The content received through the intelligent dialogue interface is text content, audio content or video content, and when audio content or video content is received, it is converted into text through a text conversion process, and then semantic features of the text are obtained through semantic analysis; The system for triggering intelligent dialogue through real images according to claim 11.

13. the natural language model executed on the server uses a conversion model to perform machine translation, document summarization, and document generation processes to generate the dialogue content; The system for triggering intelligent dialogue through real images according to claim 11.

14. In the server, further executing a vector algorithm on the content input by the user, the user preferences, the real-time environmental information, and the one or more location-based data within each visible range, tagging the acquired text, calculating the vector of each word, and acquiring relevant content based on the vector distance between words, thereby generating the interactive content suitable for the user preferences and the real-time environmental information. The system for triggering intelligent interaction through real images according to claim 13.

15. The dialogue content generated by the natural language model executed in the chatbot based on the semantic features of the content input by the user, the user preferences, the real-time environmental information, and the one or more location-based data within each visible range includes a plurality of recommended options, a plurality of recommended audio-visual contents, and / or a plurality of friend recommendation links; The system for triggering intelligent dialogue through real images according to claim 11.

Citation Information

Patent Citations

  • Augmented reality (AR) imprinting method and system

    JP2022507502A

  • JPP7082219B

  • Shopping system using virtual reality technology capable of improving the real feeling of consumer shopping and increasing consumption pleasure

    TW202026996A

  • Method and system of an augmented / virtual reality platform

    US20210209676A1

  • System and method for location relevant augmented reality based communication

    WO2021064747A1