Method and system for triggering intelligent dialogue through reality image
Through real-life image triggering intelligent dialogue system, location information and real-life image calculate location data, link images, and start chat robot conversations, solving the problem that existing chat robots cannot respond in real time, and achieving dialogue content generation that is closer to user needs and environment.
Patent Information
- Application Number
- CN202410097798.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-29
AI Technical Summary
Existing natural language chatbots cannot provide answers related to the user's current status in real time, and lack the combination with the real environment to cope with all needs.
Through real-life image triggering intelligent dialogue methods and systems, the server uses the server to receive the location information of the user device and real-life image requests, calculate the placement data within the visual range, and mark the link image in the real-life image interface and provide intelligent dialogue link points, and start the chat robot for dialogue.
It realizes the generation of dialogue content that meets the user's personal needs and current environmental characteristics based on user location and environmental information, and improves the real-time response ability and environmental adaptability of the chat robot.
Smart Images

Figure CN120386834A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for enabling intelligent conversation in a browsing interface, particularly a method and system for triggering intelligent conversation with real-world images, which provides an intelligent conversation service when a user uses real-world images to browse the surrounding environment. Background Art
[0002] In the current rapidly developing artificial intelligence (AI) in various fields, one type is a natural language chatbot that can process natural language and automatically generate content, such as the Chat Generative Pre-trained Transformer developed by OpenAI, known as ChatGPT in English. Such natural language chatbots use generative artificial intelligence technology. After training on a large amount of data, they can generate new data associated with the original data and build an intelligent model through deep learning (such as generative adversarial networks (GANs)).
[0003] Taking ChatGPT as an example, ChatGPT is trained by learning a large amount of network information and can communicate with users in natural language. However, the common content in response to users is the standard answers obtained through learning and cannot respond in real time to provide answers related to the user's current state. Although it is a natural language chatbot, it lacks content related to the user and in line with the real situation.
[0004] Furthermore, in addition to the above defects, the services provided by current natural language chatbots are only for general discussions and cannot meet all needs. For example, because they are not combined with the real environment in application, they cannot provide effective responses according to the user's current situation. Summary of the Invention
[0005] To provide a novel way to trigger intelligent conversation, the present disclosure proposes a method and system for triggering intelligent conversation with real-world images. The system includes a server and a database. The server enables a user device to connect and obtain a linked image with location-based data and related information displayed in the real-world image, and provides a link for triggering intelligent conversation in the real-world image.
[0006] In the method for triggering intelligent conversation with real-world images, after receiving the location information and real-world image request transmitted by the user device, the system calculates multiple visible ranges corresponding to the location information from multiple perspectives based on the location information, and then queries the database according to the real-world image request and the calculated multiple visible ranges to obtain one or more pieces of location-based data.
[0007] After transmitting the link information of one or more pieces of location-based data within the visible range to the user device, start an augmented reality image interface on the user device, mark one or more link images linking the one or more pieces of location-based data at the respective spatial positions of the one or more pieces of location-based data obtained within the visible range in the augmented reality image interface, and provide an intelligent dialogue connection point.
[0008] When the user touches the intelligent dialogue connection point, an intelligent dialogue request is generated, starting an intelligent dialogue program between the server and the user device, importing a chatbot, and at the same time obtaining location information and one or more pieces of location-based data within the visible range, and starting an intelligent dialogue interface on the user device to conduct a dialogue.
[0009] Furthermore, the augmented reality image interface displays the real-world image of the visible range captured by the user device's camera. After combining with one or more link images marked at one or more spatial positions, an augmented reality image is formed.
[0010] Furthermore, in the augmented reality image interface, the spatial positions of each piece of location-based data can mark the corresponding link images and text descriptions, and there is a correlation between the location-based data linked by each link image and one of the environmental objects within the visible range. Among them, by querying the map data and audio-visual data in the database, one or more environmental objects within the visible range and the corresponding location-based data for each environmental object are obtained. In the augmented reality image interface, the link images and text descriptions can be accurately marked at the spatial positions corresponding to each piece of location-based data.
[0011] Furthermore, in the intelligent dialogue program, the chatbot generates location-based dialogue content related to the location of the user device based on one or more pieces of location-based data within the visible range and one or more identified environmental objects.
[0012] According to the embodiment, in the intelligent dialogue program, it is executed by the chatbot, including receiving the content input by the user through the intelligent dialogue interface, obtaining the semantic features of the content input by the user, and obtaining user data and real-time environmental information. According to the semantic features that match the content input by the user, the user preferences obtained from the user data, and the content of the real-time environmental information, and also referring to one or more pieces of location-based data within each visible range, a natural language model is run to generate a dialogue content. After importing the dialogue content into the intelligent dialogue program, the dialogue content can be output at the intelligent dialogue interface.
[0013] Furthermore, the content received through the intelligent dialogue interface is text, voice, or audio-visual content. If the received content is voice or audio-visual content, it is converted into text through a text conversion program, and then the semantic features of the text are obtained through semantic analysis.
[0014] In the server, further perform a vector algorithm on the content input by the user, the user's preferences, the real-time environment information, and the one or more pieces of location-based data within each visible range, annotate the obtained text, calculate the vectors of each word, obtain relevant content based on the vector distances between words, and thereby generate the conversation content that conforms to the user's preferences and the real-time environment information.
[0015] Furthermore, the conversation content generated by the natural language model running in the chatbot based on the semantic features of the content input by the user, the user's preferences, the real-time environment information, and the one or more pieces of location-based data within each visible range includes providing multiple recommended options, multiple recommended audio-visual contents, and / or multiple recommended friend links.
[0016] According to the embodiment, the server provides an external system interface for connecting to one or more external systems to obtain real-time environment information, enabling the chatbot to further generate location-based conversation content according to the real-time environment information.
[0017] To enable a further understanding of the features and technical content of the present invention, please refer to the following detailed description and diagrams of the present invention. However, the provided diagrams are only for reference and illustration and are not used to limit the present invention. Description of the Drawings
[0018] Figure 1 Showing an embodiment diagram of the intelligent conversation system architecture triggered by real-world images;
[0019] Figure 2 Showing an embodiment diagram of the data structure of the natural language model running in the intelligent conversation system triggered by real-world images;
[0020] Figure 3 Showing an embodiment flowchart of the intelligent conversation method triggered by real-world images;
[0021] Figure 4 Showing an embodiment flowchart of the intelligent conversation method triggered by real-world images in another scenario;
[0022] Figure 5 Showing one of the embodiment flowcharts of natural language information processing;
[0023] Figure 6 Showing another embodiment flowchart of natural language information processing;
[0024] Figure 7 Showing a schematic diagram of the real-world image interface;
[0025] Figure 8 Showing a schematic diagram of an augmented reality interface embodiment;
[0026] Figure 9Schematic diagram showing an embodiment of an augmented reality interface;
[0027] Figures 10 to 12 Example diagram of a graphical user interface showing the running of an intelligent dialogue process. Detailed implementation manners
[0028] The following are specific embodiments to illustrate the implementation manners of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments. Various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the concept of the present invention. Additionally, the drawings of the present invention are only simple schematic illustrations and are not drawn according to actual sizes. This is stated in advance. The following implementation manners will further detail the related technical content of the present invention, but the disclosed content is not intended to limit the protection scope of the present invention.
[0029] It should be understood that although terms such as "first", "second", "third", etc. may be used in this article to describe various components or signals, these components or signals should not be limited by these terms. These terms are mainly used to distinguish one component from another, or one signal from another. Additionally, the term "or" used in this article should, depending on the actual situation, possibly include any one or a combination of multiple of the associated listed items.
[0030] The present disclosure proposes a method and system for triggering an intelligent dialogue with real - world images. In this method, the user device is used to turn on the camera function and start the augmented reality interface, which can capture and display the current environmental image. According to the location information transmitted to the server, corresponding location - based data can be obtained from the server, and link images of content are marked at the spatial positions associated with the location - based data on the augmented reality interface. In one way, in the augmented reality interface, a real - world image within a visible range captured by the camera, such as the surrounding environmental image, is displayed. After combining with one or more link images marked at one or more spatial positions, an augmented reality image is formed. In this way, when the user operates the user device and views the environmental image on the display, they can also see, through augmented reality (AR) technology, the link images marked at the spatial coordinates and some text information. In particular, a link point is added for the user to touch - press to start the intelligent dialogue program, achieving the purpose of triggering the intelligent dialogue method with real - world images.
[0031] According to an embodiment, the intelligent dialogue system triggered by real - world images can provide social media services through a cloud server via a network, allowing users to join and share text, pictures, and audio - visual content. Thus, the location - specific information that can be obtained in real - world images can include text, pictures, and audio - visual content shared by users. Further, the cloud server provides chatbots in various fields, enabling users to have conversations through the services provided by the cloud server. Among them, artificial intelligence technologies are used, including chatbots trained by learning data in various fields using machine learning algorithms and natural language processing (NLP) technologies to provide dialogue services. By learning the activities of users in social media, user preferences are obtained, so that the chatbots can provide dialogue content that better meets the personal needs of users and the characteristics of the current environment based on the semantics of user conversations, user preferences, and real - time environmental information obtained.
[0032] Figure 1 Show an embodiment diagram of the architecture of the intelligent dialogue system triggered by real - world images. The system architecture shown in the figure mainly includes the intelligent dialogue system triggered by real - world images implemented by a server, a database, and related software and hardware on the server side, and an application program executed by a user - side device.
[0033] The figure shows a cloud server 100 implemented by a computer system, a database, and a network. Various functional modules are implemented through the cooperation of software and hardware. As shown, a natural language processing module 101 for processing natural language information, and the natural language processing module 101 implements a chatbot with natural language processing capabilities; a machine learning module 103 runs machine learning algorithms. In addition to training natural language models, it can also learn information about user preferences by deep learning the activities of users on the network, enabling the chatbot to provide dialogue content that meets user preferences; the cloud server 100 provides an external system interface module 105, which runs circuits and related application software for connecting to external systems (such as external system 111 and external system 112) (such as via network 10) and obtaining data through an application programming interface (API); the cloud server 100 provides a user interface module 107. Through the network connection function in the user interface module 107, the user device 150 is connected to the cloud server 100, and a web server that provides network services can be run, enabling the application program executed in the user device 150 to obtain the services provided by the cloud server 100.
[0034] According to an embodiment, the cloud server 100 implements an augmented reality (AR) module 109 using augmented reality technology. It can respond to the virtual reality image requests generated by the user device 150, enabling the user to view the nearby environment images through the display when starting the augmented reality program, and at the same time see the objects in the augmented reality provided by the cloud server 100 combined in the augmented reality images. For example, when a virtual reality image interface 115 is initiated in the user device 150 and operates through the augmented reality module 109 of the cloud server 100, a linked image of the location-based data marked at one or more spatial positions within the visible range is obtained, along with the corresponding information. For instance, when the user starts the camera of the user device 150 to obtain and display the surrounding environment images, the location information of the user device 150 at that moment is also transmitted to the cloud server 100, enabling the software program in the cloud server 100 to determine the user's location and the surrounding environment objects visible in the space, such as buildings, roads, and scenery. By querying the database, the location-based data associated with these environmental objects can be obtained, and then the user device 150 is provided with the ability to combine the linked images corresponding to each location-based data in the augmented reality images and display them at the corresponding spatial coordinates.
[0035] According to the architecture shown in the figure, the cloud server 100 is equipped with a built-in or externally connected database. Through the cloud server 100, data services are provided. As shown in the figure, the audio-visual database 110 provides the user device 150 with access to the audio-visual content uploaded and shared by various users stored therein via the network 10, which can cover text and patterns. The user database 120 stores user data, including the user's personal information, uploaded text, patterns, and audio-visual content, as well as the activity data obtained by the user in the network services provided by the cloud server 100, such as network activities like browsing content, tracking, liking, sharing, and subscribing. Based on this, a user profile can be formed. Moreover, as the conversation content continues to be generated over time, the user database 120 can store and update the user data according to the time dimension, including recording the user's historical conversation records, which become the conversation records for the machine learning algorithm in the natural language model to learn from. The vector database 130 records the structured information of various texts, patterns, and audio-visual content after vectorization calculations, which can be used to compare and find various data that conform to the user's personalization.
[0036] The database may also include a map database 140, which is used to provide users with queries for context - aware data associated with specific geographical locations or spatial coordinates. For example, it is associated with countries, administrative regions, scenic spots, landmarks, and user - marked locations at specific geographical locations. In particular, in addition to data on planar geographical locations, context - aware data with spatial coordinates can also be recorded. For example, if a restaurant is located on a certain floor of a building, the restaurant will be assigned a geographical location (such as latitude and longitude) and spatial coordinates with height information, such as spherical coordinates described by radial distance (γ), polar angle (θ), and azimuth angle (φ), or Cartesian coordinates described by the X - axis, Y - axis, and Z - axis. Further, since the map database 140 can query environmental objects with height information, the system can determine the environmental objects that the user can see through the user device 150 based on the height information (such as altitude) of the location where the user device 150 is located and the shooting direction (a combination of azimuth angle and polar angle) of the user device 150, forming the user's visible range.
[0037] According to the system architecture schematic diagram shown in the figure, the cloud server 100 can also obtain data from external systems through the network 10 or connections under specific protocols, such as the external system one 111 and the external system two 112 shown schematically in the figure, which are, for example, servers established by governments or enterprises to provide open data, so that the cloud server 100 can obtain data that meets requirements and real - time information, such as real - time weather, real - time traffic conditions, real - time news, and real - time location - related network information, through the external system interface module 105 using the application programming interfaces provided by the respective external systems.
[0038] The user device 150 executes an application program that can obtain the services provided by the cloud server 100. For example, if the cloud server 100 provides a social media service, the user device 150 executes the corresponding social media application program to obtain the social media service through the user interface module 107. And in particular, the cloud server 100 provides a natural - language chatbot through the natural - language processing module 101, enabling users to have a conversation with the chatbot through the intelligent dialogue interface initiated in the augmented - reality video interface 115. On the other hand, the cloud server 100 can learn the activity data of users using various services of the cloud server 100 through the machine - learning module 103, including obtaining the activity data of users using the social media application program and the augmented - reality video interface 115, so that the machine - learning module 103 can learn the user's interest characteristics and establish user data.
[0039] It should be noted here that various text, pattern, and audio-visual contents obtained by the cloud server 100 are unstructured information, which can be converted into vectorized data through encoding to facilitate obtaining the meaning therein and is conducive to data search. Further, the vectorized data can be used to compare with the user's search keywords, and a distance function is used to calculate the distance between the search keywords and the vectorized data in the database. The closer the distance, the more similar the data, enabling the user to search for data through the vector database 130.
[0040] According to the embodiment, the vector database 130 of the cloud server 100 can support multi-modal search services such as text and video. Structured information is provided therein. For example, after various text, pattern, and audio-visual contents are texturized, vectorized data is calculated through vector algorithms, which can be applied to search services and can also be used in natural language processing programs. The natural language processing program will use a natural language model to map the vectorized data to a vector space. Taking the words input by the user as an example, word vectors are obtained after vector calculation.
[0041] According to the embodiment, in the method for triggering an intelligent conversation with a real-world image proposed in the present disclosure, the function of conducting the intelligent conversation is implemented as a chatbot running on the cloud server 100. The chatbot can converse with the user in natural language, including text and voice. In addition to responding to the information input by the user, it can also obtain the user's data in advance through the cloud server before the conversation, from which the user's personality and habits can be obtained. Additionally, it can also learn the real-time status from external systems (111, 112), such as obtaining the local weather and news according to the user's location, so that the replied content can not only target the user's preferences but also reflect the real situation.
[0042] For example, when the user mentions their needs during the conversation with the chatbot proposed in the present disclosure, such as stating that it is already meal time, this chatbot will provide suggestions on meals and restaurants based on the meals preferred by the user learned in the past, the environmental objects in the real-world image captured by the user's device at present, as well as the user's current location or the location information mentioned by the user in the conversation, rather than simply obtaining answers from the already learned database.
[0043] Further, various domain-trained chatbots can be set in the cloud server running the method for triggering an intelligent conversation with a real-world image. When the user indicates the need for further information during the conversation, the chatbot can introduce chatbots in related domains (such as business / product / domain robots of a certain restaurant, food street, night market), enabling the chatbots in related domains to continue to converse with the user in natural language and provide more professional and accurate conversation content.
[0044] The intelligent dialogue system triggered by real - world images provides intelligent dialogue services through the real - world image interface, uses machine learning methods to learn user preference data from the conversation history and the user's activities on social media, and forms structured data in the system. For reference Figure 2 Shown is an embodiment diagram of the data structure for running a natural language model in the intelligent dialogue system triggered by real - world images, where the data is divided into social media platform data 21, user data (such as user profile) 23, and user activity data 25.
[0045] The social media platform data 21 belongs to the non - public data in the system. The intelligent dialogue system triggered by real - world images obtains viewer data 211 of users accessing various contents provided by the cloud server, and creator data 212 of users providing various contents in the system. It also includes business data 213 for the system to provide enterprises to establish enterprise data for advertising placement. And because the system can provide location - based services, it will obtain various location - related data 214.
[0046] The user data 23 belongs to the publicly available data in the system, covering data edited by the user himself. It includes viewer data 231 obtained by the system from various user activity data, which may include user preference data (interest data) as a viewer obtained through machine learning, such as recent preference data, historical preference data, and location - related preference data.
[0047] The creator data 232 in the user data 23 is related data of the user as a creator, covering the system's machine - learning of creator preference types and creator location - related data. For example, it may include data of the user as a creator, and the learned creator preference types and locations, including geographical locations or specific locations of a place.
[0048] The business data 233 in the user data 23 is when the user is an enterprise, including the enterprise business type and its product characteristics obtained by the system through machine learning.
[0049] The user activity data 25 belongs to the non - public data in the system, which includes statistical data of the user's activities in various services provided by the cloud server, and data obtained through machine learning, mainly including viewer data 251, creator data 252, and business data 253.
[0050] The viewer data 251 is the viewing rate, viewing time of the user using the services provided in the cloud server, and activity data such as following, liking, commenting, subscribing, etc.; the creator profile 252 is the statistical data when the user is in the role of a creator, such as the followers of the channel or account, the number of views of the created content, and the account viewing rate, etc.; the commercial data 253 is the followers, the number of content views, and the overall impression data, etc. obtained when the user is an enterprise.
[0051] The social media platform data 21, user data 23, and user activity data 25 collected and learned by the cloud server above become the basis of the dialogue service proposed in this disclosure implemented by using natural language processing and generative artificial intelligence technologies. In the cloud server, the processing circuit performs calculations on the above various data, thereby implementing a chatbot that can provide personalized and real-time services according to the user's needs.
[0052] According to the embodiment, in the cloud server, the natural language model running therein can first perform a vector algorithm on the content input by the user through the dialogue interface, the user's preferences, and the real-time environment information to annotate the obtained text, calculate the vectors of each word, and obtain relevant content by querying the database according to the vector distance between the words, thereby generating dialogue content that conforms to the user's preferences and real-time environment information. During the online dialogue process, a transformer model can be used to perform procedures such as machine translation, document summarization, and document generation on the texturized data. Then, the semantics of the user in the dialogue can be obtained, enabling the chatbot to generate dialogue content.
[0053] In the intelligent dialogue program started in the intelligent dialogue system triggered by the real-world image, the actual operation can be started through a real-world image interface (such as a social media application program) executed in the user device to open an intelligent dialogue interface (such as a pop-up window or an intelligent dialogue page), which displays the dialogue content between the chatbot and the user.
[0054] Figure 3 Show the flowchart of the embodiment of the method for triggering intelligent dialogue with real-world images. The method for triggering intelligent dialogue with real-world images is executed in a server, such as Figure 1 The shown cloud server can provide augmented reality and intelligent dialogue services through the network, enabling the user device installed with the corresponding application program to trigger intelligent dialogue during the operation of the augmented reality image.
[0055] When the application executed by the user device starts the augmented reality image interface and generates an augmented reality image request to the server, after the server receives the location information and the augmented reality image request transmitted by the user device (step S301), it can calculate multiple visible ranges corresponding to the location information at multiple viewpoints based on the location information (step S303). After the user device transmits the location information of its current location, the software program in the server can calculate the multiple visible ranges that can be seen from the user's location at different viewpoints, and then query the database according to the augmented reality image request and the calculated multiple visible ranges (such as the map database described in Figure 1 ), and it can obtain the environmental objects that can be seen in the image captured by the user device's camera, such as buildings, landmarks, scenic spots, and can also query the context-sensitive data within the visible range. Based on this, it can determine the spatial positions and related text descriptions of the linked images that can mark the context-sensitive data, etc. (step S305).
[0056] After that, the query results of one or more pieces of context-sensitive data, etc. within the visible ranges at different viewpoints are transmitted to the user device (step S307). In the augmented reality image interface initiated by the user device, one or more linked images of one or more pieces of context-sensitive data within the visible range are marked at the corresponding spatial positions, and an intelligent dialogue connection point is also provided.
[0057] Here, it should be noted that when the user device turns on the augmented reality image interface, the image displayed is the augmented reality image of the surrounding area captured by the user device's camera. At this time, the server or the software program executed in the user device can calculate the visible range according to the current location, that is, the height reflected by the location (such as on a building, on a mountain, etc.), and the shooting direction of the user device can reflect the viewing angle. At this time, the visible range is displayed on the user device screen, and at the same time, location information and an augmented reality image request are generated and transmitted to the server. The server can determine the environmental objects that can be seen within the visible range. For example, the location where the user stands may have a building blocking the objects behind, and the objects behind the building do not belong to the environmental objects within the visible range. In this way, the user device can download the relevant information of the environmental objects in the augmented reality image from the server, including the link information of one or more pieces of context-sensitive data within the visible range (such as the link address of the context-sensitive data in the database). By combining the augmented reality image within the visible range captured by the user device's camera and one or more linked images marked at one or more spatial positions, an augmented reality image is formed.
[0058] In this way, the user can operate the user device to move the surrounding images being viewed, and the visible range will also change dynamically, and the environmental objects and the linked images of the context-sensitive data marked at the spatial positions will also change accordingly.
[0059] When a user taps a smart conversation link displayed on the real-world image interface, an intelligent conversation interface is initiated. Simultaneously, the server receives the generated intelligent conversation request from the user device (step S309), initiating the intelligent conversation process between the user device and the server (step S311). This process then introduces a chatbot into the intelligent conversation process, obtaining location information and one or more pieces of location-specific data within the user's visual range, allowing the user to engage in conversation with the chatbot within the intelligent conversation interface.
[0060] Because the proposed intelligent conversation program is activated within a real-world image interface, the chatbot running within it can generate contextually appropriate conversation content related to the user's device's location based on one or more pieces of contextual data and one or more identified environmental objects within the current visual range. In other words, the proposed intelligent conversation program utilizes generative artificial intelligence technology to implement conversational services. The proposed chatbot utilizes a natural language model, enabling not only natural language conversation with the user but also, by taking into account environmental objects and contextual data in the current real-world image, the user's current geographic location, user preferences, and real-time environmental information, enabling conversations that are more relevant to the user's current context.
[0061] Figure 4 A flowchart of an embodiment of a method for triggering intelligent dialogue using real-world images in another display mode is shown.
[0062] According to the illustrated process embodiment, the user operates the software program executed in the user device to start a graphical user interface, such as Figure 7 The diagram shows an embodiment of browsing place-suitability data on a map interface 70. This example shows a user interface with an electronic map as the background, through which place-suitability data in different geographical areas can be browsed. As shown in the figure, audio and video link points 701, 702, and 703 are marked at different locations. Below the interface, some functions provided by the software program are also provided, such as play 711, dialogue 712, assistant 713, search 714, and return to the user homepage 715.
[0063] exist Figure 4 In the process shown, the server receives a signal from the user device that the user has clicked on a point of interest (POI) (step S401). The server queries the database based on the point of interest clicked by the user and obtains information about the point of interest, such as audio and video content, map data, conversation groups, etc. (step S403). Figure 7The displayed audio - video link points 701, 702, 703 can initiate the playback of the browsing page of the text, pattern, or audio - video data linked to this point of interest, and the location - based data can be displayed on this browsing page. If the augmented reality connection point 705 (AR) is selected, the user device activates the camera to capture the surrounding images, thereby starting the augmented reality image mode and initiating the augmented reality image interface. In the augmented reality image mode, the query results obtained by the receiving server querying the database can mark the link patterns of the location - based data at this point of interest at one or more spatial positions in the augmented reality image interface (step S405).
[0064] When the user device activates the augmented reality image mode, reference can be made to Figure 8 The augmented reality interface 80 displayed on the user device by capturing the surrounding images through the camera. In this augmented reality interface 80, as shown, the audio - video link points 801, 802 mark the corresponding link images according to the spatial positions of each piece of location - based data, and may include text descriptions. Through these link images, the location - based data is associated with one of the environmental objects within the visible range. In this example, there are 3 link points displayed below the augmented reality interface 80, all with thumbnails, such as the audio - video link points 801, 802 and the dialogue link point 803 shown in the figure.
[0065] After the user touches the dialogue link point 803 in the augmented reality interface 80, the intelligent dialogue program can be initiated. At this time, the server receives an intelligent dialogue request from the user device (step S407), and then initiates the intelligent dialogue program between the user device and the server (step S409). At this time, a chatbot is introduced into the intelligent dialogue interface initiated by the user device, and at the same time, the location information of the user device and one or more pieces of location - based data within the visible range are obtained, and the intelligent dialogue interface can be initiated in the user device for dialogue.
[0066] Reference can be continued to Figure 9 The displayed augmented reality interface 90. In this example, the visible range captured by the user device camera has a building covering most of the screen. Through the calculation of the software program in the server, it can be known that a certain perspective of the user's location forms a visible range with this environmental object. Therefore, limited by the environmental object, the link pattern displayed through the augmented reality image interface 90, such as the audio - video link point 901 marked on this building. Another dialogue link point 902 is provided, and there is also a prompt text 903 for providing intelligent dialogue services.
[0067] According to the embodiment, in the intelligent dialogue program, the chatbot executes as Figure 5 The flowchart of the embodiment of natural language information processing shown.
[0068] When the user clicks on one of the dialogue link points on the augmented reality image interface (such as Figure 8As shown in the dialogue connection point 803), the cloud server receives the selection of the intelligent dialogue (step S501) and starts the intelligent dialogue program (step S503). Then, an intelligent dialogue interface is initiated, allowing the user to input text, patterns, or specific audio-visual content through the intelligent dialogue interface (for example, input a link to share audio-visual content), so that the cloud server receives the content input by the user through the user interface module (step S505). According to the embodiment, the intelligent dialogue program is implemented as a chatbot using a natural language model, which can converse with the user through the intelligent dialogue interface and execute natural language information processing methods for each piece of content input by the user. The intelligent dialogue interface will provide an input field for the user input content and a dialogue display area for displaying the dialogue content output by the chatbot and the user input content.
[0069] At this time, the cloud server obtains the content input by the user through the user interface module. The content received through the dialogue interface can be text, voice, or audio-visual content. When the received content is voice or audio-visual content, after being converted into text by the text conversion program, semantic analysis can be performed to obtain semantic features (step S507). During the execution of the above procedures, the cloud server obtains user data from the user database and real-time environment information from an external system (which can be obtained through the external system interface module 105 shown) (step S509). Figure 1 displayed external system interface module 105)
[0070] After that, it is possible to determine (or filter through querying the database) the content that conforms to the semantic features of the user input content, the user preferences obtained from the user data, and the real-time environment information (step S511), and after being processed by the natural language model running in the intelligent dialogue program, generate the dialogue content (step S513). Then, the dialogue content is imported into the intelligent dialogue program, and the dialogue content is output at the dialogue interface (step S515). The above steps S505 to S513 can be repeated in the process.
[0071] Furthermore, when the natural language model of the cloud server operates, it uses the database or system memory to record information in multiple dimensions, which may include the historical dialogue records under the same intelligent dialogue program. So that before generating the dialogue, such as in step S511, in addition to considering the semantic features, user preferences, and real-time environment information of the user in the dialogue, the chatbot can also consider the historical dialogue records of the user in the current intelligent dialogue program (step S517), so that the dialogue content generated by the natural language model (step S515) is dialogue content that conforms to the current situation.
[0072] For example, the historical dialogue records of the user in the same intelligent dialogue program often contain the current situation, which can reflect the user's current emotions and needs. Therefore, such asFigure 1 In the illustrated embodiment, the natural language processing model 101 in the cloud server 100 can use the natural language model to simultaneously consider the user's semantics, user preferences, real-time environmental information, and historical conversation records. When generating conversation content, it can continue the same conversation context. For example, it can continue the same conversation topic, and when generating the conversation in natural language, the tone of the language can be the same as that of the previous conversation content (reflecting the user's emotions: joy, anger, sorrow, happiness, etc.), enabling the chatbot to learn the user's emotional expression way through the historical conversation records.
[0073] Refer again to Figure 6 the flowchart of another embodiment of natural language information processing shown.
[0074] In Figure 6 the shown process, the user starts the intelligent conversation program through the application (step S601) and converses with the chatbot, so that the system receives the conversation content input by the user (step S603) to further obtain the semantic features of the user. According to the embodiment, the natural language processing module in the cloud server can perform a transformation calculus (transformer) and vector operation to obtain the semantic features (step S605).
[0075] Here, it should be noted that natural language information processing can use artificial intelligence technology to learn natural language, understand the natural language (natural language understanding), and then perform text classification and grammar analysis. When processing the conversation content input by the user, a deep learning method of a transformation model (transformer model, proposed by the Google Brain team in 2017) can be used to process the natural language content with time sequence input by the user. If the input content is not text, it needs to be texturized first to obtain text. In this way, in the online conversation program, this transformation model can be used to perform machine translation, document summarization, document generation, etc. TM After obtaining the semantic features of the user's conversation content, combined with the user preferences and the user's current location obtained by the system or the location that the user is interested in parsed from the conversation content, and obtain the real-time environmental information from the external system in real time according to this location (step S607). Among them, the real-time environmental information can include any combination of real-time weather, real-time traffic conditions, real-time news, and real-time location-related network information (such as POIs on the map, evaluations of POIs, etc.) obtained from one or more external systems in real time.
[0076]
[0077] After that, the system will use the vector database to calculate the closest answer based on the user's semantic features, user preferences, and real-time environment information, or plus the historical conversation records (step S609). Here, it should be noted that the data in the vector database is structured information obtained by using vector algorithms, enabling the system to obtain words with relatively similar semantics from the retrieved content according to the vector distance. Here is an example: the vector distance between the two words "computer" that appear in the conversation content and the word "calculation" in the database is relatively close, while the vector distance between "computer" and "running" is relatively far.
[0078] In this embodiment, for the content input by the user, the content the user is interested in, and the real-time environment information, and historical conversation records can be added according to requirements. After performing the vector algorithm on the obtained text, calculate the vectors of each word, and obtain relevant content according to the vector distance between the words, and thus generate conversation content that conforms to the user's preferences and real-time environment information. Further, according to the embodiment, when performing the vector algorithm on the historical conversation records recorded in the cloud server, conversation content that can conform to the user's current mood can be generated, such as being able to continue the same topic in the historical conversation records and using words equivalent to the analyzed mood.
[0079] Furthermore, the system also queries the audio-visual database based on the above information to obtain the corresponding audio-visual content, plus the user's location and the environmental objects and location-based content displayed through the augmented reality video interface (step S611). The chatbot will generate conversation content using natural language processing and generative artificial intelligence technologies (step S613) and output the conversation content on the conversation interface (step S615). And, in one embodiment, during the chat process, the system will continuously perform the above steps, enabling the chatbot to converse with the user through natural language (text or voice) and provide content (audio-visual, text) that the user is interested in and is real-time.
[0080] For the relevant embodiments in the intelligent conversation program, reference can be made to Figure 10 the displayed conversation interface 1000, Figure 11 the displayed conversation interface 1100, and Figure 12 the displayed conversation interface 1200, etc. Each example of the displayed conversation interface will provide an input field for the user to input content, and a conversation display area for displaying the conversation content output by the chatbot and the content input by the user.
[0081] For the relevant legends, reference can be made to Figure 10 , Figure 10The display dialogue interface 1000 contains several conversation contents 1001, 1002, 1003 between the user and the chatbot. Moreover, the chatbot can query the database based on the user semantic features obtained from the conversation content 1002 to provide recommended video and audio content 1004. Below the dialogue interface 1000, an input field 1005 is provided for the user to further input conversation content.
[0082] Another mode is, as Figure 11 shown in the dialogue interface 1100. This example shows that when starting an online conversation program, the system directly provides natural language conversation contents 1101, 1102, 1104 according to the user's preferences and real-time information, and directly provides recommended video and audio content 1103. Then the user can use the input field 1105 in the dialogue interface 1100 to respond to the above conversation contents.
[0083] According to another embodiment, the conversation content generated by the natural language model running in the chatbot based on the semantic features of the content input by the user, the user's preferences, and the real-time environment information may also include providing multiple recommended options, multiple recommended video and audio contents, and / or multiple recommended friend links. Please refer to Figure 12 the example shown.
[0084] In the online conversation program, Figure 12 the shown dialogue interface 1200 includes the chatbot generating conversation content 1201 based on the obtained user semantic features. In this example, the semantics makes the chatbot determine that the user is making a choice on a specific matter, so several recommended options 1202 are provided. And particularly, the chatbot provides the user with recommended options according to the real-time environment information obtained by the system from an external system. For example, the chatbot can provide recommended options 1202 according to the real-time climate, road conditions, time, and the user's location. Among them, if the time is exactly meal time and the user's eating habits are referred to, meal options can be provided according to the restaurants that are currently open near the user's location.
[0085] Correspondingly, if the user expresses the desire to watch video and audio content, the recommended options 1202 can be multiple recommended video and audio contents; if the user expresses the desire to find friends with similar interests, the recommended options 1202 can be multiple recommended friend links.
[0086] Further, the user then uses the input field 1206 to respond to these recommended options 1202 and inputs the conversation content 1203, enabling the chatbot to respond to the conversation content 1204 according to the semantics of the conversation content 1203 and to propose multiple recommended contents 1205 according to the semantics of the above conversation content. Continuing with the above example, when the user responds that they want one of the meals, the chatbot can provide restaurant options for the meal the user wants to eat based on the real-time weather, road conditions, and the user's location obtained by the system from an external system. If the weather is bad and there is a traffic jam on the road, correspondingly, it should recommend restaurant options that are convenient for the user to go to.
[0087] According to the above embodiments, different from the chatbot implemented by the current generative artificial intelligence that can only respond to user questions with information obtained from training with past data, the chatbot implemented by the natural language information processing method proposed in this disclosure can generate conversation content that meets the user's needs based on the semantic features parsed from the user's conversation, plus the user preferences learned by the system from the user activity data, and the real-time environmental information.
[0088] In summary, according to the above embodiments, the process and related descriptions of triggering the intelligent dialogue program through the augmented reality image interface are described, thereby providing the user with the ability to obtain context-aware data through the augmented reality image and to obtain conversation content that meets the user's needs within the relevant location through the intelligent dialogue program.
[0089] The content disclosed above is only the preferred feasible embodiment of the present invention, and does not limit the scope of the patent application of the present invention. Therefore, all equivalent technical changes made by using the description and drawings of the present invention are included in the scope of the patent application of the present invention.
Claims
1. A method for triggering intelligent conversations with real - world images, which is executed in a server, and is characterized in that, The method for triggering an intelligent conversation with real - world images includes: Receiving a location information and a real - world image request transmitted by a user device; Calculating a plurality of visible ranges corresponding to the location information from multiple perspectives based on the location information; Querying a database according to the real - world image request and the calculated plurality of visible ranges to obtain one or more pieces of location - based data; Transmitting link information of the one or more pieces of location - based data within each visible range to the user device, where the user device initiates a real - world image interface, marks one or more link images linking the one or more pieces of location - based data at the spatial positions of the one or more pieces of location - based data within the visible range in the real - world image interface, and provides an intelligent conversation connection point; Receiving an intelligent conversation request generated by triggering the intelligent conversation link point displayed on the real - world image interface from the user device, and starting an intelligent conversation program; and In the intelligent conversation program, importing a chatbot, obtaining the location information and the one or more pieces of location - based data within the visible range at the same time, and initiating an intelligent conversation interface in the user device to conduct a conversation.
2. The method for triggering an intelligent conversation with real-world images as claimed in claim 1, wherein, The real - world image interface displays a real - world image of the visible range captured by the user device using a camera, combined with the one or more link images marked at one or more spatial positions, to form an augmented reality image.
3. The method for triggering an intelligent conversation with real-world images as claimed in claim 2, wherein, In the real - world image interface, the spatial positions of each piece of location - based data mark the corresponding link image and a text description, and there is a correlation between the location - based information linked by the link image and one of the environmental objects within the visible range.
4. The method for triggering an intelligent conversation with real-world images according to claim 3, wherein After touching one of the link images, the corresponding location - based information is obtained, and the location - based data is displayed in a browsing page.
5. The method for triggering an intelligent conversation with real-world images according to claim 3, wherein In the intelligent conversation program, the chatbot generates location - based conversation content related to the location of the user device according to the one or more pieces of location - based data within the visible range and one or more identified environmental objects.
6. The method for triggering an intelligent conversation with real - world images as claimed in any one of claims 1 to 5, wherein In the intelligent conversation program, the following steps are performed by the chatbot: Receiving the content input by a user through the intelligent conversation interface; Obtaining the semantic features of the content input by the user; Obtaining the user data and real - time environment information; According to the semantic features of the content that conforms to the content input by the user, the user preferences obtained from the user data, and the content of the real - time environment information, and also referring to the one or more pieces of location - based data within each visible range, running a natural language model to generate a conversation content; and Importing the conversation content into the intelligent conversation program and outputting the conversation content at the intelligent conversation interface.
7. The method for triggering an intelligent conversation with real-world images according to claim 6, wherein The content received through the intelligent conversation interface is text, voice, or audio - visual content. When the received content is voice or audio - visual content, it is converted into text through a text - conversion program, and then the semantic features of the text are obtained through semantic analysis.
8. The method for triggering an intelligent conversation with real-world images as claimed in claim 6, wherein, The natural language model running on the server uses a conversion model to perform procedures such as machine translation, document summarization, and document generation to generate the conversation content.
9. The method for triggering an intelligent conversation with real-world images as claimed in claim 8, wherein, In the server, a vector algorithm is further executed on the content input by the user, the user's preferences, the real-time environment information, and the one or more pieces of location-based data within each visible range, the obtained text is annotated, vectors of each word are calculated, and relevant content is obtained based on the vector distances between words, and thus the conversation content that conforms to the user's preferences and the real-time environment information is generated.
10. The method for triggering an intelligent conversation with real - world images as claimed in claim 6, wherein, The conversation content generated by the natural language model running in the chatbot according to the semantic features of the content input by the user, the user's preferences, the real-time environment information, and the one or more pieces of location-based data within each visible range includes providing multiple recommendation options, multiple recommended video and audio contents, and / or multiple recommended friend links.
11. An intelligent dialogue system triggered by real - world images, characterized in that, The system described above includes: A server, provided with a database, executing a method for triggering intelligent conversation with real-world images, including: Receiving a location information and a real-world image request transmitted by a user device; Calculating multiple visible ranges corresponding to the location information from multiple perspectives based on the location information; Querying the database according to the real-world image request and the calculated multiple visible ranges to obtain one or more pieces of location-based data; Transmitting link information of the one or more pieces of location-based data within each visible range to the user device, where the user device initiates a real-world image interface, and marks one or more link images linking the one or more pieces of location-based data at the spatial positions of the one or more pieces of location-based data within the visible range, and provides an intelligent conversation connection point; Receiving an intelligent conversation request generated by triggering the intelligent conversation link point displayed on the real-world image interface from the user device, and starting an intelligent conversation program; and In the intelligent conversation program, importing a chatbot, and simultaneously obtaining the location information and the one or more pieces of location-based data within the visible range, and initiating an intelligent conversation interface in the user device for conversation.
12. The intelligent dialogue system triggered by real - world images according to claim 11, wherein, The real-world image interface displays a real-world image of the visible range captured by the user device using a camera, combined with the one or more link images marked at one or more spatial positions, to form an augmented reality image. Among them, the spatial positions of each piece of location-based data are marked with corresponding link images and a text description, and there is a relevance between the location-based data linked by the link image and one of the environmental objects within the visible range.
13. The intelligent dialogue system triggered by real - world images as described in claim 12, wherein By querying the map data and video and audio data in the database, one or more environmental objects within the visible range and the location-based data corresponding to each environmental object are obtained, so that the link images and the text description are accurately marked at the spatial positions corresponding to each piece of location-based data.
14. The intelligent dialogue system triggered by real - world images according to claim 13, wherein, In the intelligent conversation program, the chatbot generates a location-based conversation content related to the location of the user device according to the one or more pieces of location-based data within the visible range and the one or more identified environmental objects.
15. The intelligent dialogue system triggered by real - world images according to claim 14, wherein The server provides an external system interface for connecting to one or more external systems to obtain real-time environment information, so that the chatbot further generates the location-based conversation content according to the real-time environment information.
16. The intelligent dialogue system triggered by real - world images according to any one of claims 11 to 15, characterized in that, In the intelligent dialogue program, the following steps are executed by the chatbot: Receiving the content input by a user through the intelligent dialogue interface; Obtaining the semantic features of the content input by the user; Obtaining the user data and real-time environment information; According to the semantic features of the content that conforms to the content input by the user, the user preferences obtained from the user data, and the content of the real-time environment information, and also referring to the one or more pieces of location-based data within each visible range, running a natural language model to generate a dialogue content; and Importing the dialogue content into the intelligent dialogue program and outputting the dialogue content through the intelligent dialogue interface.
17. The intelligent dialogue system triggered by real - world images according to claim 16, wherein The content received through the intelligent dialogue interface is text, voice, or audio-visual content. When the received content is voice or audio-visual content, it is converted into text through a text conversion program and then the semantic features of the text are obtained through semantic analysis.
18. The intelligent dialogue system triggered by real-world images as described in claim 16, wherein The natural language model running on the server uses a conversion model to perform procedures such as machine translation, document summarization, and document generation to generate the dialogue content.
19. The intelligent dialogue system triggered by real - world images according to claim 18, wherein, In the server, further performing a vector algorithm on the content input by the user, the user's preferences, the real-time environment information, and the one or more pieces of location-based data within each visible range, annotating the obtained text, calculating the vectors of each word, and obtaining relevant content based on the vector distances between the words, and generating the dialogue content that conforms to the user's preferences and the real-time environment information accordingly.
20. The intelligent dialogue system triggered by real - world images as described in claim 16, wherein The dialogue content generated by the natural language model running in the chatbot according to the semantic features of the content input by the user, the user preferences, the real-time environment information, and the one or more pieces of location-based data within each visible range includes providing multiple recommended options, multiple recommended audio-visual contents, and / or multiple recommended friend links.