Information processing method and related device
By collecting and tracking data to learn user preferences, and selecting appropriate vertical applications or models to process user intent, the problem of low information delivery efficiency in existing technologies is solved, enabling more efficient dialogue between users and electronic devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to efficiently provide information based on user needs, resulting in inefficient interactions between users and electronic devices.
By collecting event tracking data, we learn users' usage preferences in various vertical domains, select the most popular vertical domain applications or models to process user intent, and output corresponding responses.
It improves the efficiency of user interaction with electronic devices, ensuring that the information provided better meets user needs.
Smart Images

Figure CN121809641A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to information processing methods and related devices. Background Technology
[0002] With the rapid development of electronic technology, the application of visual question answering (VQA) is becoming increasingly widespread. For example, smart devices such as mobile phones and tablets can create intelligent virtual assistants through large language models that support VQA. Users can engage in one or more rounds of question-and-answer dialogues with the virtual assistant, enabling the virtual assistant to perceive the user's intent and present the information the user needs.
[0003] How to better provide users with the information they need in response to their questions and demands is a problem that urgently needs to be solved by people in this field. Summary of the Invention
[0004] The purpose of this application is to provide an information processing method and related apparatus. Implementing this method, an electronic device can learn the user's usage preferences for large language models or applications in various vertical domains based on embedded data collected during the user's historical use of the electronic device. During visual question-and-answer sessions between the user and the electronic device, the electronic device can select the user's preferred large application (model) within the vertical domain to process the user's intent, based on the current user intent's domain, and output the response obtained from that large application (model) to the user. In this way, the electronic device's responses based on user intent can maximize the satisfaction of user needs and improve the efficiency of the dialogue between the user and the electronic device.
[0005] The aforementioned and other objectives will be achieved through the features described in the independent claims. Further implementations are illustrated in the dependent claims, the specification, and the drawings.
[0006] In a first aspect, this application provides an information processing method, comprising: determining a first user intent based on a first user input; determining a first model or a first application based on embedded data and a first vertical domain to which the first user intent belongs, wherein the embedded data characterizes the frequency of use of each model or application when the user realizes the user intent in each vertical domain, and the first model or the first application is a model or application whose frequency of use when realizing the user intent in the first vertical domain is greater than a first threshold; invoking the first model or the first application to process the first user intent, and outputting the obtained first processing result.
[0007] In this method, the electronic device can learn the user's preferred applications or large language models for a specific vertical domain based on the embedded data collected during the user's historical use of the electronic device. Then, when the user engages in VQA with the electronic device, the electronic device can determine the user's current intended user intent based on the first user input, determine the first vertical domain to which this intent belongs, and combine this with the frequency of various applications or models used in the electronic device's history of achieving user intents within the first vertical domain. It then determines the first model or application used more than a first threshold to achieve the user intent within that first vertical domain, and processes the first user intent using this first model or application to obtain the first user result. Understandably, the first processing result is obtained from the user's preferred models or applications within the vertical domain to which the first user intent belongs.
[0008] Optionally, the specific value of the first threshold can be 1000, 2000 or other values, and this application does not limit it.
[0009] Optionally, when there are multiple models or applications whose frequency of use is greater than the first threshold when implementing the user intent of the first vertical domain, the first model and the first application can be the model or application with the highest frequency of use when implementing the user intent of the first vertical domain.
[0010] Optionally, the aforementioned data may record the number of times each model or application is used when an electronic device realizes user intent in various vertical domains.
[0011] Optionally, the vertical domain to which the user's intent pertains may include, but is not limited to, food, fashion, beauty, general knowledge, and medical categories. Optionally, in some embodiments, the vertical domain can be further subdivided according to the user's specific needs. For example, the aforementioned food vertical domain can be further refined into a cooking learning vertical domain, a restaurant exploration vertical domain, etc. This application does not limit this. Specifically, the specific division of the vertical domain can be predetermined at the time of manufacture of the electronic device. The method of dividing the vertical domain is the same when collecting data and when subsequently dividing it according to the user's intent.
[0012] In conjunction with the first aspect, in one possible implementation, when the first model is invoked to process the first user intent, the first processing result includes one or more of the following: a first overview summary, at least one first content snapshot, and at least one first related prompt. The first overview summary is a summary obtained by the first model from the search results obtained from the search of the user intent. The at least one content snapshot is a snapshot corresponding to at least one link contained in the search results. The at least one first related prompt is a prompt generated by the first model after predicting subsequent user intents based on the search results. Alternatively, when the first application is invoked to process the first user intent, the first processing result includes a first application service card, the first application service card corresponding to a first application service, and the first application service being an application service provided by the first application for processing the first user intent.
[0013] In this embodiment, the electronic device can use the first model to process user intent within the first vertical domain. In this case, the first processing result obtained after processing by the first model may include the first overview summary, at least one first content snapshot, and at least one first related prompt. The overview summary is the textual description obtained by the first model after summarizing and refining the searched content. The at least one first content snapshot is a snapshot generated by the first model based on the searched webpage. Each snapshot can display a portion of the corresponding webpage content (e.g., key titles and / or images), and each snapshot can respond to user actions, such as a click, causing the electronic device to display the full content of the webpage corresponding to that snapshot. The at least one first related prompt is a prompt message generated by the first model for the searched content. This prompt message indicates the user's potential next intent. Any first related prompt can respond to user actions, such as a click, allowing the electronic device to obtain the user intent contained in the related prompt, so that the user can quickly engage in the next round of dialogue with the electronic device.
[0014] The aforementioned electronic device can select the first application to process the user's intent within the first vertical domain. In this case, the first application can select a service from its available application services that matches the user's intent, process it, and return the processing result to the electronic device. Specifically, the processing result can be returned to the electronic device in the form of a first application service card. Optionally, when the electronic device displays the first application service card, it can display all or part of the content obtained after the application service processes the user's intent. For example, when the first user intent is to navigate to a certain location, the first application service card can be a navigation service card provided by a map application. After the user clicks the first application service card, the electronic device can directly display the navigation service interface, and the content displayed in the navigation service interface can include route maps and other content provided by the navigation service after searching for the location (without requiring the user to input the location again and click the navigation control in the navigation service interface).
[0015] In conjunction with the first aspect, in one possible implementation, the first user input includes a first image and first text, and determining the first user intent based on the first user input includes: extracting entities contained in the first image to obtain first entity information, wherein the first entity information is an image-type entity or a text-type entity; recognizing the first text to obtain a second user intent, wherein the second user intent is a user intent lacking slot information; and using the first entity information to fill the missing slot information in the second user intent to obtain the first user intent.
[0016] In this embodiment, the first user input includes both an image and text. In this case, to obtain accurate user intent, the electronic device will separately recognize the first image and the first text to determine the entity contained in the image and the vague user intent. This vague user intent needs to be combined with the entity information extracted from the first image to become a definite user intent. For example, assuming the entity information contained in the input first image is a dog, and the user inputs the first text "What is this?", it is understandable that based solely on the first text "What is this?", it can only be determined that the user intent is to understand information related to a certain object, but it cannot determine what that object is. Therefore, it is necessary to combine the entity extracted from the first image to understand the user's current specific user intent.
[0017] Specifically, the user can input the first text via keyboard or voice, and the user can input the first image via the "Anywhere Door" function, or directly input the first image in the dialog box through the corresponding file access control. This application does not limit the specific input method for the first image and first text included in the aforementioned first user input.
[0018] Optionally, the vague intent obtained from the first text can be determined by certain fields contained in the text. For example, when the user input text includes fields such as "navigation" or "how to get there," it can be determined that the current user intent may be to navigate to a certain location. When the user input text includes fields such as "make a phone call" or "call," it can be determined that the current user intent may be to call a certain contact. Understandably, the field content corresponding to each user intent can have one or more other forms of expression, and this application does not limit this.
[0019] Optionally, after determining the vague user intent, the electronic device can pre-reserve a slot for each vague intent. This slot can be used to accommodate one or more entities extracted from the first image, in conjunction with the vague intent. Figure 1 The user intent is then transformed into a slotted user intent. For example, when the user intent is to navigate to a certain place, the electronic device reserves a destination slot in the user intent. In this case, the vague user intent can be represented by the text "navigate to (destination)", where "destination" is a pronoun that will be replaced by the address-type entity extracted from the image later.
[0020] In conjunction with the first aspect, in one possible implementation, the aforementioned tracking data also characterizes the number of times the user performs various user intentions for each type of entity data. The method further includes: updating the number of times the user performs the second user intention for the first entity type in the aforementioned tracking data based on the first entity type to which the first entity information belongs and the second user intention.
[0021] In this embodiment, the electronic device can update the number of times the user performs the second user intent on the first entity type based on the first entity type to which the first entity information belongs and the second user intent. This allows the electronic device to more accurately learn the user intents that frequently appear for various types of entity data from the tracking data. Thus, when the user performs VQA with the electronic device later, even if the user only inputs an image, the electronic device can determine the user intent that the user might have for the entities in the image from the tracking data, improving the efficiency of VQA between the electronic device and the user.
[0022] In conjunction with the first aspect, in one possible implementation, the first user input includes only the second image, and the aforementioned tracking data also characterizes the number of times the user performs various user intentions for each type of entity. Determining the first user intention based on the first user input includes: extracting entities contained in the second image to obtain second entity information; determining at least one third user intention based on the tracking data and the second entity type to which the second entity information belongs, and generating at least one second associated prompt based on the at least one third user intention, wherein the at least one third user intention includes one or more user intentions performed most frequently by the user for entity information of the second entity type; and, in response to the user's click operation on the third associated prompt in the at least one second associated prompt, combining the second entity information with the fourth user intention corresponding to the third associated prompt to obtain the first user intention.
[0023] In this embodiment, the first user input only includes the second image. In this case, the system first identifies the second image, determines the second entity information contained in the image, and, based on the obtained entity information, retrieves the number of times various user intentions under the second entity type contained in the second image are realized from the embedded data. Based on at least one third intention frequently realized by the user under the second entity type, at least one corresponding second associated prompt is generated for the user to choose from. When the user selects one of the third associated prompts, the ambiguous intention implied by the associated prompt can be concatenated with the second entity information to obtain the specific user intention.
[0024] Assuming the entity information in the second image is a clothing-type entity, and the user has not entered any text, the electronic device will first identify the clothing-type entity from the second image. Then, the electronic device can obtain from the aforementioned embedded data the number of times the user has historically performed various intentions when the entity type is clothing. Assuming the user's most frequently performed intentions are "clothing search" (30 times) and "entity elimination" (20 times), the electronic device can combine these most frequently performed intentions with the clothing-type entity and provide the user with related prompts on the interface, such as "Which brand is this clothing?", "Other clothing items to match this clothing?", or "Eliminate the clothing in the image" (i.e., at least one second user prompt corresponding to at least one third user intention), for the user to choose from. After the user selects a related prompt, the electronic device can obtain the user intention corresponding to that prompt and combine this user intention with the entity extracted from the second image to obtain the first user intention.
[0025] In conjunction with the first aspect, in one possible implementation, in response to the user's click operation on the fourth associated prompt, the number of times the user has performed the fourth user intent on the second entity type in the above-mentioned tracking data is updated based on the second entity type to which the second entity information belongs and the fourth user intent.
[0026] In this embodiment, if the user only inputs an image and expresses their intent by clicking on the associated prompt, the electronic device will record the user intent expressed for the second entity in the second image and update the user preference information stored in the event tracking data. For example, if the user selects the associated prompt "Which brand of clothing is this product?", the user intent for the clothing type entity will be "Search for which brand of clothing is the clothing in the image." Accordingly, the electronic device can update the number of times the user performs a "clothing search" for clothing entities recorded in the event tracking data from the original N times to (N+1) times. In this way, when the user performs VQA with the electronic device later, even if the user only inputs an image, the electronic device can determine the user intent that the user may have for the entity in the image from the event tracking data, improving the efficiency of VQA between the electronic device and the user.
[0027] Secondly, this application provides an information processing method applied to an electronic device. The method includes: determining a fifth user intent based on a second user input; decomposing the fifth user intent into at least two sub-user intents, the at least two sub-user intents including a first sub-user intent and a second sub-user intent; determining a second model or a second application based on embedded data and a second vertical domain to which the first sub-user intent belongs; determining a third model or a third application based on the embedded data and a third vertical domain to which the second sub-user intent belongs; the embedded data characterizing the frequency of use of each model or application when the user realizes the user intent in each vertical domain; and the second model... Alternatively, the second application mentioned above uses a model or application with a frequency greater than the first threshold when implementing the user intent of the second vertical domain; the third model or the third application mentioned above uses a model or application with a frequency greater than the first threshold when implementing the user intent of the third vertical domain; the second model or the second application is invoked to process the first sub-user intent to obtain a second processing result; the third model or the third application is invoked to process the second sub-user intent to obtain a third processing result; some or all of the multiple processing results obtained from processing the at least two sub-user intents are displayed, and the multiple processing results include the second processing result and the third processing result.
[0028] Understandably, some user intents are divisible. This divisibility manifests in the fact that to achieve a given user intent, a user often needs to simultaneously fulfill multiple sub-user intents. For example, suppose a user currently has a user intent of "visiting the Forbidden City." To achieve this intent, the user might need to fulfill sub-user intents including, but not limited to, learning about the Forbidden City (e.g., its location, opening hours, and whether tickets are required), browsing nearby restaurants, creating travel plans / guides, booking tickets / hotels, and creating schedules / alarms for the travel plan. These sub-user intents belong to various verticals, such as general knowledge, food, and travel. Furthermore, the user preference models and applications differ depending on the specific vertical the sub-user intent belongs to.
[0029] Therefore, in this method, the electronic device can decompose the aforementioned fifth user intent into at least one sub-user intent. Optionally, the multiple sub-user intents obtained from the decomposition can belong to different vertical domains. The electronic device can, according to the user's application (model) usage preferences for each vertical domain, send each of the at least one sub-user intent to the corresponding third-party application (large language model) for processing, and obtain the processing results of multiple applications (large language models) on the sub-user intents, and output all or part of them to the user. In this way, the electronic device can provide richer information in the response content provided to the user during the VQA process, and the response content can better cover the user's actual needs, improving the VQA efficiency between the user and the electronic device.
[0030] In conjunction with the second aspect, in one possible implementation, the aforementioned multiple processing results include a first type of processing result and a second type of processing result. The first type of processing result is a fourth processing result obtained by processing the third sub-user intent using at least one fourth model. The second type of processing result includes a fifth processing result obtained by processing the fourth sub-user intent using at least one fourth application. The fourth processing result includes one or more of a second overview summary generated based on the first sub-user intent, at least one second content snapshot, and at least one fifth related prompt. The fifth processing result includes an application service card corresponding to at least one application service provided by the at least one second application for the second sub-user intent. Displaying some or all of the aforementioned multiple processing results includes: filtering the multiple processing results and displaying the fourth and fifth processing results.
[0031] In this embodiment, the vertical domains to which the at least one sub-user intent belongs may be different. In some vertical domains, the user's preference may be a model provided by the application, while in other vertical domains, the preference may be an application service provided by the application. The content of the processing results obtained by the model and the application service after processing the user intent is different. Therefore, in this embodiment, the multiple processing results obtained by the electronic device based on the at least one user sub-intent may include the processing results obtained by the model (e.g., the fourth model mentioned above) or the processing results obtained by the application service (e.g., at least one application service provided by the second application for the second sub-user intent). All results will be further filtered by the electronic device, and the displayed results after filtering will simultaneously provide processing results from the model and application service (e.g., the fourth and fifth processing results mentioned above). This ensures that the final displayed results come from diverse sources and also provides specific application service cards to the user, enabling the user to successfully locate deeper application services in VQA, further improving the efficiency of VQA between the electronic device and the user.
[0032] In conjunction with the second aspect, in one possible implementation, the second user input includes a third image and second text. Determining the fifth user intent based on the second user input includes: extracting entities contained in the third image to obtain third entity information, wherein the third entity information is an image-type entity or a text-type entity; recognizing the second text to obtain a sixth user intent, wherein the sixth user intent is a user intent lacking slot information; and using the third entity information to fill the missing slot information in the sixth user intent to obtain the fifth user intent.
[0033] In this embodiment, the second user input includes both an image and text. In this case, to obtain accurate user intent, the electronic device will separately recognize the third image and the second text to determine the entity contained in the image and the vague user intent. This vague user intent needs to be combined with the entity information extracted from the third image to become a clear user intent. For example, assuming the entity information contained in the input first image is a dog, and the user inputs the second text "What is this?", it is understandable that based solely on the first text "What is this?", it can only be determined that the user intent is to understand information related to a certain object, but it cannot determine what that object is. Therefore, it is necessary to combine the entity extracted from the first image to understand the user's specific current user intent.
[0034] Specifically, users can input the second text via keyboard or voice, and can input the third image via the "Anywhere Door" function, or directly input the third image through the corresponding file access control in the dialog box. This application does not limit the specific input method for the third image and second text included in the second user input.
[0035] In conjunction with the second aspect, in one possible implementation, the aforementioned tracking data also characterizes the number of times the user performs various user intentions for each type of entity data. The method further includes updating the number of times the user performs the sixth user intention for the third entity type in the aforementioned tracking data based on the third entity type to which the third entity information belongs and the sixth user intention.
[0036] In this embodiment, the electronic device can update the number of times the user performs the sixth user intent for the third entity type based on the third entity type and the sixth user intent in the aforementioned tracking data. This allows the electronic device to more accurately learn the user intents that frequently appear for various types of entity data from the tracking data. Thus, when the user performs VQA with the electronic device later, even if the user only inputs an image, the electronic device can determine the user intent that the user might have for the entities in the image from the tracking data, improving the efficiency of VQA between the electronic device and the user.
[0037] In conjunction with the second aspect, in one possible implementation, the second user input includes only the fourth image, and the aforementioned embedded data also characterizes the number of times the user performs various user intentions for each type of entity data. Determining the fifth user intention based on the second user input includes: extracting entities contained in the fourth image to obtain fourth entity information, where the third entity information is an image entity or a text entity; determining at least one seventh user intention based on the embedded data and the fourth entity type to which the fourth entity information belongs, and generating at least one fifth associated prompt based on the at least one seventh user intention, where the at least one seventh user intention includes one or more user intentions performed most frequently by the user for entity information of the fourth entity type; and in response to the user's click operation on the sixth associated prompt in the at least one fifth associated prompt, combining the fourth entity information with the eighth user intention corresponding to the sixth associated prompt to obtain the fifth user intention.
[0038] In this embodiment, the second user input only includes the fourth image. In this case, the system first identifies the fourth image, determines the fourth entity information contained in the image, and, based on the obtained fourth entity information, retrieves the number of times various user intentions are realized under the fourth entity type contained in the fourth image from the embedded data. Based on at least one seventh intention frequently realized by the user under the fourth entity type, at least one corresponding fifth association prompt is generated for the user to choose from. If the user selects one of the sixth association prompts, the ambiguous intention implied by the sixth association prompt can be concatenated with the fourth entity information to obtain the specific user intention.
[0039] In conjunction with the second aspect, in one possible implementation, in response to the user's click operation on the sixth associated prompt, the number of times the user has historically implemented the eighth user intent for the fourth entity type is updated in the aforementioned tracking data, based on the fourth entity type to which the fourth entity information belongs and the eighth user intent.
[0040] In this embodiment, if the user only inputs the fourth image and expresses their user intent by clicking the associated prompt, the electronic device will record the user intent expressed for the second entity in the second image and update the user preference information stored in the event tracking data. For example, if the user selects the associated prompt "Which brand of clothing is this product?", then the user intent for the clothing type entity is "Search for which brand of clothing is the clothing in the image." Accordingly, the electronic device can update the number of times the user performs "clothing search" for clothing entities recorded in the event tracking data from the original N times to (N+1) times. In this way, when the user performs VQA with the electronic device later, even if the user only inputs an image, the electronic device can determine the user intent that the user may have for the entity in the image from the event tracking data, improving the efficiency of VQA between the electronic device and the user.
[0041] Thirdly, this application provides an electronic device comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, wherein the one or more processors invoke the computer instructions to cause the electronic device to perform a method in the first aspect or any possible implementation of the first aspect, or a method in the second aspect or any possible implementation of the second aspect.
[0042] Fourthly, this application provides a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform a method as described in the first aspect or any possible implementation thereof, or a method as described in the second aspect or any possible implementation thereof.
[0043] Fifthly, this application provides a computer program product containing instructions that, when the computer program product is run on an electronic device, causes the electronic device to perform the method as described in the first aspect or any possible implementation of the first aspect, or the method as described in the second aspect or any possible implementation of the second aspect.
[0044] In a sixth aspect, this application provides a computer-readable storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform a method as described in the first aspect or any possible implementation thereof, or a method as described in the second aspect or any possible implementation thereof. Attached Figure Description
[0045] Figure 1 An architecture diagram of a visual question-answering system provided in this application embodiment;
[0046] Figure 2 User interface diagrams involved in some visual question-and-answer processes provided in the embodiments of this application;
[0047] Figure 3 An architecture diagram of a visual question-answering system provided in this application embodiment;
[0048] Figure 4 This application provides a schematic diagram illustrating a process for determining user intent based on images and text in an embodiment of the present application.
[0049] Figure 5 This application provides a schematic diagram of a process for determining user intent based on an image.
[0050] Figure 6 Some user interface diagrams related to visual question answering in the food industry provided in the embodiments of this application;
[0051] Figure 7 Some user interface diagrams related to visual question answering in the clothing vertical field provided in the embodiments of this application;
[0052] Figure 8 A user interface diagram related to a visual question-and-answer process provided in an embodiment of this application;
[0053] Figure 9An architecture diagram of a visual question-answering system provided in this application embodiment;
[0054] Figure 10 This application provides a schematic diagram illustrating the process of determining and deconstructing user intent based on images and text in an embodiment of the present application.
[0055] Figure 11 This application provides a schematic diagram illustrating a process for determining and deconstructing user intent based on an image, as exemplified in this application.
[0056] Figure 12 This application provides some user interface diagrams related to a clothing vertical domain provided in its embodiments;
[0057] Figure 13 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0058] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0059] Since the embodiments of this application involve the application of neural networks, for ease of understanding, the relevant terms involved in the embodiments of this application will be introduced below.
[0060] (1) Visual Question Answering
[0061] Visual question answering (VQA) is a technique provided by large language models that enables computers to "understand" images and answer questions about those images.
[0062] Visual question answering (VQA) technology comprises two key components: computer vision (CV) and natural language processing (NLP). CV identifies objects, scenes, and features in an image, as well as the relationships between them. NLP can be used for tasks such as text translation, sentiment analysis, and text generation. VQA combines these two fields, guiding users to formulate their needs, or combining user-initiated text reflecting their needs (which often relates to image content), and then the electronic device analyzes the image and answers the questions using natural language.
[0063] It should be noted that in some large language models used for visual question answering, the large language model can search its corresponding training set to obtain the answer when answering a question; optionally, the large language model can also collaborate with a third-party search engine to search for the question. Specifically, the answer provided by the large language model to a user's question can include, but is not limited to, one or more of the following: text, image, video, link, and snapshot.
[0064] (2) Enhanced search generation
[0065] Retrieval-augmented generation (RAG) is an advanced artificial intelligence technique that combines retrieval and generation capabilities. This technique enhances the performance of large language models in handling specific tasks, particularly those requiring access to external knowledge bases or real-time information, by integrating external knowledge bases. RAG models aim to overcome the limitations of large language models, such as limited storage capacity, difficulty in obtaining up-to-date information, and insufficient knowledge in specific domains, by integrating retrieval mechanisms to assist the model in generating more accurate, detailed, and targeted answers.
[0066] For example, in this application, the large language model used to implement visual question answering can include a RAG model. After searching for a question and obtaining various types of information corresponding to the question, the electronic device can use the RAG model to refine the obtained information and obtain an overview summary so that the user can quickly understand the whole picture of the answer to the current question.
[0067] (3) Any Door Function
[0068] In this application, the "Anywhere Door Function" is an interactive method provided by an electronic device based on the user's needs for cross-application and cross-device information transmission. Specifically, it can be a functional user interface provided by the electronic device based on the user's needs for cross-device or cross-application information transmission, also known as the "Anywhere Door Interface." As its name suggests, under any interface displayed by the electronic device, when the user needs to transmit information across applications, the user only needs to select the information, long-press the information element in the interface, and drag it to the edge of the screen. The electronic device will then perform a three-dimensional transformation of the originally displayed interface, creating a visual effect of the original interface pushing inwards, and display the application service icon recommended by the electronic device for that element on the side of the screen (i.e., the Anywhere Door Interface, for details, please refer to the relevant descriptions in the subsequent embodiments, which will not be repeated here). Afterwards, the user can place the dragged information element on the icon and release it to transmit the information element to the application service corresponding to the icon and start the corresponding application service function so that the electronic device displays the corresponding application service interface.
[0069] In this application, users can drag and drop elements (including one or more images and text) from the interface displayed on an electronic device to activate the "Any Door" function, causing the application service icon corresponding to the visual question-and-answer service to appear on the Any Door interface. After the user drags and drops an element onto the application service icon and releases it, the electronic device can use the dragged element as input for the visual question-and-answer service and display the corresponding service interface. Users can browse the information replied by the electronic device on this service interface and, by clicking on related prompts or entering images / text again through the input boxes, engage in one or more subsequent rounds of dialogue with the electronic device.
[0070] (4) Vertical domain and vertical domain large model
[0071] A vertical domain, also known as a vertical field, refers to a content area that focuses on a specific part of a particular industry or topic. Vertical domains can be divided according to users' actual needs, such as the medical field, the education field, and the commodity economy field based on different industries.
[0072] Large language models used to implement VQA can be mainly divided into general-purpose large models and vertical-domain large models. Among them:
[0073] General-purpose models possess cross-task and cross-domain versatility, and are generally suitable for non-domain-specific task scenarios such as chat. They are not optimized for specific domains or tasks, but rather aim to provide broad language processing capabilities, such as the Tongyi Qianwen model launched by Alibaba Cloud and the Doubao model launched by ByteDance.
[0074] In contrast to general-purpose large models, vertical-domain large models are trained for specific industries or fields. They focus more on knowledge and tasks in one or a few specific areas, such as commodity economics, healthcare, and education. Through pre-training in specific domains and applying datasets, vertical-domain large models can provide more accurate and professional language processing services within these domains. For example, JD.com's Yanxi model is a vertical-domain large model developed for the e-commerce industry, which can be applied to multiple business scenarios such as marketing, customer service, and delivery.
[0075] In this application, the default large language model used by the electronic device can be a general large model, while the large language model provided by third-party applications installed on the electronic device can be a vertical domain large model. Optionally, the general large model used by the electronic device can be deployed on the electronic device itself or on a cloud server provided by the electronic device manufacturer. Before accurately identifying user intent, the electronic device can communicate with the user through the general large model; after accurately identifying user intent, the electronic device can identify the specific vertical domain corresponding to the user's needs and, in conjunction with user preferences, select a vertical domain large model provided by a third-party application to process the user's needs, thereby more accurately meeting the user's requirements.
[0076] (5) Data embedding
[0077] In this application, the data collected by the device may be the data collected by the electronic device each time the user launches the application to perform an image / text search.
[0078] Specifically, the application launched by the user can be any third-party application installed on the electronic device, or it can be a system application installed on the electronic device. When the user launches the application to search, the parameters entered can be images, text, or images and text. The search can be launched by dragging and dropping an element onto the application service icon contained in any door and then releasing it, or by clicking the application icon normally to enter the application's main interface, and then clicking the controls or function options in the interface to jump to the corresponding search interface and enter information to search, or by other launch methods (such as voice launch). This application does not limit this.
[0079] Each time a user performs a search, the electronic device collects a corresponding data point. This data point records the entities contained in the user's input (images and text), the user's specific intent, the vertical domain to which the intent belongs, and the application launched by the user for that search (or, in other words, the large language model used for that search). For example, suppose the user uses a social application 1 (e.g., If an image of clothing is searched, the electronic device will record a tracking data point corresponding to that search. The tracking data point includes: {Entity - Clothing; User Intent - Search for clothing in the image, Vertical Domain - Clothing, Application Used - Social Application 1}.
[0080] In the embodiments of this application, when there is enough data collected by the user, the electronic device can learn user habits from the collected data, including the user's usage preferences for applications (models) in various vertical domains, and the user's frequently realized intentions for various types of entities.
[0081] On one hand, during the dialogue between the user and the electronic device, after recognizing the user's intent, the electronic device can determine the most likely model the user will use and the application to which that model belongs based on the vertical domain to which the user's intent belongs. The electronic device will then send the specific user intent to the application, call the model provided by the application to perform a search, and the model provided by the application will obtain the corresponding content snapshot, as well as a generative summary and related suggestions based on the search results. Afterwards, the application can return this content to the electronic device for the electronic device to output to the user for browsing. For details, please refer to the relevant descriptions in the following embodiments.
[0082] On the other hand, during the question-and-answer process, if the general large language model cannot accurately understand the user's intent (e.g., when the user only inputs an image but not any text), the electronic device can extract entities from the image and determine one or more user intents that the user commonly uses for that entity based on the collected embedded data. Based on these one or more user intents, the electronic device can display corresponding related prompts on the interface for the user to choose from. After the user selects a related prompt, the electronic device can determine the user's current user intent based on the related prompt, so that the question-and-answer dialogue can proceed normally. For details, please refer to the relevant descriptions in the following embodiments.
[0083] (6) Application Service Card
[0084] Application service cards are a new form of application display page content. They can bring the content of the application page to the front of the application service card, so that users can achieve the application experience by directly interacting with the application service card, thereby achieving the goal of direct service access and reducing experience layers.
[0085] Application service cards are commonly embedded into other applications as part of the user interface (or the application can be saved to the service center using atomic services, which eliminates the need for application installation). They support basic interactive functions such as page launches and message sending. Application service cards have three main features: 1. Ease of use and visibility: Application service cards allow prominent service information to be displayed, reducing user experience issues caused by hierarchical navigation. 2. Intelligent selectability: Cards can display variable data information throughout the day and support custom application service card designs, allowing users to customize card styles. 3. Multi-device adaptability: Service cards can adapt to various devices, including mobile phones, smartwatches, and tablets.
[0086] With the rapid development of electronic technology, the application of VQA (Voice Quality Assurance) is becoming increasingly widespread. In daily use of electronic devices, users can drive large language models to engage in conversations using the voice assistants or "anywhere door" functions provided by the devices. For details, please refer to... Figure 1 and Figure 2 .
[0087] Figure 1 An example is shown in the architecture diagram of a system that implements VQA based on a large language model.
[0088] like Figure 1 As shown, the visual question-answering system 10 may include a processing module 101, a third-party question-answering module 102, and a system tool module 103. The processing module 101 may be a processor of an electronic device, and may include a central control module 101a, an image semantic understanding module 101b, and a text semantic understanding module 101c. The third-party question-answering module 102 may correspond to a general large model, specifically including a search module 103a, an overview module 103b, a snapshot module 103c, and a prompting module 103d. Optionally, if the general large model is deployed in the electronic device, the third-party question-answering module may be included in the processing module 101. The system tool module 103 may be a functional component provided by the operating system used by the electronic device. Optionally, in some embodiments, the above-mentioned system tool module may also be included in the processing module 101.
[0089] Specifically, during the process of a user interacting with an electronic device through the visual question-answering system 10, the processing module 101 first acquires user input. This user input can be one or more of an image or text. As explained above, in visual question-answering technology, the image is responsible for providing the objects, scenes, and features involved in the user's intent, as well as the relationships between them, while NLP is responsible for analyzing the user's sentiment and determining the user's intent. After acquiring the user input (assuming that the user input includes both images and text), the image semantic understanding module 101b and the text semantic understanding module 101c will respectively recognize the image and text, determine the entities contained in the image and the user's intent, and then the central control module 101a will determine whether the user's intent can be achieved solely through system tools.
[0090] When the user intent can be realized solely through system tools, the central control module 101a sends the acquired user intent along with the entity information contained in the image to the system tool module 103. The system tools provided by the system tool module 103 process the entities contained in the image and return the processing result to the central control module 101a. Optionally, the system tools provided by the system tool module may include text extraction, text translation, entity matting, entity removal, etc., and may also include other system tools. The result of image processing by the corresponding system tool module 103 may be text extracted from the image, or a new image obtained after processing the text and / or other entities in the image. Taking entity removal as an example, suppose a user inputs a landscape photo containing both buildings and pedestrians, along with the text "remove pedestrians from the image". After the image semantic understanding module 101b and the text semantic understanding module 101c process the user input, the central control module can obtain the entity information (i.e., buildings and pedestrians) contained in the image and the clear user intent (i.e., "remove pedestrian entities from the image"). The central control module 101a can determine that the user intent can be directly realized through the system tools of the electronic device. The central control module 101a will then send the entity information contained in the image and the user intent to the system tool module 103. The system tool module 103 will remove the pedestrian entities in the image according to the user intent and return the image without pedestrian entities to the central control module.
[0091] When the user intent is not realized through system tools, the central control module 101a will call the third-party question-answering module to realize the user intent. Specifically, the third-party question-answering module 102 retrieves the response information corresponding to the current user intent using the search capabilities of the general-purpose large model and returns it to the central control module 101a. Specifically, when the third-party question-answering module 102 searches based on the received user intent, the search module 103a can search based on the dataset used to train the general-purpose large model. This search method does not require internet access; it only needs to select keywords based on the user intent and retrieve relevant data from the aforementioned dataset based on those keywords. Alternatively, the search module 103a can also use the internet to search for real-time data corresponding to the user intent. Optionally, the search engine used by the search module can be one or more search engines such as Baidu, Google, and Taobao, or other search engines, depending on the type of general-purpose large model corresponding to the third-party question-answering module 102. Furthermore, after retrieving data relevant to the user's intent, the overview module 103b in the third-party question-answering module 102 can use RAG technology to analyze and summarize all the content retrieved by the search module 103a, generating a corresponding overview summary so that the user can quickly obtain the full picture of the model's response. The snapshot module 103c generates corresponding content snapshots based on the web pages, posts, product information, etc., retrieved by the search module 103a. The prompt module 103d generates one or more related prompts corresponding to the search results based on the content retrieved by the search module 103a, so that the user can quickly engage in the next round of dialogue with the general model. In other words, the response returned by the third-party question-answering module 102 to the central control module can simultaneously include the aforementioned overview summary, content snapshot, and related prompts.
[0092] Finally, after obtaining the response content returned by the third-party question-and-answer module 102, the central control module 101a can combine the response with the graphics processor and display screen in the electronic device. Figure 1 (Not shown in the image), and after the hardware such as the application processor typeset and rendered the above response content, the resulting content is displayed in the user interface for the user to browse.
[0093] The following section describes the specific process of a user interacting with an electronic device based on the visual question-and-answer system 10, using a concrete user interface as an example.
[0094] like Figure 2 As shown in (A), the user interface 21 can be a chat interface provided by an instant messaging application in an electronic device, or it can be the application interface of other applications; this application does not limit this. Optionally, the user interface 21 can be an application... The provided chat interface, user interface 21 may include image 211.
[0095] The electronic device can respond to user operations on image 211, such as Figure 2 As shown in (A), the operation of holding down image 212 and dragging it to the right side generates image 212 in the interface and displays it as shown. Figure 2 User interface 22 is shown in (B) above. Without the user releasing their finger, image 212 can move within user interface 22 with the user's fingertip, and image 212 can be displayed on top of user interface 22, meaning image 212 can cover any element in user interface 22. It should be understood that the image content displayed by image 212 is the same as the image content displayed by image 211.
[0096] Optionally, the electronic device will only display the user interface 22 if the user drags the image 212 to a distance less than a certain threshold from the right or left edge of the screen. For example... Figure 2 As shown in (B) above, the user interface 22 is the arbitrary door interface mentioned in the foregoing description. The user interface 22 may include a sub-interface 221, an icon area 222, and an image 212. Wherein:
[0097] The elements displayed in sub-interface 221 are the same as those displayed in user interface 21. In fact, sub-interface 221 can be considered as the interface obtained by the electronic device after performing a three-dimensional transformation on all elements of the original user interface 21.
[0098] Icon area 222 can be used to display multiple application service icons recommended by the electronic device based on the user-selected image 212. Optionally, each icon in icon area 222 has a corresponding text description below it, which can be used to explain the specific application name or application service name corresponding to the icon. Among them, the application service icon 222a can be an icon of VQA service, which can be used to activate the VQA function provided by the electronic device.
[0099] Then as Figure 2 As shown in (B), in the user interface 22, after the user places the dragged image 212 on the application service icon 222a and releases it, the electronic device can respond to the user operation by transmitting the image 212 to the aforementioned processing module 101 and displaying it as shown in (B). Figure 2 User interface 23 is shown in (C).
[0100] like Figure 2As shown in (C), user interface 23 is the service interface corresponding to the VQA service provided by the electronic device. Users can interact with the electronic device on this interface and obtain the information they need through the device's responses. User interface 23 may include an image 212, a conversation history display area 231, a related prompt display area 232, an input box 233, and a send control 234, wherein:
[0101] Image 212 is the same as image 212 in the aforementioned user interface 22.
[0102] The dialogue record display area 231 can be used to display the dialogue records generated when the user and electronic device engage in visual question and answer.
[0103] The associated prompt display area 232 can be used to display associated prompts generated by the electronic device based on entity information in image 212. Any associated prompt can respond to user actions, such as clicks, allowing the electronic device to learn the user intent corresponding to that associated prompt. Understandably, when the user interface 23 is first displayed, the user has only entered image 212 and not any text. Therefore, at this time, the electronic device can only obtain entity information contained in image 212 through the aforementioned image semantic understanding module, but cannot accurately identify the user intent. Therefore, after the central control module 101a obtains the entity information in the image, it can generate one or more associated prompts based on the entity information. These associated prompts can then be displayed in the associated prompt display area 232, allowing the user to quickly input their intent through clicks. For example, if the entity information contained in image 212 is food, the associated prompts generated by the central control module 101a based on this entity may include, for example, food. Figure 2 (C) shows "What is this?", "Food Exploration", etc.
[0104] Of course, users can also choose not to click on any of the related prompts in the related prompt display area 232, but instead directly enter text information in the input box 233 and click the send control 234 to express their intention. For example... Figure 2 As shown in (C), after the electronic device displays the user interface 23, the user directly enters the text "Is this Nanjing cuisine?" in the input box 233 and clicks the send control 234. Correspondingly, the electronic device can respond to this operation and display as shown in (C). Figure 2 User interface 24 is shown in (D).
[0105] like Figure 2 As shown in (D), the user interface 24 may include a conversation history display area 231, an input box 233, a send control 234, an image 241, text 242, and reply content 243, wherein:
[0106] The dialogue history display area 231, input box 233, and send control 234 are the same as the dialogue history display area 231, input box 233, and send control 234 in the aforementioned user interface 23.
[0107] The image 241 displayed in the chat history display area 231 is the aforementioned image 212, and the text 242 is the text entered by the user into the input box 233. After the user sends control 234 in the user interface 23, the aforementioned image 212 and the text "Is this Nanjing cuisine?" in the user input box 233 can be displayed as chat history in the chat history display area 231. Correspondingly, the image semantic understanding module 101b and the text semantic understanding module 101c in the aforementioned processing module 101 may respectively perform recognition processing on image 241 and text 242, and the results of the two processing will be returned to the aforementioned central control module 101a, which will further process them to obtain the current user intent. Understandably, since the user's current intent is to ask whether the object in the current image is a local specialty food, which is not an intent that can be directly achieved through system tools, the central control module will further send the entity information extracted from the image (in this application embodiment, the food in image 241) and the user intent (in this application embodiment, "search whether the food in the image is a specialty food of Nanjing") together to the three-party general model (that is, the three-party question answering module 102 mentioned above). The general model combines the above user intent and entity information to perform a search, and generates a reply based on the search results and returns it to the central control module 101a.
[0108] The response content 243 refers to the response returned to the central control module 101a after the general model performs a search by combining the obtained user intent and image entity information. Specifically, the response content 243 may include an overview summary 243a, a content snapshot 243b, and related prompts 243c. The summary 243a is the text description obtained by the summary module 103b after summarizing and refining the content searched by the search module 103a. The content snapshot 243b is a snapshot generated by the snapshot module 103c based on the web pages searched by the search module 103a. Each snapshot can display part of the content of the corresponding web page (such as key titles and / or images), and each snapshot can respond to user operations, such as a click, so that the electronic device displays the full content of the web page corresponding to the snapshot. The related prompt 243c is the prompt information generated by the prompt module 103d based on the content searched by the search module 103a. This prompt information indicates the user's possible next intention. Any related prompt can respond to user operations, such as a click, so that the electronic device can obtain the user intention contained in the related prompt, so that the user can quickly start the next round of dialogue with the general model. Of course, the user can also continue to directly enter text information in the input box 233 and click the send control 234 to enter their intention again.
[0109] like Figure 2 As shown in (D), the user can click on the associated prompt 243c1 to cause the electronic device (central control module 101a) to receive the text "Which qingtuan in Nanjing is delicious?" contained in the associated prompt 243c1, and display it as shown in (D). Figure 2 User interface 25 is shown in (E) in the diagram. Accordingly, the user intent corresponding to the associated prompt is returned to the third-party general model. The third-party general model searches for the user intent and organizes the search results in a similar manner as described above. After obtaining the corresponding response content, it is returned to the central control module 101a, which then outputs the response content to the user interface 25 for the user to browse.
[0110] like Figure 2 As shown in (E), the user interface 25 may include text 251 and response content 252 from the electronic device to the user based on data obtained from the general big model. The text 251 is the text contained in the associated prompt 243c1; the response content 252 is the response content returned to the central control module 101a after the three-party general big model performs a search based on the obtained user intent and image entity information. Similarly, the response content 252 may include an overview summary 252a, a content snapshot 252b, and an associated prompt 252c. For details, please refer to the aforementioned description of the response content 243, which will not be repeated here.
[0111] Afterwards, the user can continue to communicate with the electronic device in the aforementioned manner until the conversation ends.
[0112] As can be seen from the foregoing explanation, after the central control module 101a understands the user's intent, regardless of the intent itself, the electronic device uses the same general model to search for that intent. The responses returned by the electronic device to the central control module are also obtained through the same search engine. Therefore, the content that these general models can reply with mainly comes from publicly available general knowledge bases on the internet. However, in reality, users may have different preferences for information sources in different scenarios. For example, when a user searches for food-related content, they might prefer to use consumer review applications (e.g., [example application]). Searching within e-commerce apps (such as e-commerce platforms) is common practice, while searching for clothing and styling-related content might be done through e-commerce apps. or Searching on JD.com or Taobao is possible, but two different users may have different information source preferences for the same type of question. For example, when searching for clothing in images, one user might prefer Xiaohongshu, while another might prefer JD.com. Therefore, when conducting VQA, consistently using a general, large-scale model for search responses may not provide users with the information they truly need. This is especially true when a user's intent falls within a specific vertical domain and they have a clear preference for information sources within that domain. The information the user desires may need to be obtained from a company's internal private knowledge base, and the general, large-scale model may not meet the user's needs. In other words, during the Q&A process, there is a high possibility of irrelevant answers or responses that do not match the user's actual needs. For example, suppose a user wants to search for content related to Nanjing Qingtuan shops on Xiaohongshu. However, as can be seen from the responses displayed in user interfaces 23 and 24, the snapshots provided by the general, large-scale model are all webpage snapshots and do not provide snapshots of notes shared by bloggers on Xiaohongshu. In this case, the electronic device may not be able to meet the user's true intent during the VQA process.
[0113] To address the shortcomings of the aforementioned VQA methods, this application provides an information processing method. Implementing this method, an electronic device can learn the user's usage preferences for applications (large language model) across various vertical domains based on embedded data collected during the user's historical use of the device. During visual question-and-answer sessions between the user and the electronic device, the device can select the user's preferred application (vertical domain large model) within the vertical domain to process the user's intent and output the response obtained from that application (vertical domain large model) to the user. In this way, the electronic device's responses based on user intent can maximize the satisfaction of user needs and improve the efficiency of the dialogue between the user and the electronic device.
[0114] Next, combine Figures 3-7 This section describes the specific process by which electronic devices implement the aforementioned information processing method. It should be noted beforehand that, because this method requires the electronic device to combine historical data collected during the user's use of the device to determine the user's application (model) preferences across various vertical domains, in specific usage scenarios, the electronic device can default to a specific model when it first leaves the factory. Figure 1 as well as Figure 2 The method shown (i.e., the default uniformly adopts a general large model) performs VQA with the user until the running time of the electronic device reaches a certain duration threshold (or the number of embedded data points collected by the electronic device reaches a certain quantity threshold). After that, the electronic device can use the information processing method provided in the embodiments of this application to implement VQA with the user.
[0115] Figure 3 An example is shown in the architecture diagram of a system that implements VQA based on a large vertical domain model.
[0116] like Figure 3 As shown, the visual question answering system 30 may include a processing module 301, a set of third-party large language models 302, and a system tool module 303. Wherein:
[0117] The processing module 301 may be a processor of an electronic device. The processing module 301 may include a central control module 301a, an image semantic understanding module 301b, a text semantic understanding module 301c, and a user preference understanding module 301d.
[0118] The third-party large language model set 302 may include multiple large language models. These large language models may include vertical domain large models provided by third-party applications installed on electronic devices. These vertical domain large models may be deployed on cloud servers provided by the service providers of the third-party applications, as well as general large models deployed on electronic devices or on the cloud servers of electronic devices. Each large language model may include one or more modules such as a search module, an overview module, a snapshot module, and a prompt module. For ease of explanation, it is assumed in the embodiments of this application and subsequent embodiments that the large language models in the language model set 302 all include a search module, an overview module, and a snapshot module.
[0119] The system tool module 303 may be a functional component provided by the operating system used by the electronic device; optionally, in some embodiments, the system tool module 303 may also be included in the processing module 301.
[0120] Specifically, during the process of a user interacting with an electronic device through the visual question-answering system 30, the processing module 301 can first obtain user input. This user input can be one or more of an image or text. Since the text input by the user is responsible for conveying the current user intent to the electronic device, in this embodiment, the processing module 301 determines the user intent differently depending on whether the user input includes both an image and text, or whether the user input only contains an image.
[0121] 1) When the user input includes both images and text, the processing module 301 determines the user's intent.
[0122] When user input includes both images and text, the image semantic understanding module 301b and the text semantic understanding module 301c will respectively recognize the image and text to determine the user's intent and the entities contained in the image. It's important to understand that the user intent determined by the text semantic understanding module 301c is only a vague intent. This vague intent needs to be combined with the entity information extracted from the image by the image semantic understanding module 301b to become a clear user intent. For example, suppose the input image contains the entity information of a dog, and the user inputs the text "What is this?". Understandably, based solely on the text "What is this?", we can only determine that the user's intent is to understand information about an object, but not what that object is. Therefore, we need to combine the entity information extracted from the image to understand the user's specific intent. During the search, the central control module 301a will combine the entity information extracted from the image with the vague user intent extracted from the text to obtain the specific user intent, and then send the user intent to the large language model for searching.
[0123] Figure 4 An example is shown where the central control module 301a combines entity information with vague user intent to obtain a specific user intent.
[0124] like Figure 4 As shown, entity information 401 is the entity information extracted from the image by the image semantic understanding module 301b, for example... Figure 4 The image shown includes entities such as address (add), number, person, clothing, food, etc. Understandably, each image input by the user may include one or more of these entities.
[0125] Semantic information 402 can be the fuzzy intent extracted by the text semantic understanding module 301c from the text input by the user. Specifically, the user can input text via keyboard or voice. This fuzzy intent can be determined by certain fields contained in the text. For example, when the user input text includes fields such as "navigation" or "how to get there," the text semantic understanding module 301c can determine that the current user intent may be to navigate to a certain location; when the user input text includes fields such as "call" or "call," the text semantic understanding module 301c can determine that the current user intent may be to call a certain contact; when the user input text includes fields such as "eliminate" or "remove," the text semantic understanding module 301c can determine that the current user intent may be to eliminate a certain entity in an image; when the user input text includes fields such as "buy" or "how to buy," the text semantic understanding module 301c can determine that the current user intent may be to buy a certain product; when the user input text includes fields such as "what is this," the text semantic understanding module 301c can determine that the current user intent may be to inquire about the relevant information of a certain item. Understandably, the content of each user intent field can have one or more other forms of representation, and this application does not limit this.
[0126] Correspondingly, after the text semantic understanding module 301c determines the fuzzy user intent, the text semantic understanding module 301c can pre-reserve a slot for each fuzzy intent. This slot can be used to accommodate one or more entities extracted from the user input image, so as to transform the fuzzy intent in the semantic information into a slotted user intent. Figure 4 Taking the slotted intent 403 as an example, when the user intent is to navigate to a certain place, the text semantic understanding module 301c can reserve a destination slot in the user intent. In this case, the fuzzy user intent can be represented by the text "navigate to (destination)," where "destination" is a pronoun and will be replaced by the address-type entity extracted from the image later. Similarly, when the user intent is to inquire about information about a certain entity in an image, the text semantic understanding module 301c can reserve an object slot in the user intent. In this case, the fuzzy user intent can be represented by the text "(the object in the image) is what," where "the object in the image" is also a pronoun and will be replaced by the entity extracted from the image later. Electronic devices can convert other fuzzy intents into slotted user intents in the same way; examples are not provided here. This is understandable, but not limited to... Figure 4 The expression used for the slotted user intent described in the document can also be expressed in other ways in electronic devices. For example, "navigate to (destination)" can also be expressed as "how to get to (destination)". This application does not limit the expression of the slotted user intent.
[0127] Subsequently, after obtaining the user intent with the slot and the entity information contained in the image, the central control module 301a can associate the user intent with the corresponding entity in the image to concatenate the two into a clear and specific user intent. Figure 4 Taking user intent 404 as an example, if the user intent with slot obtained by the central control module 301a is "navigate to (destination)", and the address-type entity information obtained by the central control module 301a from the image is "add" (add represents a specific address, which can exist in the image input by the user in the form of text), then the specific user intent obtained by the central control module 401a by concatenating the two is "navigate to add". As another example, if the user intent with slot obtained by the central control module 301a is "(object in the image)", and the clothing-type entity information obtained by the central control module 301a from the image is "obj3" (obj3 represents a specific entity information, which can exist in the image input by the central control module 301a in the form of an image, and when the entity information is extracted and sent to the large language model, the entity obj3 still exists in the form of an image), then the specific user intent obtained by the central control module 401a by concatenating the two is "how to buy obj3" (at this time, obj3 still exists in the form of an image). Similarly, for other types of entity information, such as... Figure 4 The central control module 301a can combine the entity information of the number type (tel), the entity information of the person type (obj1), the entity information of the object type (obj2), and other types of entity information with the acquired user intent with slots to obtain a clear and specific user intent. Examples will not be given here.
[0128] Understandably, even for the same type of entity, a user's intent regarding that entity may differ in different VQA processes. For example, suppose in this VQA session, the user's intent regarding an address-type entity in an image is to navigate to that address; however, in the next VQA session, the user's intent regarding the same address-type entity extracted from the image may change, no longer navigating to that address, but rather viewing surrounding information (such as nearby supermarkets, hospitals, parking lots, etc.). Therefore, after the central control module 301 determines the user's specific intent, it also records the correspondence between this user intent and the entity type. In this way, even when the user only inputs an image, the central control module 301 can obtain entity information from the user's input image and determine the type of that entity. Furthermore, the central control module 301 can determine one or more intents that the user frequently performs for the aforementioned entity types based on user habits, and output related prompts to the user based on these intents. The user can click on the related prompt to input their intent, and the central control module can then obtain the user intent corresponding to the clicked related prompt. Taking user behavior 405 as an example, assuming that the specific user intent obtained by the central control module 301a in this VQA based on the geological entity information add is "navigate to add", then the correspondence between entity information and user intent can be expressed as "entity class: address - user intent: navigation".
[0129] Understandably, the above correspondence between user intents and entities can be sent to... Figure 3 In the user preference understanding module 301d, the user preference understanding module 301d records and updates the information. For details, please refer to the subsequent sections. Figure 5 Related explanations.
[0130] 2) When the user input only contains an image, the processing module 301 determines the user's intention.
[0131] When the user input only contains an image, the image semantic understanding module 301b first identifies the entities contained in the image. Based on the obtained entities, it retrieves the frequency of various user intentions under that entity type from the user preference understanding module 301d. Based on the intentions that users frequently perform under the entity type, it generates corresponding related prompts for the user to choose from. When the user selects a certain related prompt, the ambiguous intention implied by the related prompt is concatenated with the aforementioned entities to obtain the specific user intention.
[0132] Figure 5 An example is shown where the central control module 301a combines entity information with vague user intent to obtain a specific user intent.
[0133] like Figure 5As shown, entity information 501 is the entity information extracted from the image by the image semantic understanding module 301b. User preference understanding module 301d records user habits 502, which specifically refers to the number of times the user has performed various user intentions for each type of entity during the user's historical use of electronic devices.
[0134] Since the user has not entered any text, the central control module 301a cannot determine the user's intent at this moment. However, the central control module 301a can combine the types of entity information obtained from the user's input image by the image semantic understanding module 301b, and the user habits 502 obtained from the user preference understanding module, to determine the total number of times the user has historically fulfilled various intents. Table 1 below exemplarily shows some of the user habits recorded in the user preference understanding module 301d:
[0135] Table 1 - Entity Types - Number of times each intent is implemented for each entity type
[0136]
[0137]
[0138] As can be seen from Table 1, the types of entity information defined by electronic devices can include, but are not limited to, address, number, human body, clothing, food and other entity types shown in Table 1. The common intentions to be realized for each type of entity can include, but are not limited to, text extraction, navigation, dialing, entity elimination, common sense search (including but not limited to asking what the entity is and what its function is), clothing search (including but not limited to asking what brand the entity is and what clothing it is paired with), and food search (including but not limited to asking what the entity is a specialty of and how it is made).
[0139] The central control module 301a can obtain the number of times the user has historically performed various intentions under that entity type from the user preference understanding module 301d, based on the type of entity contained in the user input image, and generate corresponding related prompts 503 based on the user intentions that have been performed more frequently. These related prompts can be displayed on the user interface for the user to select. If the user selects a certain related prompt, the central control module 301a can combine the user intention corresponding to the related prompt with the entity information 501 to obtain the user intention 504; understandably, the user intention 504 is a specific and clear user intention.
[0140] Let's take the example of an input image containing clothing-related entities, without any text input by the user. First, the image semantic understanding module 301b identifies clothing-related entities from the user's input image. Then, the central control module 301a obtains from the user preference understanding module 301d the number of times the user has historically performed various intents when the entity type is clothing. Based on Table 1, assuming that for clothing-related entities, the most frequently performed user intents are "clothing search" (30 times) and "entity elimination" (20 times). Therefore, even without any text input, the central control module 301a can combine the most frequently performed user intents for clothing-related entities and provide the user with related prompts on the interface, such as "Which brand is this clothing?", "Other clothing items to match this clothing?", and "Eliminate clothing in the image," for the user to choose from. After the user selects a related prompt, the central control module 301a can obtain the user intent corresponding to that prompt and combine it with the entity extracted from the image to obtain the specific user intent.
[0141] Correspondingly, the central control module 301a records the user intent for this instance of targeting the clothing type entity and updates the user preference information stored in the user preference understanding module 301d. For example, if the user selects the associated prompt "Which brand is the clothing product?", then the user intent for this instance of targeting the clothing type entity is "Which brand is the clothing product in the search image?". Accordingly, the number of times the user performs a "clothing search" will be updated from the original 30 times to 31 times in the entity information for clothing recorded in the user preference understanding module 301d. Optionally, in subsequent use, when the user launches a large language model or application to search for an image in any way, the central control module 301a will update the information in Table 1 above by combining the embedded data collected during the search (i.e., the entity type in the image and the user intent for the entity).
[0142] After determining the specific user intent, the central control module 301a will determine whether the acquired user intent can be implemented through system tools. If the user intent can be implemented solely through system tools, the central control module 301a will send the acquired user intent along with the entity information contained in the image to the system tool module 303. The system tool module 303 will then process the entities contained in the image using its provided system tools, and return the processing result to the central control module 301a. For details, please refer to the aforementioned related explanations; they will not be repeated here.
[0143] When a user's intent is not realized through system tools, the central control module 301a determines the vertical domain to which the user's intent belongs and, in conjunction with the user habits recorded in the user preference understanding module 301d, identifies the application (large language model) that the user habitually uses (which may be the most frequently used) within that vertical domain. Then, the central control module 301a can send the user's intent and the entity information extracted from the image to the application (large language model) that the user habitually uses, which then processes the user's intent.
[0144] Optionally, if a user prefers a large language model when expressing a user intent belonging to a certain vertical domain, the result returned to the central control module 301a after processing the user intent by the large language model may include an overview summary, content snapshot, and related prompts generated by the large language model based on the search results. If a user prefers an application when expressing a user intent belonging to a certain vertical domain, the application may process the user intent based on a certain application service it provides, and return the corresponding application service card as the processing result to the central control module 301a. If the user clicks on the application service card in the interface, the electronic device can directly jump to the application service interface. Optionally, the content displayed in the application service interface may include the content obtained after the application service processes the user intent. For example, when the user intent is to navigate to a certain location, the aforementioned application service card may be a navigation service card provided by a map application. After the user clicks on the application service card, the electronic device can directly display the navigation service interface, and the content displayed in the navigation service interface may include the route map and other content provided by the navigation service after searching for the aforementioned location (without requiring the user to enter the aforementioned location again in the navigation service interface and click the navigation control).
[0145] Specifically, the central control module 301a can determine the vertical domain to which the user intent belongs based on the keywords and entity information contained in the user intent. For example, when the user intent contains keywords such as "food" or "delicious" and the entity information contained is food, the central control module 301a can determine that the vertical domain to which the user intent belongs is "food"; when the user intent contains keywords such as "brand" or "store" and the entity contained is clothing, the central control module 301a can determine that the vertical domain to which the user intent belongs is "fashion"; when the user intent contains keywords such as "what is it" and the entity contained is clothing, the central control module 301a can determine that the vertical domain to which the user intent belongs is "general Q&A"; when the user intent contains keywords such as "cutout" or "remove," the central control module 301a can determine that the vertical domain to which the user intent belongs is "food." Further examples are not provided here.
[0146] After determining the vertical domain to which the user's intent belongs, the central control module 301a can identify the application (large language model) most frequently used by the user within that vertical domain from the third-party large language model set 302, and then call that application (large language model) to process the user's intent. In other words, in Figure 3 In this context, models model1, model2, model3, model4, and model5 can be the most commonly used large language models for users in the fashion, beauty, food, general knowledge, and general Q&A verticals, respectively. Optionally, if the user's intent explicitly points to a specific application (large language model), the central control module 301a can directly call that large language model to process the user's intent without needing to determine the vertical to which the intent belongs; for example, when the user inputs the text "use..." If the provided Da Vinci model searches for what the object in the image is, then the central control module 301a can directly determine the user's current preferred language model as the Da Vinci model by combining text semantics. In this case, the central control module 301a can directly call the Da Vinci model to process the user's intent.
[0147] Specifically, the central control module 301a can send the aforementioned specific user intent to the large language model. This specific user intent can include text expressing the user's intent (e.g., "What is the object in the image?") and entities extracted from the user's input image (e.g., clothing, food, human figures, etc.). Correspondingly, any one of the three large language models in the set 302 can search for images and / or text, and each large language model can include an overview module, a snapshot module, and a prompt module to generate an overview summary, a content snapshot, and related prompts based on the searched content, respectively. For details, please refer to the aforementioned... Figure 1 and Figure 2 The relevant explanations will not be repeated here. Next, the large language model responsible for processing user intent will return the above summary, content snapshot, and related prompts to the central control module 301a. Finally, the central control module 301a can combine the graphics processor and display screen in the electronic device... Figure 3 The above summary, content snapshots, and related prompts are typed and rendered by hardware such as the application processor (as shown in the image), and the resulting content is displayed in the user interface for the user to browse.
[0148] List 2 below illustrates, for example, some user habits recorded in the user preference understanding module 301d:
[0149] Table 2: Intent Verticals - Number of times each application (model) is used within each intent vertical.
[0150]
[0151] As can be seen from Table 2, the vertical domain to which the user intent belongs may include, but is not limited to, the fashion vertical domain, beauty vertical domain, food vertical domain, life knowledge vertical domain, and general Q&A vertical domain shown in Table 2. The application (model) used to process the user intent may include, but is not limited to, the social application 1-model 1, the consumer review application 2-model 2, and the e-commerce application 3-model 3 shown in Table 2.
[0152] Based on Table 2 and the aforementioned explanations, we assume that in the fashion category, the most frequently used broad language model is social application 1 - Model 1 (38 times), and in the food category, the most frequently used broad language model is consumer review application 2 - Model 2. Based on these assumptions, we will then combine... Figure 6 and Figure 7 The user interface involved in VQA is described separately for users and electronic devices.
[0153] Figure 6 This example illustrates the specific process of a user conducting VQA with an electronic device when the user's intent pertains to the food industry.
[0154] like Figure 6 As shown in (A), the user interface 61 can be any door interface mentioned in the foregoing description, which may include a sub-interface 611, an icon area 612, and an image 613. Wherein:
[0155] Sub-interface 611 can be an electronic device that displays the previous user interface ( Figure 6 (Not shown in the image; image 613 is the interface obtained after a 3D transformation of the user interface.)
[0156] The icon area 612 can be used to display multiple application service icons recommended by the electronic device based on the user-selected image 613. Optionally, each icon in the icon area 612 has a corresponding text description below it, which can be used to explain the specific application name or application service name corresponding to the icon. Among them, the application service icon 612a can be an icon of VQA service, which can be used to activate the VQA function provided by the electronic device.
[0157] In user interface 61, after the user places the dragged image 613 onto the application service icon 612a and releases it, the electronic device responds to the user's operation and displays, as shown below. Figure 6 User interface 62 is shown in (B) of the diagram.
[0158] like Figure 6As shown in (B) in [reference], the user interface 62 is the service interface corresponding to the VQA service provided by the electronic device. The user can conduct visual question-and-answer with the electronic device in the user interface 62 and obtain the information they need from the reply content provided by the electronic device. The user interface 62 may include an image 613, a conversation record display area 621, a general prompt display area 622, an input box 623, and a send control 624, where:
[0159] The image 613 is the image 613 in the aforementioned user interface 61, and the entity contained therein is the entity "Qingtuan" of the food type.
[0160] The conversation record display area 621 can be used to display the conversation records generated when the user and the electronic device conduct visual question-and-answer. It can be understood that at this time, the user has not officially conversed with the electronic device, so there are no conversation records in the conversation record display area 621 yet.
[0161] The general prompt display area 622 can be used to display the general prompts provided by the electronic device. Any one of the general prompts can respond to a user operation, such as a click operation, to enable the electronic device to obtain the user intention corresponding to the general prompt. It should be noted that the above general prompts may not be the associated prompts recommended by the central control module 301a in combination with the user habits and the entities contained in the image 613, but the general prompts preset by the central control module 301. Optionally, regardless of whether the user inputs an image and the type of the entity contained in the image, when just entering the VQA service interface, the electronic device can display one or more general prompts in the general prompt area 622 for the user to quickly input their user intention through a click operation. Of course, the user can also not click on any of the general prompts in the general prompt display area 622, but directly enter text information in the input box 623 and then click the send control 624 to input their intention.
[0162] Such as Figure 6As shown in (B), after the electronic device displays the user interface 62, the user directly enters the text "Is this a Nanjing delicacy?" in the input box 623 and clicks the send control 624. In response to this operation, the electronic device transmits image 613 to the image semantic understanding module 301b and the text "Is this a Nanjing delicacy?" to the text semantic understanding module 301c. The image semantic understanding module 301b extracts the entity information "Qingtuan" from the image, and the text semantic understanding module 301c extracts the fuzzy intent "Is the subject in the image a Nanjing delicacy?", which is then transmitted to the central control module 301a. The central control module 301a processes both to obtain the specific user intent "Is Qingtuan a Nanjing delicacy?". Next, the central control module 301a determines that the user intent "Is searching for qingtuan (a type of glutinous rice dumpling) a Nanjing delicacy?" belongs to the food vertical domain, and identifies application 2 (model 2) as the most frequently used application (large language model) within the food vertical domain. Accordingly, the central control module 301a transmits the user intent to model 2 provided by application 2. Model 2 then searches for and responds to the user intent. After obtaining the response from model 2, the central control module 301a finally outputs the response to the user; the corresponding electronic device can then display the response as shown below. Figure 6 The user interface 63 is shown in (C) in the diagram. Optionally, application 2 can be... Model 2 can be The provided orange model;
[0163] like Figure 6 As shown in (C), the user interface 63 may include a conversation history display area 621, an input box 623, a send control 624, an image 631, text 632, and reply content 633, wherein:
[0164] The dialogue history display area 621, input box 623, and send control 624 are the same as the dialogue history display area 621, input box 623, and send control 624 in the aforementioned user interface 22.
[0165] The image 631 displayed in the chat history display area 621 is the aforementioned image 613, and the text 632 is the text that the user enters into the input box 623 in the user interface 62. After the user sends control 624 in the user interface 62, the aforementioned image 613 and the text "Is this Nanjing cuisine?" in the user input box 623 can be displayed as chat history in the chat history display area 621.
[0166] The response content 633 refers to the response content returned to the central control module 301a after the consumer review application 2 (Model 2) performs a search based on the obtained user intent. Specifically, the response content 633 may include an overview summary 633a, a content snapshot 633b, and related prompts 633c. The specific functions of the overview summary 633a, content snapshot 633b, and related prompts 633c can be found in the aforementioned section. Figure 2 The relevant explanations will not be repeated here. It should be noted that, unlike the general prompts contained in the general prompt display area in the user interface 62, any one of the related prompts contained in the related prompts 633c is recommended to the user by Model 2 provided by the consumer review application 2 based on the entity information contained in the aforementioned user intent.
[0167] Optionally, the user can quickly initiate the next round of dialogue with the electronic device by clicking any of the user intents in the associated prompts 633. Alternatively, the user can continue to enter text information again in the input box 623 of the user interface 63 and click the send control 624 to re-enter their intent. Optionally, in a subsequent round of dialogue, if the central control module 301a determines that the vertical domain to which the user intent belongs has changed (e.g., from the food vertical domain to the fashion vertical domain), the central control module 301a can redetermine the application of the user's preferences (large language model) under the new vertical domain in that round of dialogue and process the user intent in that round of dialogue using the application of the user's preferences (large language model) under the new vertical domain.
[0168] Figure 6 (D) in the diagram shows the user interface displayed by the electronic device after multiple rounds of VQA dialogue between the user and the electronic device. Figure 6 As shown in (D), the user interface 64 includes text 641, response content 642, text 643, and response content 644. Here, it is defined that the user inputting image 631 and text 632, and the electronic device outputting response content 633, constitutes the first round of dialogue. Correspondingly, after the electronic device displays the user interface 63, if the user continues to input text content 641 and the electronic device outputs response content 642, this constitutes the second round of dialogue. If the user inputs text 643 and the electronic device outputs response content 644, this constitutes the third round of dialogue. The electronic device only displays the user interface 64 at the end of the third round of dialogue.
[0169] During the second round of dialogue between the user and the electronic device, the text 632 entered by the user contains the phrase "Which qingtuan (a type of glutinous rice dumpling) is delicious in Nanjing?" Based on the aforementioned explanation, the central control module 301a can infer from this text content that the current user intent (i.e., "searching for which qingtuan is delicious in Nanjing") belongs to the food vertical. Therefore, the electronic device can again call application 2 (model 2) to process the aforementioned user intent, obtain the data returned by model 2 based on the user intent, organize it into response content 642, and output it to the user for viewing. Understandably, if the central control module 301a determines that the user intent in this round of dialogue belongs to the fashion vertical rather than the food vertical, the electronic device can no longer call model 2 provided by the consumer review application 2 to process the user intent, but instead call model 2 provided by the social application 2 to process the user intent.
[0170] Furthermore, during the user's VQA process, the electronic device can review information from previous rounds of dialogue in each round and combine this information with the images and / or text input by the user in this round to determine the user's intent and respond accordingly. For example, in the third round of dialogue between the user and the electronic device, the user inputs text 643 containing the phrase "Choose the first one." Understandably, with only text 643 stored, the central control module 301a cannot understand which store the field "first one" in text 643 refers to. However, by combining this with the response content 642 output by the electronic device in the second round of dialogue, which contains specific information about multiple stores recommended by Model 2 for the user, the central control module 301a can determine that the aforementioned "first one" refers to "xxx cake shop."
[0171] It should also be noted that in some embodiments, if the central control module 301a determines that the current user intent no longer requires searching using a third-party large language model, then the central control module 301a may not call the third-party large language model for searching, but instead recommend various application services available to the user based on the user intent. For example, in the third round of dialogue, if the user inputs "Let's choose the first one," the user's emotion in this text is not "question" but "decision," and the relevant information about "xxx cake shop" has already been output to the user in the second round of dialogue, then the central control module 301a can combine the information of "xxx cake shop" to output a response content 644 containing multiple application service cards to the user. Specifically, since "xxx cake shop" is the shop name, the application services corresponding to the application service cards in the response content 644 may include instant messaging applications. Queue management services, coupon viewing services offered by consumer review apps, and travel navigation apps... The ride-hailing service and consumer review platform provided The provided store sharing service, etc.; for example, the application services corresponding to the application service cards included in response content 644 can be respectively The queuing service provided The coupon viewing service is provided. The ride-hailing service provided and Services such as store sharing are provided.
[0172] Understandably, the types of application services recommended by the central control module 301a to the user based on the user's intent may vary depending on the vertical domain to which the user's intent belongs. For example, in the beauty vertical domain, in the final round of dialogue, the application services corresponding to the application service cards provided by the central control module 301a may include store viewing services provided by e-commerce applications, product viewing services provided by e-commerce applications, note viewing services provided by social applications, etc. This application does not limit this.
[0173] In this embodiment, image 613 and text 623 can be collectively referred to as "first user input". The user intent "searching whether qingtuan is a Nanjing delicacy" can be referred to as "first user intent", the food vertical domain to which "searching whether qingtuan is a Nanjing delicacy" belongs can be referred to as "first vertical domain", the model 2 provided by application 2 can be referred to as "first model", and the response content 633 can be referred to as "first processing result". The overview summary 633a, content snapshot 633b, and related prompt 633c can be referred to as "first overview summary", "at least one first content snapshot", and "at least one first related prompt", respectively.
[0174] In this embodiment, image 613 may be referred to as "first image" and text 623 may be referred to as "first text". Entity information 401 may be referred to as first entity information and slotted intent 403 may be referred to as "second user intent".
[0175] Figure 7 This example illustrates the specific process of a user conducting VQA (Vehicle Quality Assurance) with an electronic device when the user's intent pertains to the fashion category.
[0176] like Figure 7 As shown in (A), user interface 71 is the service interface corresponding to the VQA service provided by the electronic device. Users can engage in visual question-and-answer sessions with the electronic device in user interface 71 and obtain the information they need from the responses provided by the electronic device. User interface 71 may include an image 711, a dialogue history display area 712, a general related prompt display area 713, an input box 714, and a send control 715, wherein:
[0177] Image 711 is the image that is subsequently sent to the processing module 30 (image semantic understanding module 301b) as user input. The entity contained in image 711 is the clothing type entity "handbag".
[0178] The dialogue log display area 712 can be used to display the dialogue log generated during the visual question-and-answer session between the user and the electronic device. Understandably, at this time, the user has not yet formally engaged in a dialogue with the electronic device, so there is no dialogue log in the dialogue log display area 712 yet.
[0179] The general association prompt display area 713 can be used to display general association prompts provided by the electronic device, where any one of the association prompts can respond to user actions, such as clicks, enabling the electronic device to obtain the user intent corresponding to that association prompt. For details, please refer to the aforementioned... Figure 6 The relevant explanations for the Zhongtong General prompt display area 622 will not be repeated here.
[0180] As explained above, user input can consist of only images. For example... Figure 6 As shown in (A), after the electronic device displays the user interface 72, the user does not click on any of the general association prompts in the general association prompt display area 713, nor does the user directly enter any text in the input box 714. Instead, the user directly clicks the send control 724. The electronic device can then respond to this operation by transmitting the image 711 to the image semantic understanding module 301b. The image semantic understanding module 301b extracts the entity information "handbag" from the image, which is then transmitted to the central control module 301a. The central control module 301a recognizes that the entity information "handbag" is a clothing type entity. The central control module 301a can then combine this with the user habits recorded in the user preference understanding module 301d to obtain one or more of the user's most frequently used intentions regarding clothing-related entities. Here, we assume that the two most frequently used intentions are extracted, and these two intentions... Figure 1 One is "entity removal," and the other is "brand search." Therefore, the central control module 301a will generate two corresponding related prompts based on the two most common user intentions regarding clothing-related entities, and output the entity extraction results and these two related prompts so that the electronic device can display them as shown. Figure 7 User interface 72 is shown in (B) of the diagram.
[0181] like Figure 7 As shown in (B), the user interface 72 may include a conversation history display area 712, an input box 714, a send control 715, an image 721, image description information 722, and related prompts 723, wherein:
[0182] The dialog history display area 712, input box 714, and send control 715 are the dialog history display area 712, input box 714, and send control 715 in the user interface 72.
[0183] Image 721 is the same as the aforementioned image 711. At this time, image 721 is displayed as a history dialogue record in the dialogue record display area 712.
[0184] Image description information 722 is descriptive text generated by the central control module 301a for the entity types contained in image 721. Image description information 722 can be used to prompt the user to input brief information about the entities contained in the image.
[0185] The associated prompt information 723 refers to the associated prompts generated by the central control module 301a based on the entity types contained in the image 721 and the user's common intents regarding those entity types. This includes associated prompt 723a and associated prompt 723b. Associated prompt 723a is generated by the central control module 301a based on the intent "query brand," and may contain the text "What brand is this bag?". Associated prompt 723b is generated by the central control module 301a based on the intent "eliminate entity," and may contain the text "Eliminate the handbag in the image." Any associated prompt can respond to a user action, such as a click, enabling the central control module 301a to obtain the intent corresponding to that associated prompt and combine it with the entity information extracted from the image 721 to arrive at the accurate user intent. Optionally, the user can also enter text expressing their intention in the input box 714 again and send it to the control 715, so that the central control module 301a can learn the user's current intention from the input text and combine it with the entity information extracted from the image 721 to obtain the accurate user intention.
[0186] like Figure 7As shown in (B), in response to the user's operation on the associated prompt 723a, the central control module 301a can obtain the intent "query brand" corresponding to the associated prompt 723a, and combine it with the aforementioned entity information "handbag" to obtain the specific user intent "query the brand of the handbag in the picture". As can be seen from the foregoing description, after obtaining the specific user intent, the central control module 301a can determine that the vertical domain to which the user intent belongs is the fashion vertical domain. Furthermore, based on the user habits obtained from the user preference understanding module 301d, the central control module 301a can determine that, when the user intent belongs to the fashion vertical domain, the application (large language model) most frequently used by the user is Model 1 provided by the social application 1. Accordingly, the central control module 301a will transmit the aforementioned user intent to Model 1 provided by the social application 1, where Model 1 will search and respond to the user intent. After obtaining the response content from the Da Vinci model, the central control module 301a will finally output the response content to the user; the corresponding electronic device can display as shown in the image. Figure 7 User interface 73 is shown in (C).
[0187] like Figure 7 As shown in (C), the user interface 73 may include a conversation history display area 712, an input box 714, a send control 715, text 731, and reply content 732, wherein:
[0188] The dialogue history display area 712, input box 715, and send control 724 are the same as the dialogue history display area 712, input box 715, and send control 714 in the aforementioned user interface 22.
[0189] The text 731 displayed in the chat history display area 712 is the same as the text "What brand is this bag?" contained in the aforementioned associated prompt 723a. After the user sends control 724 in the user interface 72, the text "What brand is this bag?" contained in the associated prompt 723a can be displayed as chat history in the chat history display area 712.
[0190] The response content 732 refers to the response content returned to the central control module 301a by Model 1, which provides the social application 1 after performing a search based on the obtained user intent. Specifically, the response content 732 may include an overview summary 732a, a content snapshot 732b, and related prompts 732c. The specific functions of the overview summary 732a, content snapshot 732b, and related prompts 732c can be found in the aforementioned section. Figure 2 The relevant explanations will not be repeated here. Similarly, unlike the general prompts contained in the general prompt display area in user interface 71, any one of the related prompts contained in related prompts 732c is recommended to the user by the Da Vinci model based on the entity information contained in the user intent mentioned above.
[0191] Optionally, the user can quickly initiate the next round of dialogue with the electronic device by clicking any user intent in the associated prompt 732c. Alternatively, the user can continue to enter text information again in the input box 714 of the user interface 73 and click the send control 724 to re-enter their intent. Optionally, in a subsequent round of dialogue, if the vertical domain to which the user intent of the central control module 301a belongs changes (e.g., from the fashion vertical domain to the beauty vertical domain), the central control module 301a can redetermine the application (large language model) of user preferences under the new vertical domain in that round of dialogue and process the user intent in that round of dialogue using the application (large language model) of user preferences under the new vertical domain. Examples will not be provided here.
[0192] In this embodiment, image 721 can be referred to as "first user input". The user intent "to query the brand of the handbag in the image" can be referred to as "first user intent", the clothing vertical domain to which "to query the brand of the handbag in the image" belongs can be referred to as "first vertical domain", model 1 provided by application 1 can be referred to as "first model", and the response content 732 can be referred to as "first processing result". The overview summary 732a, content snapshot 732b, and related prompt 732c can be referred to as "first overview summary", "at least one first content snapshot", and "at least one first related prompt", respectively.
[0193] In this embodiment, image 721 may be referred to as "second image". Entity information 501 may be referred to as second entity information, association prompt 503 may be referred to as "at least one second association prompt", and user intent 504 may be referred to as "at least one third user intent".
[0194] Understandably, some user intents are divisible. This divisibility manifests in the fact that, in order to achieve a specific user intent, a user often needs to simultaneously fulfill multiple sub-user intents. Please refer to [reference needed] for details. Figure 8 .
[0195] like Figure 8 As shown in (A), assuming a user currently has a user intent of "visiting the Forbidden City," then to achieve this intent, the user may need to fulfill multiple sub-user intents, including but not limited to: learning about Forbidden City information (e.g., its location, opening hours, and whether tickets are required), checking nearby restaurants, creating travel plans / guides, booking train tickets / hotels, and creating schedules / alarms for the travel plan. These sub-user intents belong to various verticals, such as general knowledge, food, and travel, and the user preference models and applications differ depending on the specific vertical.
[0196] However, when users utilize Figure 3 In the process of VQA between the visual question answering system and the electronic device, when the electronic device obtains a user intent that can be divided based on the user input, it does not split the user intent into multiple sub-user intents. Instead, it directly classifies the user intent into a certain vertical domain and, in conjunction with the user's preferred large language model or application for that vertical domain, sends the user intent to the user's preferred large language model or application for processing. After obtaining the processing result of the large language model or application on the user intent, the processing result is output to the user as the answer content.
[0197] by Figure 8 Taking the user interface 81 shown in (B) as an example, the user interface 81 can be the user interface displayed by the electronic device when the electronic user conducts VQA with the electronic device. At this time, the user has completed the first round of dialogue with the electronic device. The user interface 81 can include an image 811, text 812, and response content 813. The image 811 and text 812 are the image and text input by the user. The response content 813 is the content that the electronic device replies to the user based on the user's intent after determining the user's intent based on the image 811 and text 812.
[0198] As described above, after the user sends image 811 and text 812 in the dialogue interface, the electronic device can respond to the operation by transmitting image 811 to the image semantic understanding module 301b and transmitting the text "How to travel to this place" to the text semantic understanding module 301c. The central control module 301a will combine the entity information "Forbidden City" extracted from the image by the image semantic understanding module 301b and the fuzzy intent "How to travel to the place in the picture?" extracted by the text semantic understanding module 301c to determine that the user's intent is "Search for how to travel to the Forbidden City". Next, the central control module 301a determines that the user intent "search how to travel to the Forbidden City" belongs to the tourism vertical domain, and identifies the most frequently used application (large language model) within the tourism vertical domain as Model 1 provided by Application 1. Accordingly, the central control module 301a transmits the user intent to Model 1 provided by Application 1, where Model 1 searches for and responds to the user intent. After obtaining the response content 813 from Model 1, the central control module 301a finally outputs the response content 813 to the user, and the corresponding electronic device can display as shown below. Figure 8 User interface 81 is shown in (B) of the diagram.
[0199] However, the specific information contained in response content 813 shows that when searching for the intent to "search for ways to travel to the Forbidden City" based solely on Model 1, the electronic device only obtains and outputs some rough travel plans and travel-related posts. It lacks crucial information and services that users might need to achieve "traveling to the Forbidden City." For example, users need to know the Forbidden City's opening hours, combine the weather and their own itinerary to create a travel plan, determine their mode of transportation, purchase tickets to the Forbidden City, and book hotels. The information and services required to perform these operations are not present in response content 813. This is because the central control module 301, in understanding the user intent to "travel to the Forbidden City," did not combine user information and the actual travel process to understand and break down the user intent into multiple sub-user intents. This resulted in the user intent being limited to a single vertical domain, leading to a single source of information for the electronic device and a limited range of application services. While users can also request the electronic device to combine with one or more other applications (large language model) to obtain the information or services they need and output them to the user in subsequent rounds of dialogue, this undoubtedly reduces the efficiency of VQA between the user and the electronic device.
[0200] Therefore, based on the shortcomings of the aforementioned visual question answering systems, this application also provides another visual question answering system. This system can decompose user intents, and the resulting multiple sub-user intents can belong to different vertical domains. The electronic device can, according to the user's application (model) usage preferences for each vertical domain, send these multiple sub-user intents to corresponding third-party applications (large language models) for processing, and obtain the processing results of the multiple applications (large language models) on the sub-user intents. Then, the local large model in the electronic device can summarize and outline the processing results returned by the multiple third-party applications (large language models), generate corresponding related prompts based on the processing results returned by the multiple third-party applications (large language models), and output them to the user. In this way, the information in the response can be richer, the response content can better cover the user's actual needs, and improve the VQA efficiency between the user and the electronic device.
[0201] Figure 9 An example of the architecture diagram of the above-described visual question answering system is shown.
[0202] like Figure 9 As shown, the visual question-answering system 92 may include an electronic device 90 (device side) and a cloud server 91 (cloud side). The electronic device may include a processing module 901 and a local application set 902, and the cloud server 91 may include a large language model 911 and a local large model 912. Wherein:
[0203] The processing module 901 can be a processor of an electronic device. The processing module 901 can include a central control module 901a, an image semantic understanding module 901b, a text semantic understanding module 901c, and a user preference understanding module 901d.
[0204] The local application collection 902 may include third-party applications or system applications installed on electronic devices. These system applications may include system tool modules, which may be functional components provided by the operating system used by the electronic device.
[0205] The large language model 911 can be a large language model provided by a third-party application in the electronic device, while the local large model 912 can be a large model trained by the manufacturer of the electronic device. It can include an overview module and a prompt module. After obtaining search results from multiple third-party models for different user intentions (i.e., multi-party processing results), the overview module can summarize and refine the multi-party processing results and generate a corresponding overview. The prompt module can predict the user's potential intentions in the next round of dialogue based on the multi-party processing results and generate corresponding related prompts based on these intentions, enabling the user to quickly engage in the next round of dialogue with the electronic device through these related prompts.
[0206] Model / application set 902 may include multiple large language models and multiple applications installed on electronic devices. The large language models in model / application set 902 may include vertical domain large models provided by third-party applications installed on electronic devices (generally deployed on cloud servers provided by the service provider of the third-party application), and general large models deployed on electronic devices or on the cloud servers of electronic devices. Any one of these large language models may contain a search module for searching user intents. The applications in model / application set 902 may be local applications installed on electronic devices or third-party applications, which provide application services to realize user intents.
[0207] Specifically, during the process of a user interacting with an electronic device through the visual question-answering system 92, the processing module 901 can first obtain user input. This user input can be one or more of an image or text. Understandably, in this embodiment, the processing module 901 determines the user's intent and breaks it down differently depending on whether the user input includes both images and text or only an image.
[0208] 1) When the user input contains both images and text, the processing module 901 determines the user intent and breaks down the user intent.
[0209] Figure 10An example is shown where, when user input includes both images and text, processing module 901 determines and breaks down the user's intent.
[0210] like Figure 4 As shown, entity information 1001 is the entity information extracted from the image by the image semantic understanding module 901b. Semantic information 1002 can be the fuzzy intent extracted from the user-input text by the text semantic understanding module 901c. Correspondingly, after the text semantic understanding module 301c determines the fuzzy user intent, the fuzzy intent in the semantic information is converted into a slotted user intent 1003. The central control module 901a can associate the slotted user intent with the corresponding entity in the image to concatenate the two into a clear and specific user intent 1004. This process can be referred to in detail above. Figure 3 The relevant explanations will not be repeated here.
[0211] After obtaining a clear user intent 1004, the electronic device can decompose the aforementioned user intent by combining the user habits learned from the user preference understanding module 901d.
[0212] It should be noted in advance that, in this embodiment, the user habits 1006 obtained from the user preference understanding module 901d include not only the user's common user intentions for various types of entity data and the user's preferences for applications (large language models) in various vertical domains, but also the contextual relationships generated when the user uses electronic devices to realize user intentions in the past, such as the time and location of the electronic device processing various user intentions. In other words, during the user's historical use of electronic devices, the collected data points each time the user uses a certain large model or application to realize each user intention can include the time when the electronic device processes the user intention, the location of the electronic device, and another or more user intentions processed by the electronic device before processing the user intention. This information can be used as reference information when the central control module 901a decomposes user intentions. For example, assuming that the user first uses the Da Vinci model of Xiaohongshu to process the user intention of "searching for travel guides to a certain place", and then the user uses the Xiaochengzi model provided by Dianping to process the user intention of "searching for food near a certain place", it means that the user intention of "searching for travel guides" and the user intention of "searching for food" usually occur together. For example, suppose a user frequently uses a ride-hailing service provided by a map application to take a taxi home at a certain time and in a certain location (e.g., 6 pm near the company). This means that the user's usual intent at that location is "take a taxi to add" (add is the home address).
[0213] Therefore, after obtaining the explicit user intent 1004, the user can combine the user habits recorded in the user habits record 1005 with one or more pieces of information such as the user's location and current time to decompose the user intent 1004 and obtain the user intent list 1005. Understandably, the user intent list 1005 can contain multiple sub-user intents, for example... Figure 10 The sub-user intents 1005a, 1005c, 1005d, and 1005e are user intents that the electronic device may need to implement in order to realize user intent 1004, and the vertical domains to which each sub-user intent belongs may not be completely the same.
[0214] Subsequently, the central control module 901a can determine the vertical domain to which the user intent list 1005 may contain multiple sub-user intents, and collect user usage preferences for each vertical domain's application (large language model) from the user habits 1006. Further, the central control module 901a can send the multiple sub-user intents in the user intent list 1005 to the corresponding third-party large language models or applications for processing, and obtain the processing results of the sub-user intents from multiple applications (large language models). Optionally, the processing results of each application (large language model) for the received sub-user intents may include a snapshot of the content obtained from the user intent search.
[0215] by Figure 9 For example, the sub-user intents and the vertical domains to which each sub-user intent belongs, as decomposed by the central control module 901a, include sub-user intent int1 - vertical domain dom1, sub-user intent int2 - vertical domain dom2, sub-user intent int3 - vertical domain dom3, sub-user intent int4 - vertical domain dom4, and sub-user intent int5 - vertical domain dom5. The models preferred by the user in each vertical domain are vertical domain dom1 - model1, vertical domain dom2 - model2, vertical domain dom3 - model3, vertical domain dom4 - application APP1, and vertical domain dom5 - application APP2, respectively. Therefore, sub-user intent int1 will be sent to model1 for processing, sub-user intent int2 will be sent to model2 for processing, and so on.
[0216] Specifically, the processing results of each large language model for the corresponding sub-user intent are sent to the local large model 912; while the application services provided by each local application based on the corresponding sub-user intent are returned to the central control module 901a in the form of application service cards. Next, the overview module in the local large model 912 can summarize the processing results of multiple parties (multiple large language models) to obtain an overview summary, while the prompt module can generate corresponding related prompts based on the processing results of multiple parties.
[0217] Finally, the summary returned by the local large model 912, the summary of the content obtained from the search of each large model mentioned above, and the application service cards returned by the local application will be used as result set 1008 (i.e. Figure 9 The response content is returned to the central control module 901a along with the response content. The central control module 901a can be combined with modules such as the GPU, display screen, and application processor in electronic devices. Figure 9 (Not shown in the image) This information is displayed on the screen for the user to view. Optionally, the local large model 912 can also select snapshots of the content returned by various language models from the multi-processing results and output them to the user along with the above information.
[0218] Understandably, users may not use all the content output by the electronic device, but may only select and browse a portion of the content last output by the electronic device (including service cards, content snapshots, and associated prompts, all of which correspond to specific user intentions). Therefore, in some embodiments, the central control module 901a will also update the user habits recorded in the user preference understanding module 901d based on the user intention corresponding to the content actually selected by the user. In this way, during subsequent dialogues, when decomposing similar user intentions, the central control module 901a can more likely decompose the sub-user intentions that the user actually needs to achieve.
[0219] 2) When the user input only contains an image, the processing module 901 determines the user intent and breaks down the user intent in a certain way.
[0220] Figure 11 An example is shown where the processing module 901 determines and breaks down the user intent when the user input only contains an image.
[0221] When the user input only contains an image, the image semantic understanding module 901b first identifies the image to determine the entity information 1101 contained in the image. Based on the obtained entities, it retrieves the user habits 1102 from the user preference understanding module 901d. The user habits 1102 record the number of times various user intentions are realized under the entity type to which the entity information 1101 belongs. Then, based on the user's frequently realized intentions for the entity type to which the entity information 1101 belongs, the central control module 301a can generate corresponding related prompts 1103 for the user to choose from. When the user selects a certain related prompt, the fuzzy intention implied by the related prompt is concatenated with the aforementioned entity information to obtain the specific user intention 1104.
[0222] After obtaining a clear user intent 1004, the electronic device can decompose the aforementioned user intent by combining the user habits 1102 learned from the user preference understanding module 901d.
[0223] Similarly, user habits 1102, in addition to including user intentions commonly used by users for various entity data and user preferences for applications (large language models) in various vertical domains, can also include the contextual relationships of user intentions generated by electronic devices in the user's historical use of electronic devices, such as the time and location of the electronic device processing various user intentions. Therefore, after obtaining the aforementioned explicit user intention 1004, the user can combine the user habits recorded in user habits 1101 with one or more pieces of information such as the user's location and current time to decompose user intention 1004, obtaining a user intention list 1105 including sub-user intentions 1105a-1105e. Then, the central control module 901a can determine the vertical domain to which multiple sub-user intentions may belong in the user intention list 1005, and collect the user's application (large language model) usage preferences for each vertical domain from user habits 1106. Furthermore, the central control module 901a can send multiple sub-user intents from the user intent list 1105 to the corresponding third-party large language models or applications for processing. Specifically, the processing results of each large language model for the corresponding sub-user intent are sent to the local large model 912; while the application services provided by each local application based on the corresponding sub-user intent are returned to the central control module 901a in the form of application service cards. Next, the overview module in the local large model 912 can summarize the processing results from multiple parties (i.e., multiple large language models) to obtain an overview summary, while the prompt module can generate corresponding related prompts based on the processing results from multiple parties.
[0224] Finally, the summary returned by the local large model 912, the summary of the content obtained from the search of each large model mentioned above, and the application service card returned by the local application will be used as result set 1107 (i.e. Figure 9 The response content is returned to the central control module 901a along with the response content. The central control module 901a can be combined with modules such as the GPU, display screen, and application processor in electronic devices. Figure 9 (Not shown in the image) This information is displayed on the screen for the user to view. Optionally, the local large model 912 can also select snapshots of the content returned by various language models from the multi-processing results and output them to the user along with the above information. For details, please refer to the aforementioned... Figure 10 The relevant explanations will not be repeated here.
[0225] Figure 12 An example shows an electronic device based on Figure 9 The visual question-answering system shown here interacts with the user through a user interface designed for VQA.
[0226] like Figure 12 As shown, user interface 12 can be the user interface displayed by the electronic device when the electronic user engages in VQA (Voice-to-User Questions) with the electronic device. At this point, the user has completed the first round of dialogue with the electronic device. User interface 12 can include image 121, text 122, overview information 123, WeChat Official Account access service card 124, schedule creation service card 125, note snapshot 126, flight booking service card 127, store snapshot 128, and related prompts 129. Among them, image 121 and text 122 are the image and text input by the user. Overview information 123, WeChat Official Account access service card 124, schedule creation service card 125, note snapshot 126, flight booking service card 127, store snapshot 128, and related prompts 129 are the content that the electronic device replies to the user after determining the user's intent based on image 121 and text 122.
[0227] As can be seen from the foregoing description, after the user sends image 121 and text 122 in the dialogue interface, the electronic device can respond to the operation, obtain the entity information "Forbidden City" in image 121, and determine the user's intention as "searching for how to travel to the Forbidden City" based on text 122.
[0228] Next, the electronic device can break down the user intent "searching for how to travel to the Forbidden City" into multiple sub-user intents based on user habits. Here, it is assumed that the sub-user intents include "learning about Forbidden City information," "booking a visit to the Forbidden City," "creating a travel itinerary," "developing a Forbidden City travel guide," "booking flights to Beijing," and "finding restaurants near the Forbidden City." The electronic device can then determine the vertical domain to which each of these sub-user intents belongs and, according to the user's application (large language model) usage preferences for each vertical domain, send these sub-user intents to the corresponding applications (large language models) for processing, and obtain the processing results of the multiple applications (large language models) on the sub-user intents. The local large model in the electronic device can summarize the processing results returned by multiple application large language models to obtain an overview summary, and generate corresponding related prompts based on the processing results returned by multiple large language models, which are then displayed in the user interface 12. Among these:
[0229] The sub-user intent "to learn about information related to the Forbidden City" belongs to the general knowledge category. Assuming the user preference model within this category is model1 provided by the information search application APP1, the electronic device can call model1 to search for this sub-user intent. After obtaining the search results returned by model1, the electronic device can summarize the results to obtain overview information 123. Specifically, overview information 123 can be used to display some basic information about the Forbidden City to the user, including location information, opening hours, etc.
[0230] The sub-user's intent to "book a visit to the Forbidden City" belongs to the activity booking vertical domain. Assuming the user's preferred application within this vertical domain is an instant messaging application (APP2), for example... The electronic device can then use WeChat's search function to search for the sub-user's intent, retrieve the official account information returned by WeChat, and output it to the user in the form of an official account access service card 124. Specifically, the instant messaging application's official account access service card 124 can respond to user operations, such as clicking, causing the electronic device to redirect the interface to the information display interface corresponding to the official account, so that the user can make a reservation for a visit to the Palace Museum.
[0231] The user's intention to "create a travel itinerary" belongs to the scheduling planning category. Assuming the user's preferred application within this category is a scheduling app (APP3), such as a calendar, the electronic device can invoke the new schedule service provided by APP3, displaying the schedule creation service card 125 in the user interface 12. Specifically, the schedule creation service card 125 provided by the scheduling app APP3 can respond to user actions, such as clicking, causing the electronic device to redirect to the corresponding service interface (which can automatically fill in specific schedule items, such as a trip to the Forbidden City), so that the user can create a new travel itinerary.
[0232] The sub-user's intent to "create a Palace Museum travel guide" belongs to the tourism vertical domain. Assuming the user preference model within this vertical domain is model 2 provided by the social application APP4, for example... The provided Da Vinci model allows the electronic device to call model2 to search for the sub-user's intent and display the searched travel guide notes as note snapshots 126 in the user interface 12. Specifically, any snapshot in note snapshots 126 can respond to user actions, such as clicks, causing the electronic device to jump to the note display interface corresponding to that snapshot, allowing the user to quickly browse detailed travel guide notes.
[0233] The sub-user's intent to "book a flight to Beijing" belongs to the ticketing booking vertical. Assuming the user's preferred application within this vertical is a travel booking app (APP5), for example... The electronic device can then access the flight booking service provided by the travel booking application App 3 to process the sub-user's intent and display the flight booking service card 127 in the user interface 12. Specifically, the flight booking service card 127 can respond to user actions, such as a click, causing the electronic device to redirect the interface to the corresponding service page so that the user can purchase a flight to Beijing.
[0234] The sub-user intent "to learn about restaurants near the Forbidden City" belongs to the food vertical. Assuming the user preference model within this vertical is model 4 provided by a consumer review app (APP6), such as the "Little Orange" model provided by Dianping, the electronic device can call model 4 to search for this sub-user intent and display the search results as a store snapshot 128 on the user interface 12. Specifically, the store snapshot 128 can respond to user actions, such as clicks, causing the electronic device to redirect the interface to the corresponding store information display screen so that the user can learn about the store's specific details.
[0235] The related suggestion 129 is a related suggestion generated by the electronic device based on the search results of the corresponding sub-user intent based on moedl1-model4. This related suggestion can facilitate the user to quickly start the next round of dialogue with the electronic device.
[0236] Optionally, the electronic device can also summarize the search results of moedl1-model4 for their respective sub-user intents to obtain overview information 123.
[0237] Optionally, in this embodiment, for the WeChat Official Account Access Service Card 124, Schedule Creation Service Card 125, Note Snapshot 126, Flight Booking Service Card 127, Store Snapshot 128, and Related Prompt 129, regardless of which service card the user clicks, the electronic device will recognize the click as positive feedback from the user, indicating that the sub-user intent corresponding to the service card has been successfully predicted by the electronic device. This positive feedback will be stored in the electronic device as embedded data (i.e., in the aforementioned user preference understanding module 901d). In this way, when the electronic device subsequently decomposes similar user intents, it can more easily determine that the sub-user intent is the sub-user intent that the user may need to achieve (it can be considered that the correlation between the sub-user intent and the decomposed user intent will be higher), and it is also easier to provide application service cards that can achieve the sub-user intent in the response content, so as to maximize the satisfaction of the user's actual needs. For example, assuming the user subsequently clicks on the flight booking service card 127 in the user interface 12 to purchase a flight ticket, the electronic device can combine this user action to determine that for a user intent such as "travel to a certain place", the user likely does indeed have a sub-user intent of "purchasing flight / train tickets". Therefore, in the subsequent VQA process, if the electronic device recognizes the user intent as "travel to a certain place", the sub-user intent of "purchasing flight / train tickets" is more likely to be present in the list of sub-user intents obtained after the electronic device breaks down the user intent.
[0238] Optionally, the electronic device can determine the historical frequency of each sub-user intent through embedded data, and further aggregate the historical frequency of each sub-user intent to determine the display order of the content corresponding to each sub-user intent in the user interface 12, so that users can more quickly locate the information they need. Optionally, the content corresponding to the sub-user intent with a higher historical frequency of realization can be displayed at the top of the user interface 12. For example, assuming that the number of times the user has realized the user intent "create a schedule" is significantly greater than the number of times the user intent "find nearby restaurants" has been realized, then in the user interface 12, the schedule creation service card 125 can be displayed above the store snapshot 128.
[0239] Combining user interface 21 and user interface 81, it can be seen that the electronic device can provide richer information in the response content based on the visual question answering system 90, and the response content can cover the user's actual needs to a greater extent, which can effectively improve the VQA efficiency of the user and the electronic device.
[0240] In this embodiment of the application, the user intent "to travel to the Forbidden City" can be referred to as the "fifth user intent". "to learn about information related to the Forbidden City", "to make an appointment to visit the Forbidden City", "to create a travel itinerary", "to develop a travel guide for the Forbidden City", "to book a flight to Beijing", and "to learn about the food near the Forbidden City" can be referred to as "at least two sub-user intents". "to develop a travel guide for the Forbidden City" and "to book a flight to Beijing" can be referred to as the "first sub-user intent" and "second sub-user intent". The tourism vertical domain and the ticket booking vertical domain can be referred to as the "second vertical domain" and "third vertical domain", respectively. The note snapshot 126 can be referred to as the "second processing result" and the flight booking service card 127 can be referred to as the "third processing result".
[0241] "Creating a travel guide for the Forbidden City" and "Booking a flight to Beijing" can also be referred to as "the third sub-user intent" and "the fourth sub-user intent," respectively. Note snapshot 126 can be referred to as "the fourth processing result," and flight booking service card 127 can be referred to as "the fifth processing result."
[0242] The electronic device provided in the embodiments of this application will be described below.
[0243] The electronic device may be a mobile phone, tablet computer, wearable device, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), or dedicated camera (such as SLR camera, point-and-shoot camera), etc. This application does not limit the specific type of the electronic device.
[0244] Figure 13 The structure of the electronic device is shown as an example.
[0245] The electronic device provided in this application will now be described.
[0246] The electronic device can be a mobile phone, tablet computer, wearable device, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), or dedicated camera (such as SLR camera, point-and-shoot camera), etc. This application does not impose any limitation on the specific type of the electronic device. Specifically, the electronic device can be one of the electronic devices shown in the aforementioned icon display method.
[0247] Figure 9 The structure of the electronic device is shown as an example.
[0248] like Figure 9 As shown, the electronic device 100 may include a processor 110, internal memory 151, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, a sensor module 180, buttons 190, a display screen 194, etc. The sensor module 180 may include a pressure sensor 180A, a touch sensor 180K, etc.
[0249] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0250] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0251] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0252] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0253] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0254] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0255] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0256] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 151, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0257] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0258] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0259] Internal memory 151 can be used to store computer executable program code, which includes instructions. Internal memory 151 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 151 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 151 and / or instructions stored in memory located in the processor.
[0260] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.
[0261] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194. In this embodiment, electronic device 100 can use touch sensor 180K to detect touch operations such as tapping, long-pressing, double-tapping, and dragging on the screen.
[0262] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0263] In this application, the GPU in the electronic device can perform three-dimensional transformation on the content displayed on the display screen 194 after the user drags the information element in the interface, creating a visual effect of the original interface pushing inward on the screen, and displaying recommended application and / or service icons on the side of the screen of the electronic device.
[0264] In this embodiment, the processor 110 can be the processing module 301 in the aforementioned visual question-answering system 30 or the processing module 901 in the aforementioned visual question-answering system 90. Specifically, the processor 110 can be used to: determine a first user intent based on a first user input; determine a first model or a first application based on the embedded data and the first vertical domain to which the first user intent belongs, wherein the embedded data represents the frequency of use of each model or application when the user realizes the user intent in each vertical domain, and the first model or the first application is a model or application whose frequency of use when realizing the user intent in the first vertical domain is greater than a first threshold; call the first model or the first application to process the first user intent, and output the obtained first processing result.
[0265] This application also provides an electronic device, which includes one or more processors and a memory; wherein the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the methods shown in the foregoing embodiments.
[0266] The term "user interface (UI)" used in the specification, claims, and drawings of this application refers to the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form acceptable to the user. The user interface of an application is source code written in a specific computer language such as Java or Extensible Markup Language (XML). This source code is parsed and rendered on the terminal device, ultimately presenting user-recognizable content such as images, text, and buttons. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, images, and text. The attributes and content of controls in the interface are defined using tags or nodes, such as XML tags. <textview> 、 <imgview> 、 <videoview>Nodes define the controls contained in the interface. A node corresponds to a control or property in the interface, and after parsing and rendering, the node is presented as the content visible to the user. In addition, many applications, such as hybrid applications, often contain web pages within their interfaces. A web page, also known as a webpage, can be understood as a special control embedded in the application interface. Web pages are source code written in a specific computer language, such as Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), JavaScript (JS), etc. Web page source code can be loaded and displayed as user-readable content by a browser or a web page display component with browser-like functionality. The specific content contained in a webpage is also defined through tags or nodes in the webpage source code; for example, HTML uses tags or nodes to define the content. 、 、 <video> 、 <canvas>To define the elements and attributes of a webpage.
[0267] The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0268] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if..." or "after..." or "in response to determining..." or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is detected" can be interpreted as meaning "if determining..." or "in response to determining..." or "when (the stated condition or event) is detected" or "in response to detecting (the stated condition or event)".
[0269] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0270] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.< / canvas> < / video> < / videoview> < / imgview> < / textview>
Claims
1. An information processing method, characterized in that, include: Determine the first user's intent based on the first user's input; The first model or first application is determined based on the tracking data and the first vertical domain to which the first user intent belongs. The tracking data represents the frequency of use of each model or application when the user realizes the user intent in each vertical domain. The first model or first application is a model or application whose frequency of use when realizing the user intent in the first vertical domain is greater than a first threshold. The first model or the first application is invoked to process the first user intent, and the resulting first processing result is output.
2. The method according to claim 1, characterized in that, When the first model is invoked to process the first user intent, the first processing result includes one or more of the following: a first overview summary, at least one first content snapshot, and at least one first related prompt. The first overview summary is a summary obtained by the first model from the search results obtained from the search of the user intent. The at least one content snapshot is a snapshot obtained by the first model based on at least one link contained in the search results. The at least one first related prompt is a prompt generated by the first model after predicting subsequent user intent based on the search results. Alternatively, when the first application is invoked to process the first user intent, the first processing result includes a first application service card, the first application service card corresponds to a first application service, and the first application service is an application service provided by the first application for processing the first user intent.
3. The method according to claim 1 or 2, characterized in that, The first user input includes a first image and a first text, and determining the first user intent based on the first user input includes: The entities contained in the first image are extracted to obtain first entity information, which is either an image entity or a text entity. The first text is identified to obtain the second user intent, which is the user intent that lacks slot information; The first user intent is obtained by filling the missing slot information in the second user intent with the first entity information.
4. The method according to claim 3, characterized in that, The data collected also represents the number of times the user performs various user intentions for each type of entity data. The method further includes: Based on the first entity type to which the first entity information belongs and the second user intent, the number of times the user realizes the second user intent for the first entity type in the tracking data is updated.
5. The method according to claim 1 or 2, characterized in that, The first user input includes only the second image, and the embedded data also characterizes the number of times the user performs various user intentions for each type of entity. Determining the first user intention based on the first user input includes: The entities contained in the second image are extracted to obtain the second entity information; Based on the embedded data and the second entity type to which the second entity information belongs, at least one third user intent is determined, and at least one second associated prompt is generated based on the at least one third user intent. The at least one third user intent includes one or more user intents that the user performs most frequently for the entity information of the second entity type. In response to a user's click on a third associated prompt in at least one second associated prompt, the second entity information is combined with the fourth user intent corresponding to the third associated prompt to obtain the first user intent.
6. The method according to claim 5, characterized in that, In response to the user's click on the third associated prompt, based on the second entity type to which the second entity information belongs and the fourth user intent, the number of times the user performs the fourth user intent for the second entity type in the tracking data is updated.
7. An information processing method, characterized in that, Applied to electronic devices, the method includes: Determine the fifth user's intent based on the second user's input; The fifth user intent is broken down into at least two sub-user intents, the at least two sub-user intents including a first sub-user intent and a second sub-user intent; A second model or second application is determined based on the tracking data and the second vertical domain to which the first sub-user intent belongs; a third model or third application is determined based on the tracking data and the third vertical domain to which the second sub-user intent belongs; the tracking data represents the frequency of use of each model or application when the user realizes the user intent in each vertical domain; the second model or second application is a model or application whose frequency of use when realizing the user intent in the second vertical domain is greater than a first threshold; the third model or third application is a model or application whose frequency of use when realizing the user intent in the third vertical domain is greater than the first threshold. The second model or the second application is invoked to process the first sub-user's intent, resulting in a second processing result; The third model or third application is invoked to process the second sub-user's intent, resulting in a third processing result; Display some or all of the multiple processing results obtained from processing the at least two sub-user intentions, the multiple processing results including the second processing result and the third processing result.
8. The method according to claim 7, characterized in that, The multiple processing results include a first type of processing result and a second type of processing result. The first type of processing result is a fourth processing result obtained by processing the third sub-user intent using at least one fourth model. The second type of processing result includes a fifth processing result obtained by processing the fourth sub-user intent using at least one fourth application. The fourth processing result includes one or more of a second overview summary generated based on the first sub-user intent, at least one second content snapshot, and at least one fourth related prompt. The fifth processing result includes an application service card corresponding to at least one application service provided by the at least one second application for the second sub-user intent. The display of some or all of the multiple processing results includes: The multiple processing results are filtered, and the fourth and fifth processing results are displayed.
9. The method according to claim 7 or 8, characterized in that, The second user input includes a third image and second text, and determining the fifth user intent based on the second user input includes: The entities contained in the third image are extracted to obtain third entity information, which is either an image entity or a text entity; The second text is identified to obtain the sixth user intent, which is a user intent that lacks slot information; The missing slot information in the sixth user intent is filled with the third entity information to obtain the fifth user intent.
10. The method according to claim 9, characterized in that, The data collected also represents the number of times the user performs various user intentions for each type of entity data. The method further includes: Based on the third entity type to which the third entity information belongs and the sixth user intent, the number of times the user implements the sixth user intent for the third entity type in the data points is updated.
11. The method according to claim 7 or 8, characterized in that, The second user input includes only the fourth image, and the embedded data also characterizes the number of times the user performs various user intentions for each type of entity data. Determining the fifth user intention based on the second user input includes: The entities contained in the fourth image are extracted to obtain the fourth entity information, wherein the third entity information is an image entity or a text entity; Based on the embedded data and the fourth entity type to which the fourth entity information belongs, at least one seventh user intent is determined, and at least one fifth associated prompt is generated based on the at least one seventh user intent. The at least one seventh user intent includes one or more user intents that the user performs most frequently for the entity information of the fourth entity type. In response to a user's click on the sixth associated prompt in the at least one fifth associated prompt, the fourth entity information is combined with the eighth user intent corresponding to the sixth associated prompt to obtain the fifth user intent.
12. The method according to claim 11, characterized in that, In response to the user's click on the sixth associated prompt, based on the fourth entity type to which the fourth entity information belongs and the eighth user intent, the number of times the user has historically implemented the eighth user intent for the fourth entity type in the tracking data is updated.
13. An electronic device, characterized in that, The electronic device includes: one or more processors, memory, and a display screen; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-12.
14. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-12.
15. A computer program product containing instructions, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-12.
16. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-12.