Content generation method and electronic device
By searching for and receiving multimodal inputs through a multimodal agent, electronic devices display the results in card format, solving the problem of unsatisfactory results generated by multimodal systems, enabling convenient access to information and services, and improving the user experience.
Patent Information
- Application Number
- PCT/CN2025/101542
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-27
- Filing Date
- 2025-06-17
- Publication Date
- 2025-12-26
AI Technical Summary
When users are not satisfied with the results generated by multimodal systems, it is difficult to quickly refresh the page to obtain more specific and authoritative results, resulting in a poor user experience. Furthermore, existing systems cannot intelligently recommend relevant information and services, nor can they share results in the form of cards.
By using a search multimodal agent, electronic devices can receive voice, text, or image input, perform searches, and display multiple information and service search results in card format, enabling users to share results with fewer interactions.
It improves the user experience, allowing users to access relevant information and services more conveniently, reducing the number of interactions with the system, and increasing the efficiency of information and service acquisition.
Smart Images

Figure CN2025101542_26122025_PF_FP_ABST
Abstract
Description
Content generation method and electronic device
[0001] The present application claims priority to the Chinese patent application No. 202410807702.4, filed on June 20, 2024, entitled "Content generation method and electronic device", the priority of which is hereby claimed in its entirety. The present application claims priority to the Chinese patent application No. 202411026444.2, filed on July 27, 2024, entitled "Content generation method and electronic device", the priority of which is hereby claimed in its entirety. The contents of the aforementioned applications are hereby incorporated by reference in their entirety. TECHNICAL FIELD
[0002] The present application relates to the technical field of terminal, and in particular, to a content generation method and an electronic device. BACKGROUND
[0003] With the continuous development of computers, multi-modal systems are becoming more and more common. Users can input multi-modal information such as pictures, audio, video, and text into multi-modal systems to obtain satisfactory results. However, if the user is not satisfied with the results generated by the multi-modal system, the multi-modal system is difficult to quickly refresh the generated results to make the user obtain more specific and authoritative results in the case that the user no longer interacts with the multi-modal system. This brings a poor experience to the user. Therefore, after the user inputs multi-modal information into the multi-modal system, how to obtain the results that the user wants in the case that the number of times of multi-modal interaction between the user and the multi-modal system is as few as possible is particularly important. SUMMARY
[0004] The present application provides a content generation method and an electronic device. By using a search multi-modal Agent, the electronic device can receive a first multi-modal input, which can include a voice input and one or more of the following: a text input, an image input; the electronic device can search the first multi-modal input; the electronic device can display the search results of searching the first multi-modal input in the form of a card, the search results have multiple, and the multiple search results can include information search results and service search results; the electronic device can select the search results for sharing. In this way, the user can reduce the number of interactions with the multi-modal system, more conveniently obtain different information and services related to the input, and improve the user experience.
[0005] In a first aspect, the present application provides a content generation method, the method comprising: a first electronic device entering a first state, the first state being used for the first electronic device to receive a first multi-modal input; the first electronic device receiving the first multi-modal input, the first multi-modal input comprising a voice input and one or more of the following: a text input, an image input; the first electronic device searching the first multi-modal input; and the first electronic device displaying the search results of searching the first multi-modal input in a card form.
[0006] In some embodiments, the first electronic device can be the electronic device 100 shown in FIG. 12, the first multi-modal input can comprise a voice input (e.g., a voice operation on a selected object by a user) and one or more of the following: a text input, an image input (including but not limited to a circle selection, a pointing, a checkmark, a highlight, a line drawing, a long press, a click on a text or an image), wherein the voice input is received through a voice assistant; and the first state can be a state in which the electronic device 100 starts to search a multi-modal Agent service.
[0007] By implementing the method of the first aspect, the user can reduce the number of interactions with the multi-modal system, more conveniently obtain different information and services related to the same multi-modal input, and can select different information and services for sharing according to needs, thereby improving the user experience.
[0008] In combination with the first aspect, in some embodiments, before the first electronic device enters the first state, the method further comprises: the first electronic device displaying a first user interface, the first user interface comprising text content and image content; wherein the text input is text content selected by the user in the first user interface; and the image input is image content selected by the user in the first user interface; and detecting a first operation, the first electronic device triggering a voice assistant.
[0009] In some embodiments, the first user interface can be, for example, the user interface 1000 shown in FIG. 1A, and the first operation can be an operation of triggering the voice assistant (e.g., an operation of detecting that the user long-presses the power key for 1 second).
[0010] In this way, before the first electronic device enters the first state, the first electronic device can trigger the voice assistant to enable the user to interact with the search multi-modal Agent through the voice assistant.
[0011] In combination with the first aspect, in some embodiments, the first electronic device receiving the first multi-modal input specifically comprises: the first electronic device receiving a first input and a second input; wherein a time interval between the first electronic device ending to receive the first input and starting to receive the second input is less than a first time length; and the first input and the second input belong to the first multi-modal input.
[0012] The first input and the second input both belong to the first multi-modal input, and the first time length can be a preset time length, for example, 3 seconds.
[0013] For example, within the first time length during which the first electronic device ends receiving the first input, if the user starts the second input, the first electronic device can receive the second input, otherwise the first electronic device can stop receiving the second input.
[0014] In combination with the first aspect, in some embodiments, the first electronic device displays the search results searched based on the first multi-modal input in the form of cards, specifically including: during the process of receiving the first multi-modal input, the first electronic device first displays the first search results, and then increases the display of the second search results; the first search results and the second search results are of different types, and the types include information search results and service search results.
[0015] The first search results can be first part search results first displayed by the first electronic device during the process of receiving the first multi-modal input, and the second search results can be second part search results later displayed by the first electronic device during the process of receiving the first multi-modal input.
[0016] In this way, the first electronic device can increase the display of the search results synchronously with the increase of the received multi-modal input.
[0017] For example, during the process of receiving the first multi-modal input of the user's circle selecting a picture and performing a voice operation "what does this talk about, and is there a ticket" on the picture selected by the user, the first electronic device can first display the first search results corresponding to "what does this talk about", and then increase the display of the second search results corresponding to "is there a ticket", wherein the first search results are information search results of a musical A, and the second search results are ticket purchase service search results of the musical A.
[0018] In combination with the first aspect, in some embodiments, the search results are multiple, and the multiple search results include information search results and service search results.
[0019] In combination with the first aspect, in some embodiments, the search results searched based on the first multi-modal input are displayed in the form of cards, specifically including: the first electronic device displays the information search results and the service search results in different cards respectively; wherein the information search results of different information types are displayed in different cards respectively; wherein the service search results of different service types are displayed in different cards respectively.
[0020] In this way, the first electronic device can display the search results according to the types of the search results.
[0021] Exemplarily, the Query is "What is this about? Are there any tickets?", the first electronic device can search to obtain search results of drama music information, drama video information, ticket purchasing service, travel service, etc., and then the first electronic device can display the search results of drama music information, drama video information, ticket purchasing service, travel service, etc. in a classified manner.
[0022] In combination with the first aspect, in some embodiments, the method further includes: detecting a second operation, the first electronic device displaying a second user interface, the second user interface being used for selecting a plurality of search results; detecting a third operation, the first electronic device sharing the selected search results.
[0023] The second user interface may, for example, be the user interface 1040, which can be used for selecting a plurality of search results, the second operation can be a sharing operation (e.g., the user triggers the operation of sharing the search results through a voice instruction or a touch screen operation) for the search results, and the third operation can be an operation of selecting the search results for sharing (e.g., the user checks the search results displayed in the form of cards and clicks the sharing control).
[0024] Specifically, after displaying the search results of the first multi-modal input in the form of cards, detecting a second operation (e.g., the user triggers the operation of sharing the search results through a voice instruction or a touch screen operation), the first electronic device can display a second user interface, which can be used for selecting a plurality of search results. Detecting a third operation (e.g., the user checks the search results displayed in the form of cards and clicks the sharing control), the first electronic device can share the search results selected by the user.
[0025] In combination with the first aspect, in some embodiments, the first electronic device searches the first multi-modal input, specifically including: the first electronic device searches based on a first user intent, the first user intent being recognized by the first electronic device based on the first multi-modal input.
[0026] The first user intent can be recognized by the first electronic device based on the first multi-modal input.
[0027] In this way, the first electronic device can invoke a search engine to search based on the first user intent to obtain search results.
[0028] In combination with the first aspect, in some embodiments, the method further includes: determining, by the first electronic device, a search strategy of the search according to the first user intent and a first cache, the first cache including search results corresponding to different user intents respectively; and the search strategy specifically includes: if the first user intent corresponds to a search result in the first cache, the search result is the search result corresponding to the first user intent obtained from the first cache; otherwise, calling a search engine, and the search result is obtained by searching the first user intent by using the search engine.
[0029] The first cache can be a multi-modal cache in the search multi-modal Agent shown in FIG. 7.
[0030] The multi-modal system of the embodiments of the present application has the ability of autonomous planning, autonomous decision-making, and autonomous execution, can autonomously formulate a reasonable strategy and plan according to a user intent and a multi-modal cache, reduces the overall time delay, improves the work efficiency, and further improves the user experience.
[0031] In combination with the first aspect, in some embodiments, the type of the search engine includes one or more of the following: visual search, text search, and fusion search; the visual search is based on image search; the text search is based on text search; and the fusion search is based on image and text search.
[0032] For example, the Query is “What is the brand of the hat worn by celebrity A”, and the first electronic device can determine that the type of the search engine called according to the commodity intent and the person intent is the commodity search in the visual search and the person search in the vertical domain Box in the text search.
[0033] In combination with the first aspect, in some embodiments, the first electronic device searches the first multi-modal input, specifically including: calling, by the first electronic device, a first search engine to search the first multi-modal input, to obtain a plurality of third search results, the first search engine being a search engine called by the first electronic device according to the search strategy; matching, by the first electronic device, the first user intent with the plurality of third search results, to obtain a fourth search result, the matching degree of the fourth search result with the first user intent being greater than the matching degree of the third search result with the first user intent, and the search result displayed in the form of a card being the fourth search result.
[0034] The first search engine can be a search engine called by the first electronic device according to the search strategy, the third search result can be a search result obtained by calling the first search engine by the first electronic device to search the first multi-modal input, and the fourth search result can be a search result obtained by matching the first user intent with the plurality of third search results by the first electronic device, which is the search result of searching the first multi-modal input.
[0035] In this way, the search results displayed in the card form are more matched with the user's intention than the search results obtained by calling the search engine, and the user experience is improved.
[0036] In a second aspect, the present application provides an electronic device, comprising a processor and a memory; wherein the memory is coupled with the processor, and the memory is configured to store computer program codes, the computer program codes comprising computer instructions, which, when executed by the processor, cause the electronic device to perform the method of any one of the first aspect.
[0037] In a third aspect, the present application provides a computer readable storage medium, which stores computer programs or computer instructions, and the computer programs or computer instructions are executed by a processor to implement the method of any one of the first aspect.
[0038] In a fourth aspect, the present application provides a computer program product, which, when executed by a processor, implements the method of any one of the first aspect.
[0039] In a fifth aspect, the present application provides a chip, comprising a processor and a memory, wherein the memory is configured to store computer programs or computer instructions, and the processor is configured to execute the computer programs or computer instructions stored in the memory, so that the chip performs the method of any one of the first aspect.
[0040] The solutions provided in the second aspect to the fifth aspect are used to implement or cooperate to implement the method provided in the first aspect, and thus can achieve the same or corresponding beneficial effects as the corresponding method in the first aspect, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0041] FIGS. 1A-1C are a set of user interface schematic diagrams provided by the embodiments of the present application;
[0042] FIG. 2 is a user interface schematic diagram provided by the embodiments of the present application;
[0043] FIG. 3 is another user interface schematic diagram provided by the embodiments of the present application;
[0044] FIGS. 4A-4B are another set of user interface schematic diagrams provided by the embodiments of the present application;
[0045] FIGS. 5A-5C are another set of user interface schematic diagrams provided by the embodiments of the present application;
[0046] FIGS. 6A-6C are another set of user interface schematic diagrams provided by the embodiments of the present application;
[0047] FIG. 7 is a schematic diagram of the working principle of a content generation method provided by the embodiments of the present application;
[0048] FIG. 8A is a flow diagram of a content generation method based on searching a multi-modal Agent according to an embodiment of the present application;
[0049] FIG. 8B is another flow diagram of a content generation method based on searching a multi-modal Agent according to an embodiment of the present application;
[0050] FIG. 9 is a flow diagram of a content generation method based on a search strategy according to an embodiment of the present application;
[0051] FIG. 10 is a flow diagram of a content generation method based on a search result display strategy according to an embodiment of the present application;
[0052] FIG. 11 is a flow diagram of a content generation method according to an embodiment of the present application;
[0053] FIG. 12 is a structural diagram of an electronic device 100 according to an embodiment of the present application. DETAILED DESCRIPTION
[0054] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application.
[0055] With the continuous development of computers, multi-modal systems are becoming more and more common. Users can input multi-modal information such as pictures, audio, video, and text into multi-modal systems to obtain satisfactory results. However, if the user is not satisfied with the results generated by the multi-modal system, the multi-modal system is difficult to quickly refresh the generated results to make the user obtain more specific and authoritative results in the case that the user no longer interacts with the multi-modal system. This brings a poor experience to the user. Therefore, after the user inputs multi-modal information into the multi-modal system, how to obtain the results that the user wants in the case that the number of multi-modal interactions between the user and the multi-modal system is as few as possible is particularly important.
[0056] The current multi-modal system generates single results, cannot intelligently recommend related information and services according to the input, and does not support sharing the results in the form of cards.
[0057] Based on the above problems, the embodiment of the present application provides a content generation method and an electronic device. By using a search multi-modal Agent, the electronic device 100 can receive a first multi-modal input, which can include a voice input and one or more of the following: a text input, an image input; the electronic device 100 can search the first multi-modal input; the electronic device 100 can display the search results of searching the first multi-modal input in the form of a card, the search results can include information search results and service search results; and the electronic device 100 can select the search results for sharing. In this way, the user can reduce the number of interactions with the multi-modal system and more conveniently obtain different information and services related to the input, thereby improving the user experience.
[0058] First, a series of graphical user interfaces provided by the embodiment of the present application will be described in detail. The graphical user interface is essentially a human-computer interaction method, which is an interface control method of the electronic device 100.
[0059] As shown in FIGS. 1A-1C, in some embodiments, the electronic device 100 can display a first user interface including text content and image content. The aforementioned text input can be text content selected by the user in the first user interface, and the aforementioned image input can be image content selected by the user in the first user interface, which includes pictures and videos. After detecting a first operation (for example, detecting an operation of the user pressing the power key for 1 second), the electronic device 100 can trigger a voice assistant to enable the user to interact with the search multi-modal Agent through the voice assistant. The first user interface can be, for example, the user interface 1000 shown in FIG. 1A. After triggering the voice assistant, the electronic device 100 can enter a first state, which can be used for the electronic device 100 to receive a first multi-modal input. The first state can mean that the electronic device 100 starts the search multi-modal Agent service. For example, after detecting an operation of the user double-tapping the screen, the electronic device 100 can enter the first state and display the user interface 1010 shown in FIG. 1B. After entering the first state, the electronic device 100 can directly receive the first multi-modal input, so that the user can perform voice input, image input, and text input without other operations. The first multi-modal input can be, for example, an image input of the user circling a picture (for example, the image content 1023 shown in FIG. 1C) and a voice input of the user circling the picture, wherein the voice input is received through the voice assistant, and the state of the voice input can be identified by a voice input state icon (for example, the icon 1022 shown in FIG. 1C). The user interface in which the electronic device 100 receives the first multi-modal input can be, for example, the user interface 1020 shown in FIG. 1C.
[0060] Exemplarily, as shown in FIG. 1A, the electronic device 100 can display a user interface 1000, which can include text content (e.g., text content 1002), image content (e.g., image content 1001). Among them, the text content can be selected by the user as text input; the image content can be selected by the user as image input, which includes pictures and videos.
[0061] In some embodiments, after detecting the first operation (e.g., detecting the operation of the user pressing the power key for 1 second), the electronic device 100 can trigger the voice assistant (not shown in the figure) to enable the user to interact with the search multimodal Agent through the voice assistant. After triggering the voice assistant, the electronic device 100 can enter a first state, which can be used for the electronic device 100 to receive the first multimodal input. The first state can refer to the state in which the electronic device 100 starts the search multimodal Agent service. Exemplarily, after detecting the operation of the user double-finger pressing the screen, the electronic device 100 can enter the first state and display the user interface 1010 shown in FIG. 1B.
[0062] Exemplarily, as shown in FIG. 1B, the electronic device 100 can display a user interface 1010, which can include an input box 1011, image content (e.g., image content 1013), text content (e.g., text content 1014). Among them, the image content can be selected by the user as image input, and the text content can be selected by the user as text input. The input box 1011 can be used to receive voice input. The input box 1011 can include a voice input state icon (e.g., icon 1012), which can be used to prompt the user to directly input voice.
[0063] In some embodiments, after entering the first state, the electronic device 100 can directly receive a first multi-modal input, such that the user can make a voice input, an image input, and a text input without other operations. The first multi-modal input can be, for example, an image input of a user circled picture and a voice input of a voice operation on the user circled picture, where the voice input is received through a voice assistant. Specifically, the electronic device 100 can receive a first input, which can be, for example, an image input of a user circled picture, and a second input, which can be, for example, a voice input of a voice operation on the user circled picture. Wherein, the time interval between the end of receiving the first input and the start of receiving the second input is less than a first time length (e.g., 3 seconds), and the first input and the second input belong to the first multi-modal input. Illustratively, within 3 seconds of the end of receiving the image input of the user circled picture by the electronic device 100, if the user starts the voice operation on the user circled picture, the electronic device 100 can receive the voice input of the voice operation on the user circled picture, otherwise the electronic device 100 can stop receiving the voice input of the voice operation on the user circled picture. After receiving the first multi-modal input (e.g., the image input of the user circled picture and the voice input of the voice operation on the user circled picture), the electronic device 100 can display a user interface 1020 as shown in FIG. 1C.
[0064] Illustratively, as shown in FIG. 1C, the electronic device 100 can display a user interface 1020, which can include an input box 1021, image content (e.g., image content 1023), where the image content 1023 is image content selected by the user as an image input, and the input box 1021 can be used to receive a voice input (e.g., "what does this say, are there tickets?"). The input box 1021 can include a voice input state icon (e.g., icon 1022), which can be used to identify the state of the electronic device 100 receiving the voice input.
[0065] As shown in FIG. 2, in some embodiments, after receiving the first multi-modal input, the electronic device 100 can search the first multi-modal input. Then, the electronic device 100 can display the search results of searching the first multi-modal input in the form of cards, where the search results can be multiple, and the multiple search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, where the information search results of different information types are displayed in different cards respectively, and the service search results of different service types are displayed in different cards respectively.
[0066] Exemplarily, as shown in FIG. 2, the electronic device 100 can display a user interface 1030, which can include a plurality of search results. Among them, the plurality of search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, wherein information search results of different information types can be displayed in different cards respectively, and service search results of different service types can be displayed in different cards respectively. For example, in the user interface 1030, the search results can include information search results of a musical A and ticket purchase service search results of the musical A, wherein the information search results of the musical A can include summary information search results of the musical A and encyclopedia information search results of the musical A, and the summary information search results of the musical A, the encyclopedia information search results of the musical A, and the ticket purchase service search results of the musical A can be displayed in different cards respectively.
[0067] In some embodiments, in the process of receiving the first multi-modal input, the electronic device 100 can first display the first search results, and then increase the display of the second search results, wherein the types of the first search results and the second search results are different, and the types of the search results can include information search results and service search results. Exemplarily, in the process of receiving the first multi-modal input of the user's circled picture and the voice operation "what is this about, and is there a ticket" on the user's circled picture, the electronic device 100 can first display the first search results corresponding to "what is this about", and then increase the display of the second search results corresponding to "is there a ticket", wherein the first search results are information search results of the musical A, and the second search results are ticket purchase service search results of the musical A.
[0068] As shown in FIG. 3, in some embodiments, after displaying the search results of the first multi-modal input in the form of cards, the electronic device 100 can receive a second operation (for example, an operation triggered by the user through a voice instruction or a touch screen operation to share the search results), and in response to the operation, the electronic device 100 can display a second user interface (for example, a user interface 1040), which can be used to select a plurality of search results. After detecting a third operation (for example, an operation of checking the search results displayed in the form of cards and clicking a sharing control), the electronic device 100 can share the search results selected by the user.
[0069] Exemplarily, as shown in FIG. 3, the electronic device 100 can display a user interface 1040, which can include a plurality of search results displayed in the form of cards, a sharing control (e.g., control 1041). Among them, the plurality of search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, wherein the information search results of different information types can be displayed in different cards respectively, and the service search results of different service types can be displayed in different cards respectively. For example, in the user interface 1040, the search results can include information search results of musical A and travel service search results of musical A, wherein the information search results of musical A can include summary information search results of musical A and song information search results of musical A, and the summary information search results of musical A, the song information search results of musical A and the travel service search results of musical A can be displayed in different cards respectively. Among them, the sharing control can be used to trigger sharing of the search result selected by the user.
[0070] As shown in FIGS. 4A-4B, in some embodiments, after displaying the search results, the electronic device 100 can store the received first multi-modal input and the search results corresponding to the first multi-modal input into a first cache, and display a history record user interface (e.g., user interface 1050 shown in FIG. 4B). In the case of receiving the same or similar first multi-modal input, the electronic device 100 can find the search results corresponding to the same or similar first multi-modal input in the first cache, and directly display the search results in the form of cards.
[0071] Exemplarily, as shown in FIG. 4A, the electronic device 100 can display a user interface 1030, which can also include a history record control (e.g., control 1031), which can be used to trigger display of the cache record (history record) of the first multi-modal input.
[0072] In some embodiments, the electronic device 100 can receive an input operation (e.g., single click) of the user acting on the history record control (e.g., control 1031), and in response to the input operation, the electronic device 100 can display the user interface 1050 as shown in FIG. 4B.
[0073] Exemplarily, as shown in FIG. 4B, the electronic device 100 can display a user interface 1050, which can include a cache record (history record) of the first multi-modal input, which can be arranged in a time sequence from bottom to top, with the most recent first multi-modal input search record arranged at the top of the history record. In some embodiments, the electronic device 100 can receive an input operation (e.g., a single click) of the user acting on the cache record (history record) of the first multi-modal input, and in response to the input operation, the electronic device 100 can display the search result corresponding to the first multi-modal input. For example, the electronic device 100 can receive a single click operation of the user acting on the first multi-modal input cache record (history record) containing the voice input "what is this about, and are there tickets?", and in response to the operation, the electronic device 100 can re-display the user interface 1030 as shown in FIG. 4A.
[0074] As shown in FIGS. 5A-5C, in some embodiments, the electronic device 100 can display a first user interface including text content, image content. Wherein the aforementioned text input can be text content selected by the user in the first user interface, and the aforementioned image input can be image content selected by the user in the first user interface, the image content including pictures and videos. Upon detecting a first operation (e.g., detecting an operation of the user pressing the power key for 1 second), the electronic device 100 can trigger the voice assistant to enable the user to interact with the search multi-modal Agent through the voice assistant. The first user interface can be, for example, the user interface 1000 shown in FIG. 5A. After triggering the voice assistant, the electronic device 100 can enter a first state, which can be a state in which the electronic device 100 starts the search multi-modal Agent service. Exemplarily, upon detecting an operation of the user double-tapping the screen, the electronic device 100 can enter the first state and display the user interface 1010 shown in FIG. 5B. After entering the first state, the electronic device 100 can directly receive the first multi-modal input, so that the user can perform voice input, image input, and text input without other operations. The first multi-modal input can be, for example, an image input of the user pointing to a picture (e.g., the image content 1063 shown in FIG. 5C) and a voice input of the user performing a voice operation on the picture pointed to by the user, wherein the voice input is received through the voice assistant, and the state of the voice input can be identified by a voice input state icon (e.g., the icon 1062 shown in FIG. 5C). The user interface in which the electronic device 100 receives the first multi-modal input can be, for example, the user interface 1060 shown in FIG. 5C.
[0075] The relevant description of the user interface 1000 shown in FIG. 5A can refer to the content of the aforementioned FIG. 1A, which will not be repeated here.
[0076] In some embodiments, after detecting the first operation (e.g., detecting the operation of the user pressing the power key for 1 second), the electronic device 100 can trigger a voice assistant (not shown in the figure) to enable the user to interact with the search multimodal Agent through the voice assistant. After triggering the voice assistant, the electronic device 100 can enter a first state, which can be used for the electronic device 100 to receive a first multimodal input. The first state can refer to a state in which the electronic device 100 starts the search multimodal Agent service. Illustratively, after detecting the operation of the user double-tapping the screen, the electronic device 100 can enter the first state and display the user interface 1010 shown in FIG. 5B.
[0077] The related description of the user interface 1010 shown in FIG. 5B can refer to the foregoing content of FIG. 1B, which will not be described here again.
[0078] In some embodiments, after entering the first state, the electronic device 100 can directly receive the first multimodal input, so that the user can perform voice input, image input, and text input without other operations. The first multimodal input can be, for example, image input of the user pointing to a picture and voice input of the user performing voice operation on the picture pointed to by the user, where the voice input is received through the voice assistant. Specifically, the electronic device 100 can receive a first input, which can be, for example, image input of the user pointing to a picture, and a second input, which can be, for example, voice input of the user performing voice operation on the picture pointed to by the user. Wherein, the time interval between the end of the electronic device 100 receiving the first input and the start of the electronic device 100 receiving the second input is less than a first time length (e.g., 3 seconds), and the first input and the second input belong to the first multimodal input. Illustratively, within 3 seconds of the electronic device 100 ending to receive image input of the user pointing to a picture, if the user starts to perform voice operation on the picture pointed to by the user, the electronic device 100 can receive voice input of the user performing voice operation on the picture pointed to by the user, otherwise the electronic device 100 can stop receiving voice input of the user performing voice operation on the picture pointed to by the user. After receiving the first multimodal input (e.g., image input of the user pointing to a picture and voice input of the user performing voice operation on the picture pointed to by the user), the electronic device 100 can display the user interface 1060 shown in FIG. 5C.
[0079] Exemplarily, as shown in FIG. 5C, the electronic device 100 can display a user interface 1060, which can include an input box 1061, image content (e.g., image content 1063), where the image content 1063 is image content selected by the user as image input, and the input box 1061 can be used to receive voice input (e.g., "what does this say, are there tickets?"). The input box 1061 can include a voice input status icon (e.g., icon 1062), which can be used to identify the status of the electronic device 100 receiving voice input.
[0080] In some embodiments, after receiving the first multi-modal input, the electronic device 100 can search the first multi-modal input. Then, the electronic device 100 can display the search results of searching the first multi-modal input in the form of cards, where the search results can be multiple, and the multiple search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, where the information search results of different information types are displayed in different cards respectively, and the service search results of different service types are displayed in different cards respectively. The user interface in which the electronic device 100 displays the search results of searching the first multi-modal input in the form of cards can be, for example, the user interface 1030 shown in FIG. 2, and the related description of the user interface 1030 can refer to the foregoing content of FIG. 2, which will not be repeated here.
[0081] In some embodiments, after displaying the search results of searching the first multi-modal input in the form of cards, the electronic device 100 can receive a second operation (e.g., an operation triggered by the user through voice instruction or touch screen operation to share the search results), and in response to the operation, the electronic device 100 can display a second user interface (e.g., user interface 1040), which can be used to select multiple search results. After detecting a third operation (e.g., an operation of checking the search results displayed in the form of cards and clicking a sharing control), the electronic device 100 can share the search results selected by the user. The related description of the user interface 1040 can refer to the foregoing content of FIG. 3, which will not be repeated here.
[0082] As shown in FIGS. 6A-6C, in some embodiments, the electronic device 100 can display a first user interface including text content, image content. Wherein, the aforementioned text input can be text content selected by the user in the first user interface, and the aforementioned image input can be image content selected by the user in the first user interface, the image content including pictures and videos. Upon detecting a first operation (e.g., detecting an operation of the user long-pressing the power key for 1 second), the electronic device 100 can trigger a voice assistant to enable the user to interact with the search multimodal Agent through the voice assistant. The first user interface can be, for example, the user interface 1070 shown in FIG. 6A. After triggering the voice assistant, the electronic device 100 can enter a first state, which can be used for the electronic device 100 to receive a first multimodal input. The first state can refer to a state in which the electronic device 100 starts the search multimodal Agent service. Illustratively, upon detecting an operation of the user double-finger pressing the screen, the electronic device 100 can enter the first state and display the user interface 1080 shown in FIG. 6B. After entering the first state, the electronic device 100 can directly receive the first multimodal input, so that the user can perform voice input, image input, and text input without other operations. The first multimodal input can be, for example, an image input (e.g., the image content 1093 shown in FIG. 6C) of the user circling a video pause frame and a voice input of the user performing voice operation on the video pause frame circled by the user, wherein the voice input is received through the voice assistant, and the state of the voice input can be identified by a voice input state icon (e.g., the icon 1092 shown in FIG. 6C). The user interface in which the electronic device 100 receives the first multimodal input can be, for example, the user interface 1090 shown in FIG. 6C.
[0083] Illustratively, as shown in FIG. 6A, the electronic device 100 can display a user interface 1070, which can include text content (e.g., the text content 1072) and image content (e.g., the image content 1071). Wherein, the text content can be selected by the user as a text input; and the image content can be selected by the user as an image input, the image content including pictures and videos.
[0084] In some embodiments, upon detecting a first operation (e.g., detecting an operation of the user long-pressing the power key for 1 second), the electronic device 100 can trigger a voice assistant (not shown in the figure) to enable the user to interact with the search multimodal Agent through the voice assistant. After triggering the voice assistant, the electronic device 100 can enter a first state, which can be used for the electronic device 100 to receive a first multimodal input. The first state can refer to a state in which the electronic device 100 starts the search multimodal Agent service. Illustratively, upon detecting an operation of the user double-finger pressing the screen, the electronic device 100 can enter the first state and display the user interface 1080 shown in FIG. 6B.
[0085] Exemplarily, as shown in FIG. 6B, the electronic device 100 can display a user interface 1080, which can include an input box 1081, image content (e.g., image content 1083), text content (e.g., text content 1084). Wherein, the image content can be selected by the user as image input, and the text content can be selected by the user as text input. The input box 1081 can be used to receive voice input. The input box 1081 can include a voice input state icon (e.g., icon 1082), which can be used to prompt the user to directly input voice.
[0086] In some embodiments, after entering the first state, the electronic device 100 can directly receive a first multi-modal input, so that the user can perform voice input, image input and text input without other operations. The first multi-modal input may, for example, be image input of the user circling a video pause frame and voice input of the user performing voice operation on the video pause frame circled by the user, wherein the voice input is received through the voice assistant. Specifically, the electronic device 100 can receive a first input, which may, for example, be image input of the user circling a video pause frame, and a second input, which may, for example, be voice input of the user performing voice operation on the video pause frame circled by the user. Wherein, the time interval between the end of the electronic device 100 receiving the first input and the start of the electronic device 100 receiving the second input is less than the first time length (e.g., 3 seconds), and the first input and the second input belong to the first multi-modal input. Exemplarily, within 3 seconds of the electronic device 100 ending to receive image input of the user circling a video pause frame, if the user starts to perform voice operation on the video pause frame circled by the user, the electronic device 100 can receive voice input of the user performing voice operation on the video pause frame circled by the user, otherwise the electronic device 100 can stop receiving voice input of the user performing voice operation on the video pause frame circled by the user. After receiving the first multi-modal input (e.g., image input of the user circling a video pause frame and voice input of the user performing voice operation on the video pause frame circled by the user), the electronic device 100 can display a user interface 1090 as shown in FIG. 6C.
[0087] Exemplarily, as shown in FIG. 6C, the electronic device 100 can display a user interface 1090, which can include an input box 1091, image content (e.g., image content 1093), wherein the image content 1093 is image content selected by the user as image input, and the input box 1091 can be used to receive voice input (e.g., "what does this talk about, is there a ticket?"). The input box 1091 can include a voice input state icon (e.g., icon 1092), which can be used to identify the state of the electronic device 100 receiving voice input.
[0088] In some embodiments, after receiving the first multi-modal input, the electronic device 100 can search the first multi-modal input. Then, the electronic device 100 can display the search results of searching the first multi-modal input in the form of cards, where the search results can be multiple, and the multiple search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, where the information search results of different information types are displayed in different cards respectively, and the service search results of different service types are displayed in different cards respectively. The user interface in which the electronic device 100 displays the search results of searching the first multi-modal input in the form of cards can be, for example, the user interface 1030 shown in FIG. 2, and the related description of the user interface 1030 can refer to the foregoing content of FIG. 2, which will not be repeated here.
[0089] In some embodiments, after displaying the search results of searching the first multi-modal input in the form of cards, the electronic device 100 can receive a second operation (for example, an operation triggered by the user through a voice instruction or a touch screen operation to share the search results), and in response to the operation, the electronic device 100 can display a second user interface (for example, the user interface 1040) which can be used to select multiple search results. After detecting a third operation (for example, an operation in which the user checks the search results displayed in the form of cards and clicks a sharing control), the electronic device 100 can share the search results selected by the user. The related description of the user interface 1040 can refer to the foregoing content of FIG. 3, which will not be repeated here.
[0090] The working principle of the content generation method provided by the embodiments of the present application will be introduced below.
[0091] FIG. 7 is a schematic diagram of the working principle of the content generation method provided by the embodiments of the present application.
[0092] As shown in FIG. 7, the electronic device 100 can include a search multi-modal Agent, which is a multi-modal system. The electronic device 100 can receive a first multi-modal input through the search multi-modal Agent, then the search multi-modal Agent can generate search results based on the first multi-modal input, and finally the electronic device 100 can output the search results through the search multi-modal Agent.
[0093] The search multi-modal agent can include a multi-modal intent understanding module, a multi-modal cache, a multi-modal agent, and a multi-modal search aggregation module. In the search multi-modal agent, the multi-modal intent understanding module can identify a first user intent based on the first multi-modal input, where the user intent is a structured form of the first multi-modal input, so that the multi-modal agent can understand the user's intent. In some embodiments, the first user intent can be stored in the multi-modal cache, which is the first cache. After identifying the first user intent, the multi-modal intent understanding module can pass the first user intent to the multi-modal agent. After obtaining the first user intent, the multi-modal agent can determine a search strategy for searching the first multi-modal input according to the first user intent and the multi-modal cache, where the multi-modal cache can include different user intents each corresponding to a search result, and the search strategy can be that if the first user intent corresponds to a search result in the multi-modal cache, the search result is the search result corresponding to the first user intent obtained from the multi-modal cache; otherwise, a search engine in the multi-modal search aggregation module is called, and the search result is obtained by searching the first user intent using the search engine. In some embodiments, after determining the search strategy for searching the first multi-modal input, the multi-modal agent can call a first search engine in the multi-modal search aggregation module to search based on the first user intent, and finally obtain a plurality of third search results returned by the multi-modal search aggregation module, where the first search engine is the search engine called by the multi-modal agent to execute the search strategy. In some embodiments, the third search results can be stored in the multi-modal cache. The multi-modal cache can match the first user intent with the plurality of third search results to obtain a fourth search result, where the matching degree of the fourth search result with the first user intent is greater than the matching degree of the third search result with the first user intent. The search result displayed in the form of a card is the fourth search result. In some embodiments, the fourth search result can be stored in the multi-modal cache.
[0094] The multi-modal intention understanding module can include a Query understanding module, a Query correction module, a Query rewriting module, an intention recognition module, an intention understanding module, and an intention distribution module. The Query is a query request in the text input or the voice input, and the intention is a target object in the image input. In the case where the first multi-modal input includes the text input or the voice input, the Query understanding module can understand the semantics of the query request in the text input or the voice input; the Query correction module can correct the query request based on the semantics of the query request, such as correcting the wrong characters of the query request in the text input, correcting the accent or slip of tongue of the query request in the voice input, and the like; the Query rewriting module can rewrite the query request into a structured form to obtain the user intention, so as to be understood by the multi-modal Agent. In the case where the first multi-modal input is the image input, the intention recognition module can recognize the type of the target object in the image input, such as plants, animals, people, goods, and the like; the intention understanding module can understand the overall meaning of the image when the target object is the image; and the intention distribution module can recognize the user intention according to the type or the overall meaning of the target object, and preliminarily determine the search strategy according to the user intention.
[0095] The multi-modal cache can include a Query cache, an intent cache, a multi-turn reference cache, a Promot engine, a Promot cache, and a multi-modal result cache. The Query cache can be used to store query requests and corresponding user intents in different text inputs or voice inputs. The intent cache can be used to store target objects and corresponding user intents in different image inputs. The multi-turn reference cache can be used to multi-turn refine a target object in a user intent into specific manifestations and store the specific manifestations to which the target object is multi-turn refined. For example, the user intent is a query request "how much does this cost" in a picture and voice input of a B style shoe of a brand A. The multi-turn reference cache can first refine the target object in the user intent into a product, secondly refine the target object in the user intent into a shoe, thirdly refine the target object in the user intent into a shoe of the brand A, fourthly refine the target object in the user intent into a B style shoe of the brand A, and finally store the specific manifestations to which the target object is multi-turn refined. The Promot cache can be used to store a plurality of third search results based on a first user intent searched by a first search engine in a multi-modal agent calling a multi-modal search aggregation module. The Promot engine can be used to match the first user intent with the plurality of third search results to obtain a fourth search result. The multi-modal result cache can be used to store search results (i.e., the fourth search result) in a card form. In some embodiments, if the search multi-modal agent determines that the same or similar inputs are received through the Query cache and / or the intent cache, the search multi-modal agent can find the search results corresponding to the same or similar inputs in the multi-modal result cache. The same or similar inputs can refer to a first multi-modal input with the same user intent. For example, the first input is a picture of a B style shoe of a brand A selected by a user and a voice input "how much does this cost", and the second input is a text of a B style shoe of a brand A drawn by a user and a voice input "how much does it cost to purchase". The first input and the second input have the same user intent. After receiving the second input, the search multi-modal agent can find the search results corresponding to the first input in the multi-modal result cache.
[0096] The multi-modal Agent is an intelligent system capable of receiving, processing multi-modal data, making decisions based on a large language model, and finally executing the decision actions. It can receive and process query requests and corresponding user intents in different text inputs or voice inputs, target objects and corresponding user intents in different image inputs, and determine a search strategy for searching the first multi-modal input based on a large language model, and finally execute the search strategy. Specifically, after receiving and processing the first multi-modal input to obtain the first user intent, the multi-modal Agent can determine the search strategy for searching the first multi-modal input according to the first user intent and the multi-modal cache. In the case where the search strategy is to call the search engine in the multi-modal search aggregation module, the multi-modal Agent can call the first search engine in the multi-modal search aggregation module to search based on the first user intent, and finally obtain the multiple third search results returned by the multi-modal search aggregation module.
[0097] The multi-modal search aggregation module can include different types of search engines, such as visual search, text search, and fusion search. Among them, the visual search can search based on images, and the returned search results can be images and the text carried by the images. The visual search can include but is not limited to similar image search, commodity search, and recognition of all things (including animals and plants). The text search can search based on text, and the returned search results can be text, which can carry images (such as vertical domain search). The text search can include but is not limited to vertical domain Box (including person, encyclopedia) search, web page search, and vertical domain (including picture, video) search. The fusion search can search based on images and text. The fusion search can include but is not limited to image-text fusion search. Specifically, the type of search engine called by the multi-modal Agent can be determined according to the first user intent. For example, the Query is "What brand is the hat worn by celebrity A?", the multi-modal intent understanding module can identify that the first user intent is a commodity intent and a person intent from the Query, and then the multi-modal Agent can determine that the type of search engine called is the commodity search in the visual search and the person search in the vertical domain Box in the text search according to the commodity intent and the person intent.
[0098] The current multi-modal system has a small range of content acquisition, including only news, reports, papers, and web texts, which cannot meet the diversified needs of users. Compared with the current multi-modal system, the multi-modal system of the embodiments of the present application can support a larger range of search and content acquisition. In addition to news, reports, papers, and web texts, the search and content acquisition range of the multi-modal system of the embodiments of the present application can also include but is not limited to encyclopedia, audio, video, and website.
[0099] The following introduces a flowchart of a content generation method based on a search multi-modal Agent provided by the embodiments of the present application.
[0100] FIG. 8A illustrates a specific process of a content generation method based on a search multi-modal Agent according to an embodiment of the present application.
[0101] The method can be applied to the electronic device 100.
[0102] As shown in FIG. 8A, the method can include:
[0103] S101, triggering a voice assistant.
[0104] S102, entering a first state.
[0105] Specifically, the electronic device 100 can display a first user interface including text content, image content. After detecting a first operation (for example, detecting an operation of the user long-pressing the power key for 1 second), the electronic device 100 can trigger the voice assistant, so that the user interacts with the search multi-modal Agent through the voice assistant. The first user interface may, for example, be the user interface 1000 shown in FIG. 1A.
[0106] Specifically, after triggering the voice assistant, detecting an operation of the user starting the search multi-modal Agent service, the electronic device 100 can enter a first state, which can be used for the electronic device 100 to receive a first multi-modal input. The first state can mean that the electronic device 100 is in a state of starting the search multi-modal Agent service. Illustratively, detecting an operation of the user double-finger pressing the screen, the electronic device 100 can enter the first state.
[0107] S103a, voice input.
[0108] S104a, requesting to process the voice input.
[0109] Specifically, after entering the first state, the electronic device 100 can directly receive the first multi-modal input, so that the user can perform voice input without other operations. The first multi-modal input can include voice input, which is received through the voice assistant, and the state of the voice input can be identified by a voice input state icon (for example, the icon 1022 shown in FIG. 1C).
[0110] Specifically, after receiving the voice input, the electronic device 100 can request data processing on the voice input. For example, the electronic device 100 can request to understand the semantics of the query request in the voice input, then can correct the query request based on the semantics of the query request (e.g. correct the accent or mispronunciation of the query request in the voice input, etc.), and finally can rewrite the query request into a structured form to obtain the first user intent for the multi-modal Agent to understand. In some embodiments, the electronic device 100 can request to store the first user intent into the multi-modal cache.
[0111] S103b, text input and / or image input.
[0112] S104b, request to process the text input and / or image input.
[0113] Specifically, after entering the first state, the electronic device 100 can directly receive the first multi-modal input, so that the user can perform image input and text input without other operations. The first multi-modal input can also include text input and / or image input. The text input can be text content selected by the user in the first user interface, and the image input can be image content selected by the user in the first user interface, the image content including pictures and videos.
[0114] Specifically, after receiving the text input and / or image input, the electronic device 100 can request data processing on the text input and / or image input. For example, the electronic device 100 can request to understand the semantics of the query request in the text input, then can correct the query request based on the semantics of the query request (e.g. correct the misspelling of the query request in the text input, etc.), and finally can rewrite the query request into a structured form to obtain the first user intent for the multi-modal Agent to understand. For another example, the electronic device 100 can request to identify the type of the target object (e.g. plant, animal, person, commodity, etc.) in the image input, then can understand the overall meaning of the image when the target object is an image, and finally can identify the first user intent according to the type or overall meaning of the target object, and preliminarily determine the search strategy according to the first user intent. In some embodiments, the electronic device 100 can request to store the first user intent into the multi-modal cache.
[0115] The text input and / or image input can include but are not limited to operations such as circling, pointing, checking, highlighting, drawing lines, long pressing, clicking text or image.
[0116] S105, process the input.
[0117] S106, request the multi-modal Agent.
[0118] S107, return the third search result.
[0119] Specifically, after receiving the first multi-modal input, the electronic device 100 can perform data processing on the first multi-modal input. For example, the electronic device 100 can understand the semantics of the query request in the voice input or the text input, and then can correct the query request based on the semantics of the query request (for example, correct the accent or mispronunciation of the query request in the voice input, correct the misspelling of the query request in the text input, and the like), and finally can rewrite the query request into a structured form to obtain the first user intent for the multi-modal Agent to understand. For another example, the electronic device 100 can identify the type of the target object (for example, plant, animal, person, commodity, and the like) in the image input, and then can understand the overall meaning of the image when the target object is an image, and finally can identify the first user intent according to the type or overall meaning of the target object, and preliminarily determine the search strategy according to the first user intent. In some embodiments, the electronic device 100 can request to store the first user intent into the multi-modal cache.
[0120] Specifically, after identifying the first user intent, the electronic device 100 can deliver the first user intent to the multi-modal Agent. After obtaining the first user intent, the multi-modal Agent can determine a search strategy for searching the first multi-modal input according to the first user intent and the multi-modal cache, which can include different user intents each corresponding to a search result, and the search strategy can be specifically that if the first user intent corresponds to a search result in the multi-modal cache, the search result is the search result corresponding to the first user intent obtained from the multi-modal cache; otherwise, a search engine is called, and the search result is obtained by searching the first user intent using the search engine. In some embodiments, after determining the search strategy for searching the first multi-modal input, the multi-modal Agent can call the first search engine to search based on the first user intent, and finally obtain a plurality of third search results, wherein the first search engine is the search engine called by the multi-modal Agent when executing the above-mentioned search strategy.
[0121] S108, processing the third search result.
[0122] S109, returning the processed fourth search result.
[0123] S110, rendering the fourth search result.
[0124] S111, displaying the fourth search result.
[0125] In some embodiments, the electronic device 100 can store the plurality of third search results into the multi-modal cache.
[0126] Specifically, after obtaining the plurality of third search results, the electronic device 100 can match the first user intent with the plurality of third search results to obtain fourth search results, which have a matching degree with the first user intent greater than that of the third search results with the first user intent. The search results displayed in the form of cards are the fourth search results. In some embodiments, the fourth search results can be stored in the multi-modal cache.
[0127] Specifically, after obtaining the fourth search results, the electronic device 100 can render and display the fourth search results (i.e., search results) in the form of cards, where the search results can be multiple, and the multiple search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, where the information search results of different information types are displayed in different cards respectively, and the service search results of different service types are displayed in different cards respectively.
[0128] FIG. 8B exemplarily shows another specific process of a content generation method based on a search multi-modal Agent according to an embodiment of the present application.
[0129] The method can be applied to the electronic device 100.
[0130] As shown in FIG. 8B, the method can include:
[0131] S201, triggering a voice assistant.
[0132] S202, entering a first state.
[0133] Specifically, the electronic device 100 can display a first user interface including text content and image content. After detecting a first operation (e.g., detecting an operation of the user long-pressing the power key for 1 second), the electronic device 100 can trigger the voice assistant to enable the user to interact with the search multi-modal Agent through the voice assistant. The first user interface can be, for example, the user interface 1000 shown in FIG. 1A.
[0134] Specifically, after triggering the voice assistant, the electronic device 100 can enter a first state after detecting an operation of the user starting the search multi-modal Agent service, where the first state can be used for the electronic device 100 to receive a first multi-modal input. The first state can refer to a state in which the electronic device 100 starts the search multi-modal Agent service. Exemplarily, the electronic device 100 can enter the first state after detecting an operation of the user double-tapping the screen.
[0135] S203, receiving a first multi-modal input.
[0136] S204, processing the first multi-modal input.
[0137] Specifically, after entering the first state, the electronic device 100 can directly receive the first multi-modal input, so that the user can make voice input, image input and text input without other operations. The first multi-modal input can include voice input and one or more of the following: text input, image input, wherein the voice input is received through the voice assistant, the state of the voice input can be identified by a voice input state icon (for example, the icon 1022 shown in FIG. 1C), the text input can be text content selected by the user in the first user interface, and the image input can be image content selected by the user in the first user interface, the image content including pictures and videos.
[0138] Specifically, after receiving the first multi-modal input, the electronic device 100 can perform data processing on the first multi-modal input. For example, the electronic device 100 can understand the semantics of the query request in the voice input or the text input, and then can correct the query request based on the semantics of the query request (for example, correcting the accent or mispronunciation of the query request in the voice input, correcting the misspelling of the query request in the text input, etc.), and finally can rewrite the query request into a structured form to obtain the first user intent for the multi-modal Agent to understand. For another example, the electronic device 100 can identify the type of the target object (for example, plant, animal, person, commodity, etc.) in the image input, then can understand the overall meaning of the image when the target object is an image, and finally can identify the first user intent according to the type or overall meaning of the target object, and preliminarily determine the search strategy according to the first user intent. In some embodiments, the electronic device 100 can request to store the first user intent into the multi-modal cache.
[0139] In this way, the multi-modal system of the embodiments of the present application can support other modal inputs of instructions and unstructured instructions such as free text conversation.
[0140] S205, determining a search strategy for searching the first multi-modal input.
[0141] S206, executing the search strategy for searching the first multi-modal input.
[0142] Specifically, after identifying the first user intent, the electronic device 100 can determine, by the multi-modal Agent, a search strategy for searching the first multi-modal input according to the first user intent and a multi-modal cache, the multi-modal cache can include search results corresponding to different user intents respectively, and the search strategy can be specifically as follows: if the first user intent corresponds to a search result in the multi-modal cache, the search result is the search result corresponding to the first user intent obtained from the multi-modal cache; otherwise, a search engine is called, and the search result is obtained by searching the first user intent by using the search engine.
[0143] In some embodiments, after determining the search strategy for searching the first multi-modal input, the electronic device 100 can call a first search engine to search based on the first user intent by the multi-modal Agent, and finally obtain a plurality of third search results, wherein the first search engine is a search engine called by the multi-modal Agent when executing the above-mentioned search strategy.
[0144] S207, return the search result of searching the first multi-modal input.
[0145] S208, display the search result of searching the first multi-modal input in the form of a card.
[0146] S209, select the search result for sharing.
[0147] In some embodiments, the electronic device 100 can store the plurality of third search results in the multi-modal cache.
[0148] Specifically, after obtaining the plurality of third search results, the electronic device 100 can match the first user intent with the plurality of third search results to obtain a fourth search result, the matching degree of the fourth search result with the first user intent is greater than the matching degree of the third search result with the first user intent. Wherein, the search result displayed in the form of a card is the fourth search result. In some embodiments, the fourth search result can be stored in the multi-modal cache.
[0149] Specifically, after obtaining the fourth search result, the electronic device 100 can display the fourth search result (i.e., the search result) in the form of a card, wherein the search result can be multiple, and the plurality of search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, wherein the information search results of different information types are displayed in different cards respectively, and the service search results of different service types are displayed in different cards respectively.
[0150] Specifically, after displaying the search results of the first multi-modal input in the form of cards, detecting a second operation (e.g., the user triggers an operation of sharing the search results through a voice instruction or a touch screen operation), the electronic device 100 can display a second user interface (e.g., the user interface 1040), which can be used to select a plurality of search results. Detecting a third operation (e.g., the user checks the search results displayed in the form of cards and clicks the sharing control), the electronic device 100 can share the search results selected by the user.
[0151] In this way, the user can reduce the number of interactions with the multi-modal system, more conveniently obtain different information and services related to the same multi-modal input, and can select different information and services for sharing according to needs, thereby improving the user experience.
[0152] The following describes a flowchart of a content generation method based on a search strategy provided by an embodiment of the present application.
[0153] FIG. 9 exemplarily shows a specific flow of a content generation method based on a search strategy provided by an embodiment of the present application.
[0154] The method can be applied to the electronic device 100.
[0155] In the embodiment of the present application, after entering the first state, the electronic device 100 can directly receive the first multi-modal input, so that the user can perform voice input, image input and text input without other operations. The first multi-modal input can include voice input and one or more of the following: text input, image input, wherein the voice input is received through a voice assistant, and the state of the voice input can be identified by a voice input state icon (e.g., the icon 1022 shown in FIG. 1C).
[0156] As shown in FIG. 9, the method can include:
[0157] S301, identifying a user intent.
[0158] In some embodiments, after receiving the first multi-modal input, the electronic device 100 can identify the first user intent based on the first multi-modal input. Specifically, in the case that the first multi-modal input includes a text input or a voice input, the electronic device 100 can understand the semantics of the query request in the voice input or the text input, then can correct the query request based on the semantics of the query request (for example, correct the accent or mispronunciation of the query request in the voice input, correct the misspelling of the query request in the text input, and the like), and finally can rewrite the query request into a structured form to obtain the first user intent for the multi-modal Agent to understand. In the case that the first multi-modal input includes an image input, the electronic device 100 can identify the type of the target object in the image input (for example, plant, animal, person, commodity, and the like), then can understand the overall meaning of the image when the target object is an image, and finally can identify the first user intent according to the type or overall meaning of the target object, and determine the search strategy according to the first user intent. In some embodiments, the electronic device 100 can request to store the first user intent into the multi-modal cache. Exemplarily, the first user intent can include but is not limited to a commodity intent, an animal intent, a plant intent, a person intent, and a landmark intent. The above-mentioned first user intent can be further split, for example, the commodity intent can be split into a top, a bottom, an underwear, a shoe, a bag, a digital product, a makeup, a furniture, a medical care, a wine and beverage, an accessory, and the like.
[0159] Exemplarily, the Query is “What is the brand of the hat worn by celebrity A”, the electronic device 100 can identify that the first user intent is a commodity intent and a person intent from the Query.
[0160] S302, determining a search strategy for searching the first multi-modal input.
[0161] S303, executing the search strategy for searching the first multi-modal input.
[0162] Specifically, after identifying the first user intent, the electronic device 100 can determine the search strategy for searching the first multi-modal input according to the first user intent and the multi-modal cache through the multi-modal Agent, the multi-modal cache can include different search results corresponding to respective user intents, and the search strategy can be specifically as follows: if the first user intent corresponds to a search result in the multi-modal cache, the search result is the search result corresponding to the first user intent obtained from the multi-modal cache; otherwise, a search engine is called, and the search result is obtained by searching the first user intent using the search engine.
[0163] In some embodiments, after determining the search strategy for searching the first multimodal input, the electronic device 100 can call a first search engine to search based on the first user intent through the multimodal agent, and finally obtain a plurality of third search results, wherein the first search engine is a search engine called by the multimodal agent in executing the above search strategy.
[0164] For example, the query is "What is the brand of the hat worn by celebrity A", and the multimodal agent can determine that the type of search engine called is the product search in visual search and the celebrity search in the vertical domain box in text search according to the product intent and the celebrity intent.
[0165] S304, cache the search results of searching the first multimodal input.
[0166] S305, return the search results of searching the first multimodal input.
[0167] In some embodiments, the electronic device 100 can store the plurality of third search results in the multimodal cache.
[0168] Specifically, after obtaining the plurality of third search results, the electronic device 100 can match the first user intent with the plurality of third search results to obtain fourth search results, which have a matching degree with the first user intent greater than that of the third search results with the first user intent. The search results displayed in the form of cards are the fourth search results. In some embodiments, the fourth search results can be stored in the multimodal cache.
[0169] The multimodal system of the embodiments of the present application has the ability of autonomous planning, autonomous decision-making and autonomous execution, can autonomously formulate reasonable strategies and plans according to user intent and multimodal cache, reduces the overall time delay, improves work efficiency, and thus improves user experience.
[0170] The following introduces a flowchart of a content generation method based on a search result display strategy provided by the embodiments of the present application.
[0171] FIG. 10 exemplarily shows a specific flow of a content generation method based on a search result display strategy provided by the embodiments of the present application.
[0172] The method can be applied to the electronic device 100, which can include a search multimodal agent.
[0173] In the embodiments of the present application, the electronic device 100 can directly receive the first multi-modal input, so that the user can perform voice input, image input and text input without other operations. The first multi-modal input can include voice input and one or more of the following: text input, image input, wherein the voice input is received through the voice assistant, and the state of the voice input can be identified by a voice input state icon (for example, icon 1022 shown in FIG. 1C). After receiving the first multi-modal input, the electronic device 100 can search the first multi-modal input.
[0174] As shown in FIG. 10, the method can include:
[0175] S401, returning a search result of searching the first multi-modal input and a display strategy of the search result.
[0176] In some embodiments, after searching the first multi-modal input, the electronic device 100 can return the search result of searching the first multi-modal input and the display strategy of the search result to the end side through the search multi-modal Agent.
[0177] Specifically, in the search multi-modal Agent, the multi-modal Agent can determine the strategy of classified display of the search result according to the user intent. For example, the Query is “What is the brand of the hat worn by celebrity A”, the electronic device 100 can identify from the Query that the user intent is a commodity intent and a person intent. Further, the multi-modal Agent can determine to call commodity search in visual search and person search in vertical domain Box in text search according to the commodity intent and the person intent, to obtain search results such as commodity information, person information and shopping service. Since the electronic device 100 identifies two user intents from the Query, the electronic device 100 can finally obtain multiple search results of different types, and the multi-modal Agent can determine to classify and display the search results such as commodity information, person information and shopping service. For another example, the Query is “What is this about? Is there a ticket?”, the electronic device 100 can search to obtain search results such as drama music information, drama video information, ticket purchase service and travel service, and the multi-modal Agent can determine to classify and display the search results such as drama music information, drama video information, ticket purchase service and travel service.
[0178] S402, executing the display strategy of the search result.
[0179] S403, displaying the search result of searching the first multi-modal input in the form of a card.
[0180] S404, selecting the search result for sharing.
[0181] Specifically, after returning the search result of searching the first multi-modal input and a display strategy of the search result to the end side, the electronic device 100 can perform a strategy of search result classification display.
[0182] Specifically, after performing the strategy of search result classification display, the electronic device 100 can display the search result in the form of a card, where the search result can be multiple, and the multiple search results can include information search results and service search results. Specifically, the electronic device 100 can display the information search results and the service search results in different cards respectively, where the information search results of different information types are displayed in different cards respectively, and the service search results of different service types are displayed in different cards respectively.
[0183] Specifically, after displaying the search result of searching the first multi-modal input in the form of a card, detecting a second operation (for example, an operation of sharing the search result triggered by the user through a voice instruction or a touch screen operation), the electronic device 100 can display a second user interface (for example, the user interface 1040), which can be used to select multiple search results. Detecting a third operation (for example, an operation of checking the search result displayed in the form of a card and clicking a sharing control), the electronic device 100 can share the search result selected by the user.
[0184] The flowchart of the content generation method provided by the embodiment of the application is introduced below.
[0185] FIG. 11 exemplarily shows a specific flow of a content generation method provided by the embodiment of the application.
[0186] The method can be applied to a first electronic device, which is the electronic device 100 described above.
[0187] As shown in FIG. 11, the method can include:
[0188] S501, the first electronic device enters a first state.
[0189] The first state can be used for the first electronic device to receive a first multi-modal input.
[0190] In some embodiments, before the first electronic device enters the first state, the first electronic device can display a first user interface including text content and image content, where the text input can be the text content selected by the user in the first user interface, and the image input can be the image content selected by the user in the first user interface. Detecting a first operation, the first electronic device can trigger a voice assistant. The first user interface can be, for example, the user interface 1000 shown in FIG. 1A.
[0191] S502, the first electronic device receives a first multi-modal input.
[0192] Specifically, after entering the first state, the first electronic device can receive the first multi-modal input, which can include a voice input and one or more of the following: a text input, an image input, wherein the voice input is received through a voice assistant.
[0193] Specifically, the first electronic device can receive a first input, a second input. Wherein the time interval between the end of receiving the first input and the start of receiving the second input is less than the first time length, and the first input and the second input both belong to the first multi-modal input.
[0194] S503, the first electronic device searches the first multi-modal input.
[0195] Specifically, after receiving the first multi-modal input, the first electronic device can search based on the first user intent, which is identified by the first electronic device based on the first multi-modal input.
[0196] In some embodiments, the first electronic device can determine a search strategy of the search according to the first user intent and a first cache, the first cache including search results corresponding to different user intents respectively. Wherein the search strategy can specifically include: if the first user intent corresponds to a search result in the first cache, the search result can be the search result corresponding to the first user intent obtained from the first cache; otherwise, a search engine is called, and the search result can be obtained by searching the first user intent using the search engine. Wherein the type of the search engine can include one or more of the following: visual search, text search, fusion search, visual search can search based on images, text search can search based on text, and fusion search can search based on images and text.
[0197] In some embodiments, after determining the search strategy, the first electronic device can call the first search engine to search the first multi-modal input to obtain a plurality of third search results, the first search engine being a search engine called by the first electronic device in executing the search strategy. Then, the first electronic device can match the first user intent with the plurality of third search results to obtain a fourth search result, the matching degree of the fourth search result with the first user intent being greater than the matching degree of the third search result with the first user intent, and the search result displayed in the form of a card being the fourth search result.
[0198] S504, the first electronic device displays the search result of searching the first multi-modal input in the form of a card.
[0199] In some embodiments, in the process of receiving the first multi-modal input, the first electronic device can first present the first search result, and then increase the display of the second search result, where the first search result and the second search result are of different types, and the types of search results can include information search results and service search results.
[0200] In some embodiments, there are multiple search results, and the multiple search results can include information search results and service search results.
[0201] Specifically, after searching the first multi-modal input, the first electronic device can present the information search results and the service search results in different cards respectively. Wherein, information search results of different information types can be respectively presented in different cards, and service search results of different service types can be respectively presented in different cards.
[0202] In some embodiments, after presenting the search results of the first multi-modal input in the form of cards, a second operation is detected, and the first electronic device can display a second user interface, which can be used to select multiple search results. A third operation is detected, and the first electronic device can share the selected search results.
[0203] Finally, an electronic device provided by an embodiment of the present application is introduced.
[0204] FIG. 12 is a structural schematic diagram of an electronic device 100 provided by an embodiment of the present application.
[0205] As shown in FIG. 12, the electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charge management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a camera 193, a display screen 194, and a subscriber identity module (SIM) card interface 195, etc. Wherein, the sensor module 180 can include at least one of a pressure sensor 180A, a proximity light sensor 180G, a touch sensor 180K, etc.
[0206] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0207] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.
[0208] In some embodiments, the electronic device 100 can also include one or more processors 110.
[0209] Among them, the controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of fetching instructions and executing instructions.
[0210] In the embodiments of the present application, the search multi-modal Agent can run on the processor 110, and the search multi-modal Agent can include a multi-modal Agent. The electronic device 100 can perform data processing on the first multi-modal input through the search multi-modal Agent. For example, the electronic device 100 can understand the semantics of the query request in the voice input or the text input, and then can correct the query request based on the semantics of the query request (for example, correcting the accent or slip of tongue of the query request in the voice input, correcting the wrong characters of the query request in the text input, etc.), and finally can rewrite the query request into a structured form to obtain the first user intent for the multi-modal Agent to understand. For another example, the electronic device 100 can identify the type of the target object (for example, plants, animals, people, goods, etc.) in the image input, then can understand the overall meaning of the image when the target object is an image, and finally can identify the first user intent according to the type or overall meaning of the target object, and preliminarily determine the search strategy according to the first user intent.
[0211] In the embodiments of the present application, after the first user intention is identified, the electronic device 100 can determine, by the multi-modal agent, a search strategy for searching the first multi-modal input according to the first user intention and a multi-modal cache, the multi-modal cache can include search results corresponding to different user intentions respectively, and the search strategy can be specifically: if the first user intention corresponds to a search result in the multi-modal cache, the search result is the search result corresponding to the first user intention obtained from the multi-modal cache; otherwise, a search engine is called, and the search result is obtained by searching the first user intention by using the search engine.
[0212] In the embodiments of the present application, after the search strategy for searching the first multi-modal input is determined, the electronic device 100 can call the first search engine to search based on the first user intention by the multi-modal agent, and finally obtain a plurality of third search results, wherein the first search engine is a search engine called by the multi-modal agent when the search strategy is executed.
[0213] In the embodiments of the present application, after the plurality of third search results are obtained, the electronic device 100 can match the first user intention with the plurality of third search results to obtain a fourth search result, the matching degree of the fourth search result with the first user intention is greater than the matching degree of the third search result with the first user intention. The search result displayed in the form of a card is the fourth search result.
[0214] In the embodiments of the present application, after the first multi-modal input is searched, the electronic device 100 can return the search result of the first multi-modal input and a display strategy of the search result to the terminal side by the search multi-modal agent. In the search multi-modal agent, the multi-modal agent can determine a strategy for classified display of the search result according to the user intention. Specifically, the fourth search result (i.e., the search result) can be multiple, and the multiple search results can include information search results and service search results. The electronic device 100 can determine to display the information search results and the service search results in different cards respectively, wherein different information search results of different information types can be displayed in different cards, and different service search results of different service types can be displayed in different cards.
[0215] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the electronic device 100.
[0216] The USB interface 130 is an interface conforming to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and a peripheral device. It can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other electronic devices, such as AR devices, etc.
[0217] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input through a wireless charging coil of the electronic device 100. The charging management module 140 can charge the battery 142 while also providing power to the electronic device 100 through the power management module 141.
[0218] The power management module 141 is configured to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be used to detect battery capacity, battery cycle count, battery health status (leakage, impedance), etc. In other embodiments, the power management module 141 can also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.
[0219] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc. In the embodiments of the present application, the electronic device 100 can use the wireless communication function to search for the first multi-modal input.
[0220] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.
[0221] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the processed signals to the modem processor for demodulation. The mobile communication module 150 can also amplify signals modulated by the modem processor, and radiate the amplified signals as electromagnetic waves via the antenna. In some embodiments, at least part of the functions of the mobile communication module 150 can be provided in the processor 110. In some embodiments, at least part of the functions of the mobile communication module 150 can be provided in the same device as at least part of the processor 110.
[0222] The modem processor can include a modulator and a demodulator. The modulator can modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator can demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator can then transfer the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor can be transferred to the application processor. The application processor can output a sound signal through an audio device (not limited to the speaker 170A, the microphone 170B, etc.), or display an image or a video through the display screen 194. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be provided in the same device as the mobile communication module 150 or other functional modules, independently of the processor 110.
[0223] The wireless communication module 160 can provide a solution for wireless communication including WLAN (e.g., Wi-Fi network), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM), NFC, infrared (IR), ultrawideband (UWB), etc. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 can receive electromagnetic waves via an antenna, perform frequency modulation and filtering on the electromagnetic wave signals, and transmit the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, perform frequency modulation and amplification, and radiate the processed signals as electromagnetic waves via an antenna. For example, the wireless communication module 160 can include a Bluetooth module, a Wi-Fi module, etc.
[0224] In some embodiments, one part of the antennas of the electronic device 100 is coupled with the mobile communication module 150, and another part of the antennas is coupled with the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology.
[0225] The electronic device 100 can implement a display function, such as displaying search results in a card form, through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected with the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute instructions to generate or change display information.
[0226] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include one or N display screens 194, and N is a positive integer greater than 1.
[0227] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor, etc. The ISP is used to process data fed back by the camera 193. The camera 193 is used to capture still images or videos. The digital signal processor is used to process digital signals, which can process not only digital image signals but also other digital signals. The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. The NPU is a neural-network (NN) computing processor that processes input information quickly by drawing on the structure of biological neural networks, such as the transmission mode between human brain neurons, and can also constantly self-learn.
[0228] The external memory interface 120 can be configured to connect an external memory card, such as a Micro SD card, to extend the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, music, photos, videos, and the like are stored in the external memory card.
[0229] The internal memory 121 can be configured to store one or more computer programs including instructions. The processor 110 can execute the above-mentioned instructions stored in the internal memory 121 to cause the electronic device 100 to perform the method of generating content provided in some embodiments of the present application, various functional applications, and data processing, and the like. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, and the program storage area can also store one or more application programs (such as a gallery, contacts, and the like). The data storage area can store data created during use of the electronic device 100 (such as photos, contacts, and the like). In addition, the internal memory 121 can include a high-speed random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like. In embodiments of the present application, the electronic device 100 can store different user intents and different search results (including third search results and fourth search results) corresponding to the different user intents through the internal memory 121.
[0230] The electronic device 100 can implement audio functions through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, an application processor, and the like. For example, music playback, recording, and the like.
[0231] The audio module 170 is configured to convert digital audio information into an analog audio signal output and to convert analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110. In embodiments of the present application, the electronic device 100 can trigger a voice assistant and receive a voice input through the audio module 170. The electronic device 100 can also detect a sharing operation of the search result by the user (for example, the user triggers the operation of sharing the search result through a voice instruction) through the audio module 170.
[0232] The speaker 170A, also called "loudspeaker", is used to convert an audio electrical signal into an acoustic signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0233] The receiver 170B, also called "earpiece", is used to convert an audio electrical signal into an acoustic signal. When the electronic device 100 answers a call or a voice message, the receiver 170B can be held close to a human ear to listen to the voice.
[0234] The microphone 170C, also called "microphone", "sound collector", is used to convert an acoustic signal into an electrical signal. When making a call or sending a voice message, a user can speak into the microphone 170C through the human mouth to input an acoustic signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, in addition to collecting acoustic signals, noise reduction functions can also be realized. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, in addition to collecting acoustic signals, noise reduction, identifying the source of the sound, and realizing directional recording functions, etc.
[0235] The earphone interface 170D is used to connect a wired earphone. The earphone interface 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0236] The pressure sensor 180A is configured to sense a pressure signal and convert the pressure signal to an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display 194. The pressure sensor 180A can be of various types, such as a resistive pressure sensor, an inductive pressure sensor, a capacitive pressure sensor, etc. The capacitive pressure sensor can include at least two parallel plates of conductive material. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 can determine the intensity of the force based on the change in capacitance. When a touch operation is applied to the display 194, the electronic device 100 can detect the intensity of the touch operation based on the pressure sensor 180A. The electronic device 100 can also calculate the location of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations applied to the same touch location but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than a first pressure threshold is applied to a short message application icon, an instruction to view short messages is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold is applied to the short message application icon, an instruction to create a new short message is executed. In embodiments of the present application, the electronic device 100 can detect image input and text input operations (e.g., operations of circling, pointing, checking, highlighting, drawing a line, long pressing, clicking text or images) through the pressure sensor 180A. For example, the electronic device 100 can detect text input of a user long pressing selected text through the pressure sensor 180A.
[0237] The proximity light sensor 180G can include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode can be an infrared light-emitting diode. The electronic device 100 can emit infrared light outwardly through the light-emitting diode. The electronic device 100 can detect infrared reflected light from nearby objects using the photodiode. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that a user is holding the electronic device 100 close to the ear for a call, so as to automatically turn off the screen to achieve the purpose of power saving. The proximity light sensor 180G can also be used for automatic unlocking and locking in a holster mode or a pocket mode. In embodiments of the present application, the electronic device 100 can detect image input and text input operations (e.g., operations of pointing to text or images) through the proximity light sensor 180G. For example, the electronic device 100 can detect image input of a user pointing to a picture through the proximity light sensor 180G.
[0238] The touch sensor 180K, also referred to as a touch panel or a touch-sensitive surface, can be disposed on the display screen 194, forming a touch screen, also referred to as a touch screen display, with the display screen 194. The touch sensor 180K is configured to detect touch operations performed on or near the touch sensor 180K. The touch sensor 180K can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a location different from the location of the display screen 194. In embodiments of the present application, the electronic device 100 can detect image input and text input operations (e.g., circle selection, pointing, checkmarking, highlighting, line drawing, long pressing, clicking text or image operations) through the touch sensor 180K. For example, the electronic device 100 can detect a user's image input of circling a picture through the touch sensor 180K. In embodiments of the present application, the electronic device 100 can detect a user's operation of starting a search multi-modal Agent service (e.g., detecting a user's operation of double-tapping the screen) through the touch sensor 180K. The electronic device 100 can also detect a user's sharing operation on the search results (e.g., a user's operation of triggering sharing of the search results through a touch screen operation) through the touch sensor 180K. The electronic device 100 can also detect a user's operation of selecting a search result for sharing (e.g., a user's operation of checking a search result displayed in a card form and clicking a sharing control) through the touch sensor 180K.
[0239] The SIM card interface 195 is configured to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195, achieving contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. The same SIM card interface 195 can simultaneously insert multiple cards. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external storage cards. The electronic device 100 can interact with a network through the SIM card, achieving functions such as calling and data communication. In some embodiments, the electronic device 100 can use an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0240] In the embodiments of the present application, the device type of the electronic device 100 can be any one of a mobile phone, a tablet computer, a handheld computer, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like. The embodiments of the present application do not limit the specific type of the electronic device 100.
[0241] On the basis of the embodiments of the present application, the electronic device 100 can be designed to update the multi-modal Agent generated content in different forms through specific interaction operations, including but not limited to voice input (for example, voice operation on a user-selected object) and one or more of the following: text input, image input (including but not limited to operations such as circling, pointing, checking, highlighting, drawing a line, long pressing, clicking a text or an image), and the like. The electronic device 100 can be designed to display one or more output results for the same input in the form of a card, and the user can select the output result for sharing.
[0242] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps performed by the electronic device in the above method embodiments, or the steps performed by the human-computer interaction module and the computing module, can be implemented.
[0243] The embodiments of the present application also provide a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps performed by the electronic device in the above method embodiments.
[0244] The embodiments of the present application also provide a chip system. The chip system includes a processor coupled with a memory. The processor executes a computer program stored in the memory to implement the steps performed by the electronic device in any method embodiment of the present application. The chip system can be a single chip or a chip module composed of multiple chips.
[0245] The term "user interface (UI), interface" in the specification and drawings of the present application is a medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface of the application is the source code written by a specific computer language such as Java, extensible markup language (XML), etc. The interface source code is parsed, rendered on the terminal device, and finally presented as content that can be recognized by the user, such as pictures, texts, button controls, etc. The control (control) is also called a widget, which is the basic element of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scrollbars, pictures, and texts. The properties and contents of the controls in the interface are defined by tags or nodes, such as XML <textview> 、 <imgview> 、 <videoview>The interface is defined by nodes that specify the controls contained in the interface. One node corresponds to one control or property in the interface, and the nodes are parsed and rendered to present the content visible to the user. In addition, many applications, such as hybrid applications, also contain web pages in the interface. A web page, also referred to as a page, can be understood as a special control embedded in the interface of an application. The web page is a source code written in a specific computer language, such as hyper text markup language (HTML), cascading style sheets (CSS), JavaScript (JS), etc. The web page source code can be loaded and displayed by a browser or a web page display component similar to the function of a browser to present content recognizable to the user. The specific content contained in the web page is also defined by tags or nodes in the web page source code, such as HTML defines a page by 、 、 <video> 、 <canvas>to define the elements and attributes of the web page.
[0246] A common form of user interface is a graphic user interface (GUI), which refers to a user interface that displays in a graphical manner. It can be an icon, window, control, etc. interface element displayed in the display screen of an electronic device, wherein the control can include an icon, button, menu, tab, text box, dialog box, status bar, navigation bar, Widget, etc. visual interface element.
[0247] The above-described embodiments are merely intended for describing the technical solutions of the present application, but not to limit the present application; even though the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features therein; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0248] It should be understood that, in various embodiments of the present application, the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0249] In the above-described embodiments, according to the context, the term "when" can be interpreted to mean "if" or "after" or "in response to determining" or "in response to detecting". Similarly, according to the context, the phrase "upon determining" or "if detecting (the stated condition or event)" can be interpreted to mean "if determining" or "in response to determining" or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)".
[0250] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk) and the like.
[0251] Those of ordinary skill in the art understand that all or part of the processes in the above embodiments can be implemented by a computer program to instruct the relevant hardware, which can be stored in a computer readable storage medium. The program can include the processes of the above method embodiments when executed. The aforementioned storage medium includes ROM or random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.< / canvas> < / video> < / videoview> < / imgview> < / textview>
Claims
1. A content generation method, characterized in that, The method includes: The first electronic device enters a first state, which is used for the first electronic device to receive a first multimodal input; The first electronic device receives the first multimodal input, which includes voice input and one or more of the following: text input and image input; The first electronic device searches for the first multimodal input; The first electronic device displays the search results for the first multimodal input in the form of cards.
2. The method according to claim 1, characterized in that, Before the first electronic device enters the first state, the method further includes: The first electronic device displays a first user interface, which includes text content and image content; wherein, the text input is text content selected by the user in the first user interface; and the image input is image content selected by the user in the first user interface. Upon detecting the first operation, the first electronic device triggers the voice assistant.
3. The method according to claim 1 or 2, characterized in that, The first electronic device receives the first multimodal input, specifically including: The first electronic device receives a first input and a second input; wherein the time interval between the first electronic device ending the reception of the first input and starting the reception of the second input is less than a first duration; the first input and the second input belong to the first multimodal input.
4. The method according to any one of claims 1-3, characterized in that, The first electronic device displays the search results for the first multimodal input in card format, specifically including: During the process of receiving the first multimodal input, the first electronic device first displays a first search result, and then adds a second search result; the first search result and the second search result are of different types, including information search results and service search results.
5. The method according to any one of claims 1-4, characterized in that, There are multiple search results, including information search results and service search results.
6. The method according to claim 5, characterized in that, The display of search results for the first multimodal input in card format specifically includes: The first electronic device displays the information search results and the service search results in different cards respectively; The search results for different information types are displayed in different cards. The search results for different service types are displayed in different cards.
7. The method according to claim 5 or 6, characterized in that, The method further includes: Upon detecting the second operation, the first electronic device displays a second user interface for selecting from the plurality of search results; Upon detecting a third action, the first electronic device shares the selected search results.
8. The method according to any one of claims 1-7, characterized in that, The first electronic device searches for the first multimodal input, specifically including: The first electronic device performs the search based on a first user intent, which is identified by the first electronic device based on the first multimodal input.
9. The method according to claim 8, characterized in that, The method further includes: The first electronic device determines the search strategy based on the first user intent and the first cache, wherein the first cache includes search results corresponding to different user intents; The search strategy specifically includes: If the first user intent corresponds to a search result in the first cache, then the search result is the search result corresponding to the first user intent obtained from the first cache; otherwise, the search engine is invoked, and the search result is obtained by searching the first user intent using the search engine.
10. The method according to claim 9, characterized in that, The search engine type includes one or more of the following: visual search, text search, and fusion search; wherein, the visual search is based on images; the text search is based on text; and the fusion search is based on both images and text.
11. The method according to claim 9 or 10, characterized in that, The first electronic device searches for the first multimodal input, specifically including: The first electronic device calls a first search engine to search the first multimodal input and obtains multiple third search results. The first search engine is the search engine called by the first electronic device when executing the search strategy. The first electronic device matches the first user intent with the plurality of third search results to obtain a fourth search result. The degree of matching between the fourth search result and the first user intent is greater than the degree of matching between the third search results and the first user intent. The search result displayed in the form of a card is the fourth search result.
12. An electronic device, characterized in that, The electronic device includes a processor and a memory; wherein the memory is coupled to the processor and is used to store a computer program that, when executed by the processor, causes the electronic device to perform the method as described in any one of claims 1-11.
13. A computer storage medium, characterized in that, The computer storage medium stores a computer program that, when executed by a processor, causes the electronic device to perform the method as described in any one of claims 1-11.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it causes the electronic device to perform the method as described in any one of claims 1-11.
Citation Information
Patent Citations
Multi-modal image recognition search method, device and equipment and storage medium
CN112579868A
Multi-modal search method and device, equipment, storage medium and program product
CN113656546A
Information search method and device, equipment and medium
CN114969524A
Content generation method and electronic equipment
CN119537660A
Generating unified embeddings from multi-modal canvas inputs for image retrieval
US20230419571A1