Content display method, electronic device and computer-readable storage medium

By collaboratively displaying multimodal information through electronic devices and servers, the problem of user comprehension difficulties in existing technologies has been solved, enabling timely, accurate, and flexible information display and improving user experience.

WO2025260707A1PCT designated stage Publication Date: 2025-12-26HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/070255
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2025-01-02
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies cannot effectively present content to users in a multimodal manner, leading to inconvenience and misunderstandings for users when understanding information displayed on electronic devices.

Method used

By working together with electronic devices and servers, second-modal information, such as images, audio, and video, can be flexibly displayed based on the first-modal information viewed by the user, ensuring timely, accurate, and flexible display of information to adapt to different devices and scenarios.

Benefits of technology

It improves the efficiency of users' understanding of information displayed on electronic devices, enhances the user experience, and meets users' diverse information acquisition needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070255_26122025_PF_FP_ABST
    Figure CN2025070255_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of terminals. Disclosed are a content display method, an electronic device and a computer-readable storage medium, which can flexibly, accurately and timely display content to users in a multi-modal form, thereby improving use experience of users. In the present application, on the basis of a user browsing first content indicated by information of a first modality, an electronic device can accurately and timely display to the user second content related to the first content in a second modality, such as text, picture, audio, video, vibration, indicator light and AR / VR image. For example, on the basis of the detected change in target content within the first content that the user focuses on, the electronic device can determine and display related information in the second modality, such as picture, audio, video, vibration, indicator light, AR / VR image, etc., so as to assist the user in accurately, fully and timely understanding the first content.
Need to check novelty before this filing date? Find Prior Art

Description

Content display methods, electronic devices, and computer-readable storage media

[0001] This application claims priority to Chinese Patent Application No. 202410783845.6, filed on June 17, 2024, entitled "Method for Displaying Content, Electronic Device and Computer-Readable Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of terminal technology, and in particular to a method for displaying content, an electronic device, and a computer-readable storage medium. Background Technology

[0003] Currently, content generation technology based on generative pre-trained models is widely used in scenarios such as question answering and content creation. Taking the chat generative pre-trained transformer (ChatGPT) as an example, ChatGPT is beginning to be widely used due to its realistic natural language interaction and multi-scenario content generation capabilities.

[0004] With the diversification of user needs, pure text-based automatic content generation technology can no longer meet these needs. People have begun to explore technical solutions that can display content in multiple modal formats, such as text, images, and audio. Therefore, how to present content to users through multimodal information is a problem that needs to be solved. Summary of the Invention

[0005] This application provides a content display method, an electronic device, and a computer-readable storage medium, which can flexibly, accurately, and timely display content to users in a multimodal format, thereby improving the user experience.

[0006] To achieve the above objectives, this application adopts the following technical solution:

[0007] In a first aspect, a method for displaying content is provided, the method comprising: an electronic device acquiring information of a first modality, the information of the first modality being used to indicate first content; the electronic device displaying the first content in the first modality; the electronic device determining target content of the first content that is of interest to the user; the electronic device acquiring information of a second modality corresponding to the target content, the information of the second modality being used to indicate second content; and the electronic device performing a first operation to display the second content to the user in the second modality and / or a third modality.

[0008] As an example, the modality type of the first, second, or third modality includes any one or more of the following: image modality, video modality, audio modality, vibration modality, indicator light modality, and augmented reality (AR) / virtual reality (VR) image modality.

[0009] The solution provided in the first aspect above allows the electronic device to accurately and promptly display other modal information closely related to the target content that the user is interested in within the first content, based on the user's browsing of the first content. Furthermore, the second content can change accordingly as the target content that the user is interested in changes. Therefore, it can assist the user in accurately, fully, and promptly understanding the target content that the user is interested in, thereby improving the user experience.

[0010] As one possible implementation, the aforementioned electronic device performs a first operation to display second content to a user in a second and / or third modality, including: the electronic device displays the second content to the user in a second modality via its display component; and / or, the electronic device displays the second content to the user in a second and / or third modality via the display component of a cooperating device. In this way, based on different content display architectures, such as architectures including or excluding cooperating devices, the device used to display the second content can be flexibly selected according to the actual situation, and the modal type of the displayed second content can be flexibly determined. This not only allows for accurate and timely display of other modal information closely related to the target content to the user, improving the user experience, but also enhances the flexibility and applicability of the solution under different architectures and / or scenarios.

[0011] As one possible implementation, the aforementioned electronic device displays second content in a second and / or third modality through the display component of a cooperating device. This includes: the electronic device sending second-modality information to the cooperating device to display the second content in the second modality through the display component of the cooperating device; or, the electronic device converting the second-modality information into third-modality information and then sending the third-modality information to the cooperating device to display the second content in the third modality through the display component of the cooperating device; or, the electronic device sending second-modality information to the cooperating device to display the second content in the third modality through the display component of the cooperating device. This not only allows for accurate and timely display of other modal information closely related to the target content to the user, improving the user experience, but also enhances the flexibility and applicability of the solution under different architectures and / or scenarios.

[0012] As one possible implementation, the device information of the aforementioned collaborative device is matched with the second and / or third modality. This not only ensures that users accurately, fully, and promptly understand the information displayed on the electronic device, but also achieves optimal display of the second content, enhancing the user's viewing experience.

[0013] As one possible implementation, the information in the first modality is determined based on the first information input by the user. This allows for the display of first content to the user based on the first information input, and the accurate and timely display of other modal information closely related to the target content within that first content, based on the target content that the user is interested in. Furthermore, the second content can change accordingly as the target content changes, assisting the user in accurately, fully, and promptly understanding the target content and improving the user experience.

[0014] As an example, the first information includes any one or more of the following: images, videos, music, keywords; the first modality information includes any one or more of the following related to the first information: introduction, analysis, explanation; and / or, the first information includes a question; the first modality information includes any one or more of the following related to the question: analysis, explanation, answer, or response.

[0015] As one possible implementation, obtaining the first modality information includes: generating, searching, receiving from or reading pre-stored first modality information from a server. Thus, the first modality information can be obtained through a variety of methods, improving the flexibility and applicability of this solution.

[0016] As one possible implementation, the target content that the user is interested in in the first part is determined by the electronic device using at least one of the following methods: detecting the user's eye focus position; receiving the user's operation on the first content; detecting the position of the user's finger when swiping; and detecting content in the center of the screen. In this way, the target content that the user is interested in can be determined through a variety of active detection methods or passive reception methods, improving the accuracy of target content determination while also enhancing the flexibility and applicability of this solution.

[0017] As one possible implementation, the method further includes: the electronic device acquiring information about a second modality related to the first content; the acquisition of information about a second modality corresponding to the target content by the electronic device includes: the electronic device acquiring information about the second modality corresponding to the target content from the information about the second modality related to the first content. This can improve the timeliness of the second content changing accordingly to changes in the target content that the user is focusing on, thereby improving the user experience.

[0018] As one possible implementation, the electronic device acquires second modal information related to the first content by: generating, searching, receiving from, or reading pre-stored second modal information related to the first content from a server. Thus, second modal information can be obtained through a variety of acquisition methods, improving the flexibility and applicability of this solution.

[0019] As one possible implementation, the aforementioned second modal information related to the first content is obtained by the electronic device through the following process: the electronic device performs semantic segmentation on the first content to obtain multiple sub-contents; the electronic device identifies one or more first sub-contents that require supplementary information from other modalities; the electronic device obtains the second modal information corresponding to the one or more first sub-contents; wherein the target content overlaps with at least one character in the one or more first sub-contents. In this way, the obtained second modal information is at the sub-content granularity and has a high degree of relevance to the first content, thus accurately and fully assisting the user in understanding the corresponding word content and contextual meaning, improving the user experience.

[0020] As one possible implementation, the aforementioned electronic device acquires the second modality information corresponding to the target content, including: generating, searching, obtaining from a server, or retrieving pre-stored second modality information. Thus, the second modality information corresponding to the target content can be obtained through a variety of acquisition methods, improving the flexibility and applicability of this solution.

[0021] As one possible implementation, the aforementioned electronic device displays the second content in a second mode, including displaying the second content in a fixed display area or a non-fixed display area. This allows for flexible adjustment of the display area when displaying the second content, thereby improving the display effect. For example, the display area can be flexibly adjusted according to the device type, screen status, display screen size, interface layout, device functions, etc.

[0022] As one possible implementation, the screen of the aforementioned electronic device is a foldable screen. The electronic device displays second content in a second mode, including: displaying the second content in a non-fixed display area when the screen of the electronic device is in a folded state; and displaying the second content in a fixed display area when the screen of the electronic device is in an unfolded state. In this way, the display area of ​​the second content can be flexibly adjusted according to the screen state of the device to improve the display effect of the second content.

[0023] As one possible implementation, displaying the second content in the second modality includes: displaying the second content in a floating mode, wherein the second content does not obscure the target content; or, displaying the second content in a tiled mode, wherein the second content does not obscure the first content. This allows for flexible adjustment of the display method when the second content is displayed, thereby improving the display effect of both the first and second content.

[0024] As one possible implementation, the screen of the aforementioned electronic device is a foldable screen. The electronic device displays the second content in a second mode, including: when the screen of the electronic device is in a folded state, the electronic device displays the second content in a floating display mode; when the screen of the electronic device is in an unfolded state, the electronic device displays the second content in a tiled display mode. In this way, the display format of the second content can be flexibly adjusted according to the screen state of the device to improve the display effect of the second content.

[0025] As one possible implementation, the above method further includes: the electronic device, in response to a user's operation on the second content, displaying information in a fourth modality related to sub-content of the first content other than the target content, or displaying sub-content at a first position in the first content related to the second content. This can meet diverse user needs, such as improving the speed at which users acquire information. For example, taking a text modality as the first modality, if a user wants to primarily view information from other modalities besides the first modality, supplemented by information from the first modality for understanding, the user can directly operate the second content, such as viewing information from other modalities besides the first modality sequentially by operating the second content, thus saving the time required to read text information. Similarly, if a user wants to view the text content corresponding to the second content, the user can directly operate the second content, such as viewing the text content at the previous position, the next position, or all text content in the first content corresponding to the second content sequentially by operating the second content, in order to quickly obtain all information related to the first content. This can further improve the user's experience when acquiring information.

[0026] In a second aspect, a method for displaying content is provided, the method comprising: a server receiving first information from an electronic device; the server obtaining first modal information based on the first information, wherein the first modal information is used to indicate first content; the server obtaining second modal information based on the first modal information, wherein the second modal information is related to a portion of the first content; the server sending the first modal information to the electronic device, and sending the second modal information to the electronic device and / or a cooperating device.

[0027] As an example, the first information includes any one or more of the following: images, videos, music, keywords; the first modality information includes any one or more of the following related to the first information: introduction, analysis, explanation; and / or, the first information includes a question; the first modality information includes any one or more of the following related to the question: analysis, explanation, answer, or response.

[0028] As an example, the modal type of the first, second, or third modality includes any one or more of the following: image modality, video modality, audio modality, vibration modality, indicator light modality, and AR / VR image modality.

[0029] The solution provided in the second aspect above allows the server to obtain first content related to first information from an electronic device, as well as second content related to the first content. This enables the electronic device to accurately and promptly display second content in other modalities closely related to the target content based on the target content that the user is interested in within the first content when the user browses it. This helps the user to accurately, fully, and promptly understand the target content that the user is interested in, thereby improving the user experience.

[0030] As one possible implementation, the server obtains the information of the first modality based on the first information, including: the server generating, searching, and reading pre-stored information of the first modality based on the first information. In this way, the information of the first modality can be obtained through a variety of acquisition methods, improving the flexibility and applicability of this solution.

[0031] As one possible implementation, the server obtains second-modal information based on first-modal information, including: the server performing semantic segmentation on the first content indicated by the first-modal information to obtain multiple short sentences; the server identifying one or more first short sentences that require supplementation with information from other modalities; and the server generating, searching, or retrieving pre-stored second-modal information related to the one or more first short sentences. Thus, the obtained second-modal information is at the sub-content granularity and has a high degree of relevance to the first content, thereby accurately and fully assisting users in understanding the corresponding word content and contextual meaning, improving the user experience.

[0032] Thirdly, an electronic device is provided, comprising: a display screen for displaying an interface; a memory for storing computer program instructions; and a processor for executing the computer program instructions to support the electronic device in implementing the methods of any possible implementation of the first aspect.

[0033] Fourthly, a server is provided, comprising: a memory for storing computer program instructions; and a processor for executing the computer program instructions to support the server in implementing the methods as described in any possible implementation of the second aspect.

[0034] Fifthly, a content display system is provided, comprising: an electronic device for implementing the method as in any possible implementation of the first aspect; and a server for implementing the method as in any possible implementation of the second aspect.

[0035] As one possible implementation, the content display system described above also includes: one or more collaborative devices for displaying second content related to the first content while the electronic device is displaying the first content.

[0036] In a sixth aspect, a computer-readable storage medium is provided that stores computer program instructions that, when executed by a processor, implement the method as described in any possible implementation of the first or second aspect.

[0037] In a seventh aspect, a computer program product comprising instructions is provided, which, when run on a computer, causes the computer to implement the method as described in any possible implementation of the first or second aspect.

[0038] Eighthly, a chip system is provided, comprising processing circuitry and a storage medium storing computer program instructions; when executed by the processor, the computer program instructions implement the method as described in any possible implementation of the first or second aspect. The chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description

[0039] Figure 1 is an example of content display using multimodal methods;

[0040] Figure 2 shows another example of content display using multimodal methods;

[0041] Figure 3 is a schematic diagram of three multimodal content display architectures provided in the embodiments of this application;

[0042] Figure 4 is a schematic diagram of two multimodal content display architectures provided in the embodiments of this application;

[0043] Figure 5A is a schematic diagram illustrating three scenarios in which a user is determining browsing the first content through active detection methods, as provided in the embodiments of this application.

[0044] Figure 5B is a schematic diagram illustrating a scenario where a user's browsing of first content is determined based on the user's selection operation, according to an embodiment of this application.

[0045] Figure 6 is a schematic diagram of another multimodal content display architecture provided in an embodiment of this application;

[0046] Figure 7 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0047] Figure 8 is a flowchart illustrating the method for displaying the content provided in the embodiments of this application;

[0048] Figure 9 is a schematic diagram illustrating the process of displaying the content provided in the embodiments of this application;

[0049] Figure 10 is a schematic diagram of the second process of displaying the content provided in the embodiment of this application;

[0050] Figure 11 is a schematic diagram of an interface for displaying multimodal information in an electronic device according to an embodiment of this application;

[0051] Figure 12 is a schematic diagram of the interface for displaying multimodal information in another electronic device provided in an embodiment of this application;

[0052] Figure 13 is a schematic diagram of a scenario in which second content is displayed through a collaborative device according to an embodiment of this application;

[0053] Figure 14 is a schematic diagram of another scenario for displaying second content through a collaborative device, provided by an embodiment of this application;

[0054] Figure 15 is a schematic diagram of the electronic device provided in this application displaying a content interface based on the user's second action;

[0055] Figure 16 is a schematic diagram of the electronic device provided in this application displaying a content interface based on the user's second action;

[0056] Figure 17 is a schematic diagram of the electronic device provided in this application displaying a content interface based on the user's second action;

[0057] Figure 18 is a flowchart of an embodiment of this application for obtaining multimodal information based on first information input by a user;

[0058] Figure 19 is a schematic diagram of a content acquisition and display process provided in an embodiment of this application;

[0059] Figure 20 is a schematic diagram of another content acquisition and display process provided in an embodiment of this application. Detailed Implementation

[0060] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0061] In the following text, the terms "first," "second," etc., are used only to distinguish different descriptive objects and do not limit the position, order, priority, quantity, or content of the described objects. For example, if the described object is a "field," then the ordinal numbers before "field" in "first field" and "second field" do not limit the position or order of the "fields." "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the described object is a "level," then the ordinal numbers before "level" in "first level" and "second level" do not limit the priority of the "levels." Furthermore, the quantity of described objects is not limited by ordinal numbers and can be one or more; for example, in "first device," the number of "devices" can be one or more. In addition, objects modified by different prefixes can be the same or different. For example, if the described object is "device," then "first device" and "second device" can be devices of the same type or different types. Similarly, if the described object is "information," then "first information" and "second information" can be information with the same content or information with different content. In summary, the use of ordinal numbers and other prefixes used to distinguish the described objects in the embodiments of this application does not constitute a limitation on the described objects. The description of the described objects is given in the claims or the context of the embodiments, and the use of such prefixes should not constitute an unnecessary limitation.

[0062] Furthermore, in the embodiments of this application, "connection" can be a direct connection or an indirect connection; in addition, it can refer to an electrical connection or a communication connection; for example, the connection of two electrical components A and B can refer to A and B being directly connected, or it can refer to A and B being indirectly connected through other electrical components or connection media, or it can refer to A and B being indirectly connected through other communication devices or communication media, as long as it enables communication between A and B.

[0063] Currently, in order to meet the diverse needs of users and facilitate their accurate, comprehensive, and timely understanding of the information displayed on electronic devices, automatic content generation technology is exploring the use of multimodal formats to present content to users in order to improve their user experience.

[0064] As an example, please refer to Figure 1, which illustrates an example of content display using multimodal methods. As shown in Figure 1, when a user inputs the question "Top 10 must-see attractions in City A," the electronic device generates and displays a ranking and related information about the top 10 attractions in City A based on automatic content generation technology. For example, as shown in Figure 1, the electronic device can display the initial content of the ranking and related information about the top 10 attractions in City A to the user through interface 100, and then display subsequent content through interface 102 based on the user's swipe operation 101 on interface 100. As shown in Figure 1, the electronic device includes a scenic image of attraction 1 when displaying the ranking and related information about the top 10 attractions in City A, so that the user can further understand attraction 1 through the scenic image. However, in the example shown in Figure 1, the scenic image shown in Figure 1 is strongly related to attraction 1, but it is displayed in the related information section of attraction 10. Therefore, when the user browses the text description of attraction 1, they cannot see the scenic image of attraction 1, and when the user browses the scenic image of attraction 1, they cannot see the text description of attraction 1. In other words, in the example shown in Figure 1, the relationship between the image and the text is not close enough, which will cause inconvenience to the user and is easy to cause misunderstanding.

[0065] As another example, please refer to Figure 2, which shows another example of content display using multimodal methods. As shown in Figure 2, when a user inputs the question "Explanation of the chicken-and-rabbit problem," the electronic device generates and displays the solution method and process for "chicken-and-rabbit problem" based on automatic content generation technology. For example, as shown in Figure 2, the electronic device can display the initial part of the solution method and process for "chicken-and-rabbit problem" to the user through interface 200, and display subsequent content through interface 202 based on the user's swipe operation 201 on interface 200. As shown in Figure 2, the electronic device includes a mathematical diagram when displaying the solution method and process for "chicken-and-rabbit problem," so that the user can fully understand the solution method and process of "chicken-and-rabbit problem" by combining the mathematical diagram with the text answer. However, in the example shown in Figure 2, the electronic device displays the mathematical diagram at the first step text position and continues to reference the mathematical diagram at the third step text position. Therefore, when the user scrolls to the third step, they cannot see the mathematical diagram; if the user wants to view the mathematical diagram, they need to scroll back through the interface. In other words, in the example shown in Figure 2, users cannot understand the content by combining it with mathematical diagrams while the text is scrolling, which will cause inconvenience to users and affect the user experience.

[0066] To address the aforementioned shortcomings of conventional multimodal content display methods, this application provides a content display method. This method can display information from one or more other modalities (such as information from the second modality) based on the user's browsing of information in the first modality, such as related text, images, audio, video, vibration, indicator lights, augmented reality (AR) / virtual reality (VR) images, etc., based on the target content that the user is interested in within the information in the first modality. This helps the user to accurately, fully, and promptly understand the information in the first modality displayed by the electronic device, thereby improving the user experience.

[0067] For example, taking the first modality of information as text content, based on the solution provided in the embodiments of this application, the electronic device can accurately and promptly display second modality information such as pictures, audio, video, vibration, indicator lights, AR / VR images, etc., related to the target text, according to the user's browsing of the text content (i.e., the first modality of information), such as the target text that the user is interested in. This helps the user to accurately, fully, and promptly understand the information in the text content displayed by the electronic device, thereby improving the user experience.

[0068] In some embodiments of this application, when displaying information in the second modality to a user, the electronic device can flexibly and adaptively adjust the display area and / or display method of the information in the second modality according to the device type, screen status, display screen size, interface layout, device function, etc., so as to ensure that the user can accurately, fully and timely understand the information in the second modality displayed by the electronic device, while improving the display effect of the information in the second modality and enhancing the user's viewing experience.

[0069] In some embodiments of this application, when displaying information of the second modality to a user, the electronic device can distribute the information of the second modality to the corresponding collaborative devices for display based on device information such as device type and device function of one or more collaborative devices, thereby improving the display effect of the information of the second modality and making it easier for the user to accurately, fully and timely understand the information of the first modality displayed by the electronic device.

[0070] In some embodiments of this application, when displaying information in the second modality to a user, the electronic device can adaptively convert the modality type of the information in the second modality according to device information such as device type and device function of one or more cooperating devices, and distribute it to the corresponding cooperating devices for display, thereby improving the display effect of the information in the second modality and making it easier for users to understand the information in the first modality displayed by the electronic device accurately, fully and timely.

[0071] Furthermore, in this embodiment, the information displayed by the electronic device in the second modality is obtained by matching information from other modalities based on the information in the first modality. Therefore, the obtained information in the second modality has a high degree of correlation with the information in the first modality, thus accurately and fully assisting the user in understanding the information in the first modality and improving the user experience. For example, the information in the second modality is obtained by sequentially segmenting the text answer to the user's input question into short sentences according to the short sentence dimension. Therefore, the obtained information in the second modality is information at the short sentence granularity, which has a high degree of correlation with the text answer. Thus, it can accurately and fully assist the user in understanding the corresponding short sentences and contextual meaning, improving the user experience.

[0072] In some embodiments, the information of the second modality displayed by the electronic device is acquired by the electronic device. Taking a user inputting a question as an example, the electronic device may obtain the second modality information through the following steps: First, based on the user input question, the electronic device generates a text answer to the user input question based on a large model or obtains it through searching, reading pre-stored data, receiving data from other devices, etc.; then, the electronic device performs semantic segmentation on the text answer to obtain multiple short sentences, and determines which of the multiple short sentences need to be supplemented with information from other modalities; finally, the electronic device obtains one or more other modalities corresponding to the short sentences that need to be supplemented with information from other modalities through generating, searching, reading pre-stored data, receiving data from other devices, etc.

[0073] In some embodiments, the information of the second modality displayed by the electronic device is obtained by the server. For example, the server can obtain the information of the second modality through the following steps: First, the server receives a question input by the user from the electronic device and generates a text answer to the question input by the user based on a large model or by searching, reading pre-stored data, etc.; then, the server performs semantic segmentation on the text answer to obtain multiple short sentences, and determines which of the multiple short sentences need to be supplemented with information of other modalities; finally, the server obtains one or more other modalities corresponding to the short sentences that need to be supplemented with information of other modalities by generating, searching, reading pre-stored data, etc.

[0074] As an example, please refer to Figure 3, which illustrates three multimodal content display architectures provided in this application embodiment, taking the second modality of information being obtained by the server as an example. As shown in Figure 3(a), Figure 3(b), and Figure 3(c), the system architecture includes a server 310 and an electronic device 320.

[0075] As shown in Figures 3(a), 3(b), and 3(c), the electronic device 320 is used to receive first information input by a user, send the first information input by the user to the server 310, and after receiving information of a first modality from the server 310 in response to the first information input by the user, display the first content indicated by the information of the first modality on the screen.

[0076] As shown in Figure 3(a), in some embodiments of this application, the electronic device 320 can also display multimodal information based on interaction with the user; for example, the electronic device can accurately and timely display second-modal information such as pictures, audio, video, vibration, indicator lights, AR / VR images, etc., related to the target text in the text response, based on the first content indicated by the user browsing the first-modal information; or, based on the user's operation on the second-modal information, highlight the sub-contents in the relevant first content, etc.

[0077] As shown in Figure 3(b), in some embodiments of this application, the modal content display system architecture may also include a cooperating device 330. The electronic device 320 can send information of the second modality to the cooperating device 330 to display the second content in the second modality and / or the third modality through the cooperating device 330; or, the electronic device 320 can send information of the third modality to the cooperating device 330 to display the second content in the third modality through the cooperating device 330, wherein the information of the third modality is obtained by modal type conversion of the information of the second modality.

[0078] As shown in Figure 3(c), in some embodiments of this application, the modal content display system architecture may also include a collaborative device 330, which can directly obtain information of the second modality from the server 310 and display the second content indicated by the information of the second modality in the second modality or the third modality.

[0079] The server 310 shown in Figure 3 is used to receive first information input by the user from the electronic device 320, obtain information of a first modality in response to the first information input by the user, determine information of a second modality, send the information of the first modality to the electronic device 320 for display, and send the information of the second modality to the electronic device 320 or the cooperating device 330 for display, wherein the information of the second modality is used to supplement the information of the first modality.

[0080] In the architectures shown in Figures 3(a) and 3(b), the embodiments of this application do not limit the specific manner in which the server 310 sends the first modal information and the second modal information to the electronic device 320. For example, the server 310 may send the first modal information and the second modal information together to the electronic device 320 after obtaining them; or, the server 310 may send the first modal information to the electronic device 320 and obtain the second modal information after obtaining the first modal information, and then send the second modal information to the electronic device 320. The acquisition of the second modal information may occur before or after the electronic device 320 displays the first modal information, without specific limitation.

[0081] As an example, please refer to Figure 4, which illustrates two multimodal content display architectures provided in this application embodiment, taking the acquisition of information in the second modality by an electronic device as an example. As shown in Figure 4(a) and Figure 4(b), the system architecture includes an electronic device 320.

[0082] As shown in Figure 4(a) and Figure 4(b), the electronic device 320 is used to receive first information input by the user, obtain information of a first modality in response to the first information input by the user, display the first content indicated by the information of the first modality, determine information of a second modality, and the information of the second modality is used to supplement the information of the first modality.

[0083] As shown in Figure 4(a), in some embodiments of this application, the electronic device 320 can also display multimodal information based on interaction with the user; for example, based on the first content indicated by the user browsing the first modal information, such as the target text that the user is interested in in the text answer, the second modal information such as pictures, audio, video, vibration, indicator lights, AR / VR images, etc. related to the target text can be accurately and timely displayed to the user; or, based on the user's operation on the second modal information, the sub-contents in the relevant first content can be highlighted, etc.

[0084] As shown in Figure 4(b), in some embodiments of this application, the modal content display system architecture may also include a cooperating device 330. The electronic device 320 can send information of the second modality to the cooperating device 330 to display the second content in the second modality and / or the third modality through the cooperating device 330; or, the electronic device 320 can send information of the third modality to the cooperating device 330 to display the second content in the third modality through the cooperating device 330, wherein the information of the third modality is obtained by modal type conversion of the information of the second modality.

[0085] In the architectures shown in Figure 4(a) and Figure 4(b), the embodiments of this application are not limited to the order in which the electronic device 320 acquires the information of the second mode and displays the information of the first mode. For example, the electronic device 320 may display the information of the first mode after acquiring the information of the second mode, or it may display the information of the first mode before acquiring the information of the second mode, without specific limitation.

[0086] It should be noted that Figure 3 only illustrates the acquisition of information from the first and second modalities by the server, and Figure 4 only illustrates the acquisition of information from the first and second modalities by the electronic device. In practical applications, the acquisition of information from the first and second modalities may also be accomplished collaboratively by the electronic device and the server. For example, the acquisition of information from the first modality may be the responsibility of the electronic device, while the acquisition of information from the second modality may be the responsibility of the server; or, the acquisition of information from the first modality may be the responsibility of the server, while the acquisition of information from the second modality may be the responsibility of the electronic device.

[0087] In this application embodiment, the interaction between the electronic device 320 and the user may include, but is not limited to, active detection and passive reception.

[0088] For example, the active detection method may include, but is not limited to, camera detection, finger hover detection, screen centering detection, etc., and is not limited thereto. As an example, please refer to Figure 5A, which shows schematic diagrams of three cases of determining user browsing information of the first modality through active detection methods provided in the embodiments of this application. Among them, Figure 5A(a) shows a schematic diagram of detecting user browsing information of the first modality through camera detection, Figure 5A(b) shows a schematic diagram of detecting user browsing information of the first modality through finger hover detection, and Figure 5A(c) shows a schematic diagram of detecting user browsing information of the first modality through screen centering detection.

[0089] As shown in Figure 5A(a), the electronic device can detect the user's browsing of information in the first modality using a camera. Taking the first modality information as an example, which is the text answer to the user's input question "What animals are in the zoo?" displayed on interface 501 shown in Figure 5A(a), the electronic device can detect the target text that the user is paying attention to in the text answer using a camera. For example, the camera can detect the target text that the user is paying attention to by capturing the focus position of the user's eye.

[0090] As shown in Figure 5A(b), the electronic device can detect the user's finger placement position during the browsing of information in the first modality via the touchscreen to obtain information about the first modality. Taking the first modality information as a textual answer to the user's input question "What animals are in the zoo?" displayed on interface 502 as shown in Figure 5A(b), the electronic device can detect the user's finger placement position via the touchscreen while the user browses the text answer by sliding a progress bar, and determine the target text in the text answer that the user is interested in based on that finger placement position. For example, the touchscreen can detect the user's finger placement position using touch sensors and / or pressure sensors mounted on it.

[0091] As shown in Figure 5A(c), the electronic device can determine the content to be displayed in the center based on the distribution of information in the first modality on the interface, and then determine the user's browsing of information in the first modality based on the content displayed in the center. Taking the information in the first modality as the text answer to the user's input question "What animals are in the zoo?" displayed on interface 503 shown in Figure 5A(c) as an example, the electronic device can use the content displayed in the center of the text answer as the target text while the user is browsing the text answer.

[0092] For example, passive receiving methods may include, but are not limited to, electronic devices receiving user selection operations. User selection operations may include, but are not limited to, clicking, circling, dragging, underlining, long pressing, etc., while clicking operations may include single-clicking, double-clicking, etc.

[0093] In this embodiment, the electronic device can accurately detect the target text that the user is interested in through interaction methods such as active detection and passive reception, so that it can obtain and display the corresponding second modality information based on the target text. That is, in this embodiment, the specific content of the second modality information changes dynamically according to the target text detected by the electronic device.

[0094] In some embodiments, the electronic device can also dynamically determine the display position of the corresponding second modality information based on the specific location of the detected target text, so as to avoid the obstruction of the first modality information by the display of the second modality information and the inconvenience caused to the user, and make it convenient for the user to see the second modality information corresponding to the target text at a glance, so as to facilitate the user to obtain the second modality information related to the target text in a timely and accurate manner.

[0095] As an example, please refer to Figure 5B, which illustrates a schematic diagram of how an embodiment of this application determines the information viewed by a user in a first modality based on the user's selection operation. As shown in Figure 5B, taking the text answer to the user's input question "What animals are in the zoo?" displayed on interface 504 shown in Figure 5B as an example, the electronic device can determine the target text that the user is interested in in the text answer based on the received user selection operations such as clicking, circling, dragging, underlining, and long-pressing. Figure 5B uses circling as an example of the user's selection operation.

[0096] As an example, please refer to Figure 6, which shows another multimodal content display architecture provided by the embodiments of this application, with the architecture shown in Figure 3(a) as an example.

[0097] As shown in Figure 6, the electronic device 320 includes a user operation detection module, a display screen, a rendering module, and a communication module.

[0098] The user operation detection module shown in Figure 6 is used to detect user operations, such as detecting user input of initial information, user clicks / selections / drags / underlines / long presses, and user voice commands. For example, the user operation detection module may include, but is not limited to, one or more of the following: touch sensors, pressure sensors, microphones, etc.

[0099] The electronic device communication module shown in Figure 6 is used to send the first information of user input detected by the user operation detection module to the server 310, and to receive information of the first modality and related second modality of the first information of user input from the server 310.

[0100] The rendering module shown in Figure 6 is used to draw, render, and composite the interface based on the information of the first modality, and to display the information of the first modality to the user through the display screen.

[0101] The user operation detection module shown in Figure 6 is also used to detect target content that the user is interested in within the information of the first modality. Taking a user operation detection module that includes a touch sensor and / or a pressure sensor as an example, the module can determine the target content based on detected user actions such as clicking, circling, dragging, underlining, and long-pressing on parts of the information in the first modality. Alternatively, taking a user operation detection module that includes a microphone as an example, the module can determine the target content based on detected voice recordings from the user regarding parts of the information in the first modality.

[0102] In some embodiments, the electronic device 320 may include the camera shown in FIG6, which is used to detect target content of interest to the user in the information of the first modality. For example, the camera can detect target content of interest to the user by capturing the focusing position of the user's eye.

[0103] In this embodiment, the rendering module shown in FIG6 is further configured to draw, render, and synthesize the information of the second modality with other layers (such as the information of the first modality) based on the target content of interest to the user detected by the operation detection module and / or the camera. In some embodiments, the rendering module may display the information of the second modality corresponding to the target content to the user through a display screen.

[0104] In some embodiments, the electronic device 320 may further include other display components such as a speaker, a buzzer, and an indicator light, for displaying information about the second modality corresponding to the target content to the user. For example, a speaker may be used to display information about the audio modality corresponding to the target content to the user; a buzzer may be used to display information about the vibration modality corresponding to the target content to the user; and an indicator light may be used to display the indicator light effect corresponding to the target content to the user.

[0105] As shown in Figure 6, server 310 includes a large model module, a semantic segmentation module, a placeholder prediction model module, a search engine module, and a communication module.

[0106] The large model module shown in Figure 6 is used to obtain information of the first modality based on the first information input by the user to the electronic device, such as generating information of the first modality based on the first information.

[0107] The semantic segmentation module shown in Figure 6 is used to perform semantic segmentation on the information of the first modality generated by the large model module to obtain multiple sub-contents.

[0108] The placeholder prediction model module shown in Figure 6 is used to determine which sub-contents among multiple sub-contents need to be supplemented with information from other modalities. For example, the placeholder prediction model module can determine whether multiple sub-contents need to be supplemented with information from other modalities based on preset strategies and rules, and mark the sub-contents that need to be supplemented, such as by outputting placeholders to mark the sub-contents that need to be supplemented with information from other modalities.

[0109] The search engine module shown in Figure 6 is used to retrieve second-modal information corresponding to sub-content that requires supplementary information in other modalities through multi-channel and multi-dimensional searches. For example, the search engine module can perform multi-channel and multi-dimensional searches for second-modal information such as images, videos, audio, vibration, indicator lights, and AR / VR images on the tagged sub-content.

[0110] The communication module of the server 310 shown in Figure 6 is used to receive first information input by the user from the electronic device 320, and to send information of a first modality obtained based on the first information to the electronic device 320, and to send information of a second modality related to the information of the first modality.

[0111] The electronic devices described in this application can include, but are not limited to, any electronic device capable of displaying information. These include, but are not limited to, smartphones, netbooks, tablets, smart drawing tablets, handwriting tablets, smartwatches, smart bracelets, phone watches, smart glasses, smart cameras, handheld computers, in-vehicle computers, personal computers (PCs), personal digital assistants (PDAs), portable multimedia players (PMPs), augmented reality (AR) / virtual reality (VR) devices, smart TVs, projection devices, or motion-sensing game consoles in human-computer interaction scenarios. Alternatively, the electronic device can also be other types or structures of electronic devices with a display screen; this application is not limited to these categories.

[0112] As an example, please refer to Figure 7, which shows a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.

[0113] As shown in Figure 7, the electronic device may include a processor 710, a memory (including an external memory interface 720 and an internal memory 721), a universal serial bus (USB) interface 730, a charging management module 740, a power management module 741, a battery 742, an antenna 1, an antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, a headphone jack 770D, a sensor module 780, buttons 790, a motor 791, an indicator light 792, a camera 793, a display screen 794, etc.

[0114] The sensor module 780 may include, but is not limited to, one or more of the following: touch sensor 780A, pressure sensor 780B, temperature sensor, gyroscope sensor, barometric pressure sensor, magnetic sensor, accelerometer, distance sensor, proximity sensor, fingerprint sensor, ambient light sensor, bone conduction sensor, etc.

[0115] Processor 710 may include one or more processing units. For example, processor 710 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a flight controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0116] The processor 710 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 710 is a cache memory. This memory can store instructions or data that the processor 710 has just used or that are used repeatedly. If the processor 710 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 710, and thus improves the efficiency of the system.

[0117] In some embodiments, the processor 710 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0118] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 710 may include multiple I2C buses. The processor 710 can couple to the touch sensor 780A, charger, flash, camera 793, etc., through different I2C bus interfaces. For example, the processor 710 can couple to the touch sensor 780A through the I2C interface, enabling the processor 710 and the touch sensor 780K to communicate through the I2C bus interface, thus realizing the touch function of the electronic device.

[0119] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 710 and the wireless communication module 760. For example, the processor 710 communicates with the Bluetooth module in the wireless communication module 760 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 770 can transmit audio signals to the wireless communication module 760 via the UART interface to enable music playback through Bluetooth headphones.

[0120] The MIPI interface can be used to connect the processor 710 to peripheral devices such as the display screen 794 and the camera 793. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 710 and the camera 793 communicate via the CSI interface to enable the electronic device's shooting function. The processor 710 and the display screen 794 communicate via the DSI interface to enable the electronic device's display function.

[0121] The wireless communication function of electronic devices can be implemented through antenna 1, antenna 2, mobile communication module 750, wireless communication module 760, modem processor, and baseband processor.

[0122] In some embodiments of this application, the electronic device can communicate with the server or a collaborative device through antenna 1, antenna 2, mobile communication module 750, wireless communication module 760, modem processor and baseband processor, etc., such as sending first information input by the user to the server, receiving first mode information related to the first information from the server, receiving second mode information related to the first mode information from the server, and sending second mode information related to the first mode information to the collaborative device, etc.

[0123] Electronic devices implement display functions through GPUs, displays 794, and APs. A GPU is a microprocessor for image processing, connected to both the display 794 and the AP. The GPU performs mathematical and geometric calculations for drawing, rendering, or compositing graphics. Processor 710 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0124] An application processor (AP) is a very large-scale integrated circuit that extends audio / video functionality and dedicated interfaces onto a low-power central processing unit (CPU). In electronic devices, it plays a role in computation and calling upon other functional components. For example, an AP can be used to run applications. In some examples, an AP can integrate multiple modules such as a CPU, graphics processor, video codec, and memory subsystem to execute applications.

[0125] Display screen 794 is used to display text, images, videos, etc. Display screen 794 includes a display panel. The display panel can be made of low-temperature polycrystalline silicon (LTPS), low-temperature polycrystalline oxide (LTPO), liquid crystal display (LCD), organic light-emitting diode (OLED), active-matrix organic light-emitting diode (AMOLED), flexible light-emitting diode (FLED), Miniled, MicroLED, Micro-OLED, quantum dot light-emitting diodes (QLED), etc.

[0126] In some embodiments of this application, the AP can send corresponding text, images, videos, and other data to the GPU based on the user's operation on the interface. The GPU can draw and render layers based on the data, and then display the rendered image on the display screen 794, including but not limited to text, images, videos, AR / VR images, etc.

[0127] In some embodiments of this application, the AP can send corresponding text, images, videos and other data to the GPU based on the user's operation on the interface. The GPU can draw and render layers based on the data, and composite the obtained layers with other layers to be displayed, and then display the composited image through the display screen 794.

[0128] Electronic devices can achieve shooting functions through ISP, camera 793, video codec, GPU, display 794 and AP, etc.

[0129] Camera 793 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device may include one or N cameras 793, where N is a positive integer greater than 1.

[0130] In some embodiments of this application, the camera 793 can be used to detect target content that the user is interested in in the information of the first modality, such as by capturing the focus position of the user's eyeballs to detect target content that the user is interested in.

[0131] The external memory interface 720 can be used to connect external memory cards, such as Micro SD cards, to expand the storage capacity of electronic devices. The external memory card communicates with the processor 710 through the external memory interface 720 to perform data storage.

[0132] Internal memory 721 can be used to store computer executable program code. Exemplarily, the computer program may include an operating system program and application programs. The executable program code includes instructions. Processor 710 executes various functional applications and data processing of the electronic device by running the instructions stored in internal memory 721. Internal memory 721 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device, etc. Furthermore, internal memory 721 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 710 executes various functional applications and data processing of the electronic device by running instructions stored in internal memory 721 and / or instructions stored in memory disposed within the processor.

[0133] In some embodiments of this application, the internal memory 721 may be used to store information of the first modality acquired in relation to the first information and information of the related second modality.

[0134] In some embodiments of this application, for example, when the multimodal information is acquired by an electronic device, the internal memory 721 can also be used to store multiple sub-contents obtained after semantic segmentation of the information of the first modality, as well as to store the markers (such as placeholders) of the sub-contents that need to be supplemented with information of other modalities.

[0135] Electronic devices can implement audio functions through audio modules 770, speakers 770A, receivers 770B, microphones 770C, and access points (APs), such as music playback and recording.

[0136] The audio module 770 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 770 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 770 may be located in the processor 710, or some functional modules of the audio module 770 may be located in the processor 710.

[0137] The speaker 770A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Electronic devices can listen to music or make hands-free calls through the speaker 770A. In some embodiments of this application, the electronic device can play voice modal information through the speaker 770A.

[0138] Microphone 770C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 770C, inputting the sound signal into microphone 770C. An electronic device can have at least one microphone 770C. In some embodiments, the electronic device can have two microphones 770C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, the electronic device can have three, four, or more microphones 770C, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions. In some embodiments of this application, the electronic device can detect the user's voice commands through microphone 770C.

[0139] Touch sensor 780A, also known as a "touch panel," can be located on display screen 794. The touch sensor 780A and display screen 794 together form a touchscreen, also known as a "touch screen." Touch sensor 780A detects touch operations applied to or near it. The touch sensor can transmit raw information, such as touch position, touch pressure, touch angle, touch area, and touch duration, to the processor for subsequent operation information acquisition, input event generation, and processing. The processor can also provide visual output related to the touch operation through display screen 794. In some embodiments, touch sensor 780A may also be located on the surface of the electronic device, in a different position than display screen 794.

[0140] Pressure sensor 780B is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 780B can be disposed on display screen 794. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to the pressure sensor, the capacitance between the electrodes changes. The electronic device determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 794, the electronic device can detect the touch force, etc., through the pressure sensor. The electronic device can also calculate the touch position based on the detection signal from pressure sensor 780B. In some embodiments, pressure sensor 780B can be used to detect information such as touch position, touch force, touch angle, touch area, and touch duration of a stylus.

[0141] In some embodiments of this application, touch sensor 780A and / or pressure sensor 780B can be used to detect user selection operations such as clicking, circling, dragging, underlining, and long-pressing on a portion of the information in the first modality, so as to determine the target content that the user is interested in.

[0142] In the embodiments of this application, the touch operation detected by the touch sensor 780A can be the operation of the user on the touch screen by the user's finger, or the user's click operation, long press operation, swipe operation (such as swipe down, swipe up, swipe left, swipe right, etc.), drag operation, hover gesture operation, etc. on the touch screen by the user using touch auxiliary tools such as stylus, stylus, stylus ball, etc. This application does not limit it.

[0143] Motor 791 can generate vibration alerts. Motor 791 can be used for incoming call vibration alerts or for touch vibration feedback. For example, touch operations applied to different applications (such as taking photos, playing audio, etc.) can correspond to different vibration feedback effects. Touch operations applied to different areas of the display screen 794 can also correspond to different vibration feedback effects from motor 791. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. Touch vibration feedback effects can also be customized. In some embodiments of this application, the electronic device can use motor 791 to display vibration mode information to the user.

[0144] The indicator light 792 can display lighting effects, and different indicator light effects can be used to convey different meanings. In some embodiments of this application, electronic devices can use the indicator light 792 to display information about the indicator light mode to the user.

[0145] For a description of the charging management module 740, power management module 741, battery 742, and other modules shown in Figure 7, please refer to conventional technology; the embodiments in this application will not be described in detail.

[0146] It is understood that the structure illustrated in Figure 7 of this application does not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0147] The following will use the architecture shown in Figure 3 or Figure 6 as an example, combined with specific embodiments, to specifically introduce the method for demonstrating the content provided in the embodiments of this application.

[0148] For example, please refer to Figure 8, which shows a flowchart of a content display method provided in an embodiment of this application. As shown in Figure 8, the method may include S801-S805:

[0149] S801: The electronic device acquires information of the first mode, which is used to indicate the first content.

[0150] The first mode may include, but is not limited to, text, images, audio, video, vibration, indicator lights, AR / VR images, etc.

[0151] As one possible implementation, the information in the first modality is the first content determined by the electronic device based on the first information input by the user.

[0152] The first piece of information entered by the user can include, but is not limited to, questions, images, videos, music, keywords, etc. Keywords can be single characters, words, short phrases, sentences, etc., without any restrictions.

[0153] For example, the first information input by the user may be a question, which is used to obtain one or more of the following: analysis, explanation, answer, or response to the question. For instance, the first modality of the first information may be a textual answer to the question, i.e., the first content may be a textual answer to the question. Alternatively, the first information may include one or more of the following: images, videos, music, keywords; and the first information input by the user may be used to obtain one or more of the following: introduction, analysis, explanation.

[0154] For example, the way a user inputs the first information may include, but is not limited to, text input, voice input, etc., without specific limitations.

[0155] In the embodiments of this application, the electronic device may obtain the information of the first mode in one or more of the following ways, including but not limited to: the electronic device generating the information of the first mode, the electronic device obtaining the information of the first mode by searching, the electronic device reading the pre-stored information of the first mode, and the electronic device receiving the information of the first mode from other devices.

[0156] This process includes: electronic devices generating first modal information, such as generating first modal information based on user input and a large model; electronic devices acquiring first modal information through searching, such as acquiring first modal information based on user input through multi-channel and multi-dimensional searches; electronic devices reading pre-stored first modal information, such as reading stored first modal information based on user input from memory; and electronic devices receiving first modal information from other devices, such as receiving first modal information based on user input from a server.

[0157] For example, an electronic device establishes a communication connection with a server. The electronic device can send second information to the server through the communication connection to indicate first information, so that the server can generate, search or read first modal information about the first information based on a large model, and the electronic device can receive the first modal information about the first information it has acquired from the server.

[0158] As an example, the second piece of information is the same as the first piece of information; that is, the electronic device may directly send the first piece of information entered by the user to the server.

[0159] As an example, the second information includes the first information as well as other information; that is, the electronic device may send the first information entered by the user along with other information to the server.

[0160] As an example, the second information does not include the first information, but the first information can characterize the first information. That is, the electronic device may send the first information to the server through implication or after modifying the first information.

[0161] The embodiments of this application do not specifically limit the specific methods or forms by which electronic devices send the first information to the server.

[0162] S802: The electronic device displays the first content to the user in the first modality.

[0163] As an example, an electronic device can display first content to a user through a display component adapted to the specific modality type of the first modality. For instance, assuming the first modality is text, image, or video, the electronic device can display the first content through a display screen; or, assuming the first modality is audio, the electronic device can display the first content through a speaker.

[0164] Taking the example where the first content is a textual answer to a question entered by the user, the electronic device can display an interface that includes the textual answer to the question entered by the user.

[0165] As one example, an electronic device may display the first content through a single interface; as another example, if the first content has many sub-contents, the electronic device may display the first content through multiple interfaces. The display of these multiple interfaces may be triggered by the user's page-turning operation or by the user's finger sliding across a progress bar. This application does not limit the scope of these interfaces and their display can be determined based on the specific device functions, interface design, etc.

[0166] In some embodiments, when displaying the first content, the electronic device may also display one or more sub-contents in the first content in a display mode that is different from other content, such as including but not limited to highlighting, underlining, wavy lines, displaying special marks, etc., so that the user knows which sub-contents correspond to other modal information, and then views other modal information through further operations.

[0167] S803: The electronic device identifies the target content that the user is interested in from the first content.

[0168] As an example, an electronic device can determine the target content that the user is interested in within the first content based on the first action the user receives regarding a portion of the first content.

[0169] As an example, electronic devices can detect a user's first action through proactive detection to determine the target content that the user is paying attention to in the first content.

[0170] For example, electronic devices can use cameras to detect how a user views the primary content. For instance, a camera can detect the user's eye focus to identify the target content within the primary content. Similarly, electronic devices can use touchscreens to detect the location of a user's finger while browsing the primary content, thus identifying the target content. Furthermore, electronic devices can determine the centrally displayed content based on the distribution of the primary content, and then use this centrally displayed content to determine the user's browsing behavior, identifying it as the target content.

[0171] As an example, electronic devices can passively receive a user's first action to determine the target content that the user is interested in from the first content.

[0172] For example, an electronic device can determine the target content that the user is interested in within the first content based on user selection operations such as clicking, circling, dragging, underlining, and long-pressing received by touch sensors and / or pressure sensors. For instance, the selected portion of the content can be identified as the target content that the user is interested in. Alternatively, the electronic device can determine the target content that the user is interested in within the first content based on voice commands issued by the user regarding a portion of the first content received by a microphone.

[0173] In some examples, compared to electronic devices actively detecting and determining target content, passively receiving and determining target content can avoid the problem of electronic devices misidentifying target content that the user is interested in, thereby improving the reliability and accuracy of subsequent display of second-modal information.

[0174] S804: The electronic device acquires information of the second modality corresponding to the target content, and the information of the second modality is used to indicate the second content.

[0175] The second modality information indicates a second content that differs from the first modality information indicates a first content. The second modality information is used to assist the user in understanding the first content indicated by the first modality information. The second modality information corresponding to the target content is used to assist the user in accurately, fully, and promptly understanding the target content.

[0176] For example, the second mode may include image mode, video mode, audio mode, vibration mode, indicator light mode, AR / VR image mode, etc., without limitation.

[0177] It should be noted that, in practical applications, the embodiments of this application do not limit the specific order in which the electronic device detects the target text that the user is interested in and obtains the second modality information corresponding to the target text, but depend on factors such as the computing power of the specific device.

[0178] As one possible implementation, the electronic device may, while acquiring information of the first modality, also acquire information of the second modality related to the first content indicated by the information of the first modality. For example, it may acquire information of the second modality related to multiple sub-contents within the first content. This second modality information related to the first content includes information of the second modality corresponding to the target content. The target content overlaps with at least one of the multiple sub-contents, such as the target content being the same as at least one of the multiple sub-contents, or the target content belonging to at least one of the multiple sub-contents, or a portion of the target content being the same as at least one of the multiple sub-contents while another portion is different. In this case, S804 may include: the electronic device determining the information of the second modality corresponding to the target content from the acquired information of the second modality related to the first content. In this implementation, the information of different second modalities may correspond to different sub-contents within the first content, or they may correspond to the same sub-content within the first content, without limitation.

[0179] As one possible implementation, after executing S803, the electronic device may generate, search, read, or receive information of the corresponding second mode based on the target content that the user is interested in.

[0180] For example, an electronic device can receive information about a second modality corresponding to target content from a server. For instance, while a user is browsing first content, the electronic device can send detected target content that the user is interested in to the server, so that the server can search, generate, or read information about the second modality corresponding to the target content, and the electronic device can receive the acquired information about the second modality corresponding to the target content from the server.

[0181] Alternatively, for example, the electronic device can independently search, generate, or retrieve information corresponding to the target content in a second modality. For instance, while the user is browsing the first content, the electronic device can search, generate, or retrieve information in a second modality based on detected target content that the user is interested in.

[0182] It should be noted that the embodiments of this application do not limit the number of second modal information corresponding to the target content. For example, the target content may correspond to one second modal information or multiple second modal information. In the case where the target content corresponds to multiple second modal information, the modal types of these multiple second modal information may be the same (e.g., all are images, videos, audio, etc.) or different (e.g., respectively images, audio, video, etc.), without specific limitations. For example, taking the target content as "cuckoo cuckoo," the second modal information corresponding to this target content might be an image of a cuckoo and an audio recording of a cuckoo's call. The simultaneous display of multiple different types of second modal information can provide users with an immersive and lifelike experience.

[0183] S805: The electronic device performs a first operation to display second content to the user.

[0184] In this application embodiment, the electronic device performing the first operation may include: the electronic device displaying second content in a second modality through the display component of the electronic device; and / or, displaying the second content in a second modality and / or a third modality through the display component of a cooperating device, wherein the device information of the cooperating device matches the second modality and / or the third modality.

[0185] In some embodiments, the electronic device performing the first operation may include, but is not limited to: the electronic device displaying the second content in a second mode, the electronic device displaying the second content in a second mode through a cooperating device, the electronic device displaying the second content in a third mode through a cooperating device, the electronic device displaying the second content in both the second and third modes through different cooperating devices, the electronic device displaying the second content in the second mode and displaying the second content in the third mode through a cooperating device, and the electronic device displaying the second content in the second mode and displaying the second content in both the third and fifth modes through different cooperating devices.

[0186] Taking the structure shown in Figure 3(a) or Figure 4(a) as an example, the electronic device can display the second content in a second mode; taking the structure shown in Figure 3(b) or Figure 4(b) as an example, the electronic device can display the second content in a second mode through a cooperating device, display the second content in a third mode through a cooperating device, display the second content in both the second and third modes through different cooperating devices, display the second content in the second mode and then display it in the third mode through a cooperating device, or display the second content in the second mode and then display it in both the third and fifth modes through different cooperating devices. For example, the electronic device displays the second content in the second mode through display components, such as a display screen, speaker, vibration, indicator lights, etc., depending on the mode type of the second mode and the device type, device function, and other device information of the electronic device. Examples of the electronic device displaying the second content in the second mode are shown in Figures 9 and 10.

[0187] Taking a second modality as an example, which is a text modality, an image modality, or a video modality, when displaying the second content in the second modality, the purpose of the solution provided in this application embodiment is to ensure that the second content and its corresponding target content are displayed on the same interface. For example, they are displayed on the same interface with a distance between them less than a preset threshold, or they are displayed on the same interface in a prominent position. In other words, in this application embodiment, the target content that the user is interested in and the second content will always be displayed on the same interface. Regarding the specific display position of the second content, this application embodiment does not impose specific limitations and can be determined based on the specific device type, screen state, display screen size, interface layout, etc. Device types include smartphones, tablets, single-screen phones, multi-screen phones, etc. Screen states include folded state, unfolded state, etc.

[0188] As one possible implementation, when an electronic device or collaborative device displays second content, the display area and / or display method of the second content can be adaptively adjusted according to the device type, screen status, display screen size, interface layout, device function, etc.

[0189] As an example, displaying second content on an electronic device may include: displaying the second content in a fixed display area or a non-fixed display area. The fixed display area may be predefined by the electronic device or application, and may be located in, but is not limited to, the upper left corner, upper right corner, etc., of the screen; the non-fixed display area is a display area whose position and / or size can be flexibly changed, such as automatically changing its position and / or size, or changing its position and / or size based on user dragging or other operations.

[0190] For example, a fixed display area may be a pre-defined area by an electronic device or application that is not used to display the first content; that is, the second content does not obscure the first content. A non-fixed display area may be dynamically determined by an electronic device or application based on the interface layout and / or the location of the target content; therefore, it is also called a dynamic display area. For example, a non-fixed display area is not used to display the target content; that is, the second content does not obscure the target content.

[0191] As an example, the display of second content on an electronic device may include: the electronic device displaying the second content in a floating display mode, wherein the second content does not obscure the target content; or, the electronic device displaying the second content in a tiled display mode, wherein the second content does not obscure the first content.

[0192] As an example, the display of second content by an electronic device may include: the electronic device displaying the second content in a tiled manner in a fixed display area, wherein the second content does not obscure the first content; or, the electronic device displaying the second content in a floating manner in a non-fixed display area, wherein the second content does not obscure the target content.

[0193] As an example, assuming the screen of the electronic device is a foldable screen, the display of second content by the electronic device may include: when the screen of the electronic device is in a folded state, the electronic device displays the second content in a non-fixed display area, and / or the electronic device displays the second content in a floating display mode; when the screen of the electronic device is in an unfolded state, the electronic device displays the second content in a fixed display area, and / or the electronic device displays the second content in a tiled display mode.

[0194] Taking a mobile phone with its screen folded as an example, where the electronic device displays second content in a floating display mode on a non-fixed display area, please refer to Figure 9. Figure 9 illustrates a content display process provided by an embodiment of this application, where the first information entered by the user in the text input box on the functional interface or application interface of the electronic device is the question "What animals are in the zoo?". As shown in Figure 9, when the electronic device displays the question-and-answer application interface 900, the user can obtain any one or more of the analysis, explanation, answer, or response to the question by entering the question "What animals are in the zoo?" in the text input box on the question-and-answer application interface 900 and then clicking the search button. Then, the electronic device can send the question entered by the user in the text input box on the Q&A application interface 900 to the server, or perform operations such as correction, optimization, and polishing on the question entered by the user in the text input box before sending it to the server, so as to obtain a text answer to the question (i.e., the information of the first modality mentioned above) and related second modality information from the server; or, the electronic device can generate a text answer to the question and related second modality information based on the question entered by the user in the text input box on the Q&A application interface 900; the second modality information includes pictures of monkeys, giraffes, etc. Afterwards, the electronic device can display the text answer to the question to the user through interface 901 in Figure 9, such as the monkeys, peacocks, giraffes, gibbons, pandas, etc. in the zoo shown in interface 901, as well as related introductions to each animal.

[0195] During the user's text-based response on interface 901, the electronic device can determine the target text (i.e., the target content mentioned above) that the user is interested in through active detection or passive reception methods. As shown in Figure 9, assuming the target text that the user is interested in is "monkeys are a common name for some primates" on interface 901, the electronic device will display a monkey image (i.e., the second content mentioned above) corresponding to the target text "monkeys are a common name for some primates" in a floating display area (i.e., the non-fixed display area mentioned above) in a floating display mode through a floating window 903 on interface 902.

[0196] While the user is browsing the text answers, the electronic device can refresh the interface based on the user's swiping or page-turning actions to display all the text answers.

[0197] For example, referring to Figure 10, in response to the user's swiping or page-turning operation on interface 902, the electronic device refreshes the interface, displays the subsequent content of the text answer through interface 904 as shown in Figure 10, and stops displaying the monkey image. While the user is browsing the text answer on 904, the electronic device can continue to determine the target text in the text answer that the user is interested in through active detection, passive reception, or other methods. As shown in Figure 10, assuming the target text that the user is interested in is "a giraffe is a ruminant artiodactyl that lives in Africa" ​​on interface 904, the electronic device displays the giraffe image (i.e., the second content mentioned above) corresponding to the target text "a giraffe is a ruminant artiodactyl that lives in Africa" ​​in a floating display area (i.e., the non-fixed display area mentioned above) in a floating display mode through a floating window 906 on interface 905 in a display area not being focused on by the user (i.e., the non-fixed display area mentioned above).

[0198] As shown in Figures 9 and 10, the electronic device can accurately detect target text of interest to the user through interactive methods such as active detection and passive reception. This allows it to acquire and display the corresponding second-modal information based on the target text. Specifically, in this embodiment, the content of the second-modal information dynamically changes according to the target text detected by the electronic device. Furthermore, the electronic device can dynamically determine the display position of the corresponding second-modal information based on the specific location of the detected target text, avoiding any obstruction of the first-modal information and ensuring the user can easily see the second-modal information at a glance. This facilitates timely and accurate acquisition of second-modal information related to the target text.

[0199] Taking a tablet computer as an example, where the electronic device displays the second content in a tiled display mode in a fixed display area, please refer to Figure 11 for an example. Figure 11 shows a schematic diagram of an interface for displaying multimodal information on an electronic device according to an embodiment of this application. As shown in Figure 11, when it is detected that the target content that the user is interested in is "monkeys are common names for some primates" (i.e., the target content) on interface 1100 shown in Figure 11, and it is determined that the device type of the electronic device is a tablet computer, the electronic device can display the monkey image (i.e., the second content) corresponding to the target content "monkeys are common names for some primates" in a tiled display mode in area 1102 (i.e., the second display area mentioned above) on interface 1101.

[0200] In some embodiments, the screen of the electronic device is a foldable screen. During the process of the electronic device displaying the second content in the second mode, the electronic device may also switch the display area and / or display mode of the second content according to the change of the screen state.

[0201] As an example, during the process of an electronic device displaying second content in a first display area, if the screen of the electronic device changes from a folded state to an unfolded state, the electronic device changes from displaying the second content in a non-fixed display area to displaying the second content in a fixed display area, and / or, the electronic device changes from displaying the second content in a floating display mode to displaying the second content in a tiled display mode.

[0202] As an example, during the process of an electronic device displaying second content in a second display area, if the screen of the electronic device changes from an unfolded state to a folded state, the electronic device changes from displaying the second content in a fixed display area to displaying the second content in a non-fixed display area, and / or, the electronic device changes from displaying the second content in a tiled display mode to displaying the second content in a floating display mode.

[0203] For example, please refer to Figure 12, which shows another schematic diagram of an interface for displaying multimodal information on an electronic device according to an embodiment of this application. As shown in Figure 12, when the target content that the user is interested in is "monkeys are a common name for some primates" (i.e., the target content) on interface 1200 shown in Figure 12, and it is determined that the device type of the electronic device is a smartphone and the screen state is folded, the electronic device can display the monkey image corresponding to the target content "monkeys are a common name for some primates" (i.e., the second content mentioned above) in a floating display mode in area 1201 (i.e., the non-fixed display area mentioned above); when the screen state changes from folded to unfolded, the electronic device changes from displaying the monkey image in a floating display mode in area 1201 to displaying the monkey image in a tiled display mode in area 1203 (i.e., the fixed display area mentioned above) on interface 1202.

[0204] It should be noted that Figure 12 is merely an example of how an electronic device changes its display strategy (such as display area, display method, etc.) for displaying second content based on changes in the specific screen state. When the screen state changes, the electronic device may also adjust the layout of the first content to adapt to the change in screen state. The process of adaptive adjustment of the layout of the first content by the electronic device according to changes in screen state is not described in detail in this embodiment, but can be referred to conventional technology.

[0205] It is understandable that by determining the display strategy for the second content based on device information such as device type and screen status, the adaptability of the second content display to the specific device situation can be improved. This ensures that users can accurately, fully, and promptly understand the information displayed on the electronic device while achieving the best display effect for the second content and enhancing the user's viewing experience. Especially for scenarios where the screen status changes, the second content display strategy adaptation scheme provided in this application embodiment can make full use of the electronic device's display screen, greatly improving the information display effect.

[0206] Figures 9-12 above use the example of an electronic device displaying first content in a first mode and second content in a second mode. As mentioned above, in some embodiments, the electronic device can display the second content to the user through other electronic devices, such as displaying the second content of the second mode to the user through a collaborative device, or displaying the second content of different modes to the user through different collaborative devices, or displaying the second content of different modes to the user synchronously with one or more collaborative devices.

[0207] For example, an electronic device may interconnect with one or more collaborating devices. Device information for collaborating devices may include, but is not limited to, device type and functions. Device types include, but are not limited to, smartphones, tablets, smartwatches, indicator lights, etc.; device functions include, but are not limited to, image display, audio playback, video playback, vibration, lighting, virtual reality, and augmented reality functions.

[0208] As an example, an electronic device can decide whether to display the second content through a collaborative device, and which collaborative devices to display the second content through, based on actual circumstances, such as the specific modal type of the second mode (e.g., the first modal type) and device information such as the device type and device function of one or more collaborative devices, and distribute the second content accordingly.

[0209] For example, an electronic device may display second content in a second modality through a cooperating device. For instance, the electronic device sends second modality information to the cooperating device and instructs the cooperating device to display the second content; correspondingly, the cooperating device displays the second content in the second modality based on the received second modality information.

[0210] Taking the second modal information corresponding to the target content, which includes image modal information and audio modal information, as an example, the electronic device can determine from one or more collaborative devices whether it has the ability to display image modal information and whether it has the ability to display audio modal information, or whether it can display image modal information and audio modal information with the best effect respectively.

[0211] For example, please refer to Figure 13, which illustrates a scenario of displaying second content through a collaborative device according to an embodiment of this application. As shown in Figure 13, taking the user-input question "How does a cuckoo call?" as an example, the electronic device determines the corresponding second modality information, including a picture of a cuckoo and an audio recording of a cuckoo's call, based on the user's gaze at the target text "cuckoo cuckoo" during the browsing of the text answer corresponding to the question. Further, based on the modal types of the two second modal information (i.e., image modality and audio modality) and the device information of the electronic device's collaborative devices (i.e., smart TV and smart speaker), the electronic device decides whether to display a picture of a cuckoo on the smart TV and an audio recording of a cuckoo's call on the smart speaker. Based on this, the electronic device sends the picture of a cuckoo to the smart TV and the audio recording of a cuckoo's call to the smart speaker. Ultimately, the electronic device displays the text answer, and the smart TV and smart speaker synchronously display the picture of a cuckoo and the audio recording of a cuckoo corresponding to the target text "cuckoo cuckoo" according to the user's browsing of the text answer.

[0212] It is understandable that different devices, due to differences in hardware configuration such as device type and functions, are suited to displaying second content in different ways. For example, smart TVs have larger screens, thus displaying images and videos better; smart speakers have higher speaker power, thus playing sound better; and smartwatches, worn close to the user's body, can transmit vibration sensations, and so on. Therefore, the embodiments of this application can combine the hardware advantages of multiple electronic devices to distribute second content in a second modality, fully engaging the user's multiple senses to understand the answer, ensuring that the user fully understands the content, and greatly improving the user experience.

[0213] For example, an electronic device might display the second content in a third modality through a collaborating device. For instance, the electronic device converts the second-modal information into third-modal information and sends it to the collaborating device, instructing the collaborating device to display the second content. Correspondingly, the collaborating device displays the second content in the third modality based on the received third-modal information. Alternatively, the electronic device sends the second-modal information to the collaborating device and instructs the collaborating device to display the second content. Correspondingly, the collaborating device converts the received second-modal information into third-modal information based on its device type, device function, and other device information, and then displays the second content in the third modality. Here, the third modality differs from the second modality; the information in the third modality is determined based on the second-modal information corresponding to the target content.

[0214] As an example, an electronic device can decide whether to convert the modal type of the second modality information, and which modal type to convert, based on actual conditions such as the device type and functions of one or more collaborating devices. For instance, if the modal type of the second modality information is the first modality, the electronic device can convert the second modality information into a third modality information that matches the device type and / or functions of the collaborating devices, based on the device information of one or more collaborating devices. Ultimately, the collaborating devices will display the third modality information. For example, if a collaborating device has video playback capabilities, the electronic device can convert second modality information such as images, audio, vibration, and indicator lights into a video modality; if a collaborating device has vibration capabilities, the electronic device can convert modal information such as images, audio, video, and indicator lights into a vibration modality.

[0215] As an example, please refer to Figure 14, which illustrates another scenario of displaying second-modal information via a collaborative device according to an embodiment of this application. As shown in Figure 14, taking the user-input question "What is haptic feedback timekeeping?" as an example, the electronic device determines the corresponding second-modal information, including a video of haptic timekeeping, based on the user's gaze at the target text "A short tap represents one hour" during the process of browsing the text answer to the question. Further, based on the device information of the electronic device's collaborative device (i.e., a smartwatch), the electronic device decides that the smartwatch will display the haptic timekeeping in vibration mode. Based on this, the electronic device converts the haptic timekeeping video into a vibration mode and sends it to the smartwatch. Ultimately, this achieves the specific method where the electronic device displays the text answer, and the smartwatch synchronously vibrates once to display the haptic timekeeping to the user based on the user's browsing of the target text "A short tap represents one hour".

[0216] It is understandable that different devices, due to differences in hardware configurations such as device type and functions, are suited to displaying secondary content in different ways. For example, smart TVs have larger screens, thus displaying images and videos better; smart speakers have higher speaker power, thus playing sound better; and smartwatches, worn close to the user's body, can transmit vibration sensations, and so on. Therefore, the embodiments of this application can combine the hardware advantages of the device to perform modal type conversion, fully mobilizing the user's multiple senses to understand the answer, ensuring that the user fully understands the content, and greatly improving the user experience.

[0217] For example, electronic devices may also display the second content in a second mode and a third mode through different collaborating devices. For instance, the electronic device may send second-mode information to a first collaborating device and instruct the first collaborating device to display the second content; or the electronic device may convert the second-mode information into third-mode information and send it to a second collaborating device and instruct the second collaborating device to display the second content. Correspondingly, the first collaborating device displays the second content in the second mode based on the received second-mode information, and the second collaborating device displays the second content in the third mode based on the received third-mode information. Alternatively, the electronic device may send second-mode information to both the first and second collaborating devices and instruct them to display the second content. Correspondingly, the first collaborating device displays the second content in the second mode based on the received second-mode information, and the second collaborating device converts the second-mode information into third-mode information based on the received third-mode information and its device type, device function, and other device information, and then displays the second content in the third mode.

[0218] Alternatively, the electronic device may also display the second content in a second modality and display the second content in a third modality through a cooperating device. For example, the electronic device displays the second content in a second modality, converts the information in the second modality into information in the third modality, sends it to the cooperating device, and instructs the cooperating device to display the second content. Correspondingly, the cooperating device displays the second content in the third modality based on the received information in the third modality. Or, the electronic device displays the second content in a second modality, sends the information in the second modality to the cooperating device, and instructs the cooperating device to display the second content. Correspondingly, the cooperating device converts the information in the second modality into information in the third modality based on the received information in the second modality and its device type, device function, and other device information, and then displays the second content in the third modality.

[0219] Alternatively, the electronic device may also display the second content in a second modality and display the second content in a third and fifth modality respectively through different cooperating devices. For example, the electronic device displays the second content in a second modality, and after converting the information in the second modality into information in the third and fifth modality respectively, it sends the information in the third and fifth modality to the first and second cooperating devices respectively. Correspondingly, the first cooperating device displays the second content in the third modality based on the received information in the third modality, and the second cooperating device displays the second content in the fifth modality based on the received information in the fifth modality. Or, the electronic device displays the second content in a second modality and sends the information in the second modality to the first and second cooperating devices respectively. Correspondingly, the first cooperating device combines the received information in the second modality with its device type. The first collaborative device converts the second modal information, such as device type and device function, into third modal information and displays the second content in the third modal. The second collaborative device, based on the received second modal information and its device type and device function, converts the second modal information into fifth modal information and displays the second content in the fifth modal. Alternatively, the electronic device displays the second content in the second modal, sends the second modal information to the first collaborative device, converts the second modal information into third modal information, and sends it to the second collaborative device. Correspondingly, the first collaborative device, based on the received second modal information and its device type and device function, converts the second modal information into third modal information and displays the second content in the third modal. The second collaborative device, based on the received fifth modal information, displays the second content in the fifth modal.

[0220] The above example only illustrates how the collaborative device obtains information of the first modality or the second modality used to indicate the second content from the electronic device. In some embodiments, such as for the structure shown in Figure 3(c), the collaborative device may also obtain information of the second modality corresponding to the target content detected by the electronic device and display the second content accordingly, such as displaying the second content in the second modality, or converting the information of the second modality into information of the third modality and then displaying the second content in the third modality.

[0221] In some embodiments of this application, the electronic device may also support displaying content according to a preset user action on the second content (i.e., the content indicated by the information of the second modality). For example, the electronic device may display information of other modalities (such as information of the fourth modality) related to other sub-contents in the first content, or display sub-content at a first position in the first content (i.e., the content indicated by the information of the first modality) that is related to the second content, wherein the sub-content at the first position is different from the target content.

[0222] As an example, the second action may be a user's click, swipe up, swipe down, swipe left, or swipe right on the second content, or a user's click, swipe up, swipe down, swipe left, or swipe right on a preset position of the second content. This application does not limit the specific action method, but depends on the specific circumstances. Different methods of second actions may correspond to different operational intentions of the method.

[0223] For example, when the second action is a user's click on the second content, or a user's click on a preset position of the second content, the intention of the action may be to display the sub-content in the first content that is most recently related to the current second content (i.e., the first position mentioned above).

[0224] For example, when the second action is a user's swipe-up operation on the second content, or a user's swipe-up operation on a preset position of the second content, the intention of the operation may be to display the sub-content of the first content that is related to the current second content at the previous position (i.e., the first position mentioned above).

[0225] For example, when the second action is a user's swipe-down operation on the second content, or a user's swipe-down operation on a preset position of the second content, the intention of the operation may be to display the sub-content of the first content at the next position related to the current second content (i.e., the first position mentioned above).

[0226] For example, when the second action is a left swipe operation by the user on the second content, or a left swipe operation by the user on the preset position of the second content, the intention of the operation may be to display information of other modalities related to the previous sub-content in the first content.

[0227] For example, when the second action is a right swipe operation by the user on the second content, or a right swipe operation by the user on the preset position of the second content, the intention of the operation may be to display information of other modalities related to the next sub-content in the first content.

[0228] As an example, please refer to Figure 15, which shows a schematic diagram of an electronic device displaying a content interface based on a user's second action according to an embodiment of this application. As shown in Figure 15, when the electronic device displays sentence 1 (i.e., the target content mentioned above) and image 1 corresponding to sentence 1 (i.e., the second content indicated by the information of the second modality corresponding to the target content mentioned above) through interface 1500, in response to the user's double-click operation on image 1, the electronic device displays the text corresponding to image 1 at the position closest to the currently viewed text content (i.e., sentence 1), i.e., sentence 5 on interface 1501 shown in Figure 15.

[0229] As shown in Figure 16, when the electronic device displays sentence 5 (i.e. the target content mentioned above) and image 1 corresponding to sentence 5 (i.e. the second content indicated by the information of the second modality corresponding to the target content mentioned above) through interface 1600, in response to the user's swipe-up operation on image 1, the electronic device displays the text at the position above image 1, i.e., sentence 1 on interface 1601 shown in Figure 16.

[0230] As shown in Figure 17, when the electronic device displays sentence 1 (i.e., the target content mentioned above) and image 1 corresponding to sentence 1 (i.e., the second content indicated by the second modality information corresponding to the target content mentioned above) through interface 1700, in response to the user's left swipe operation on image 1, the electronic device displays the next sub-content corresponding to the second modality information and the second modality information corresponding to the next sub-content, as shown in Figure 17 on interface 1701, which displays sentence 2 and image 2 corresponding to sentence 2.

[0231] It should be noted that the above-described forms of second actions are merely examples, and the embodiments of this application do not limit the specific actions or the operational intent corresponding to each action. For example, the intent of the second action may also be any of the following: displaying information of the next other modality corresponding to the target content, displaying information of the previous other modality corresponding to the target content, displaying the sub-content at the nearest position in the first content corresponding to the second content, displaying all the sub-content in the first content corresponding to the second content, etc.

[0232] It's understandable that electronic devices can meet diverse user needs, such as improving the speed of information retrieval, by displaying content based on preset user actions on the second content. For example, if the first modality is text-based, and the user wants to primarily view information from other modalities besides the first, supplementing their understanding with the information from the first modality, the user can directly interact with the second content. This allows them to sequentially view information from other modalities besides the first, saving time spent reading text. Similarly, if the user wants to view the corresponding text content of the second content, they can directly interact with the second content, such as sequentially viewing the text content at the previous, next, or all positions within the first content, to quickly obtain all information related to the first content. This further enhances the user experience when acquiring information.

[0233] The following description, in conjunction with the accompanying drawings, details the specific process by which the first modality information and related second modality information are obtained according to the above embodiments of this application.

[0234] As an example, an electronic device or server can acquire multimodal information based on the first piece of information input by the user, based on the following steps 1-4:

[0235] Step 1: After obtaining the first information input by the user, generate information (such as the first content) of the first modality based on the large model.

[0236] Step 2: Perform semantic segmentation on the first content to obtain multiple sub-contents.

[0237] Step 3: Identify the first sub-content among multiple sub-contents that needs to be supplemented with information from other modalities.

[0238] Step 4: Obtain the information of the second modality corresponding to the first sub-content.

[0239] For example, taking a question as the first piece of information input by the user, please refer to Figure 18. Figure 18 shows a flowchart of obtaining multimodal information based on the first piece of information input by the user, according to an embodiment of this application. As shown in Figure 18, multimodal information based on the first piece of information input by the user can be obtained based on S1801-S1804:

[0240] S1801: After obtaining the user's input question, generate a text answer to the question based on the question-answering model.

[0241] Among them, written answers to the above questions include any one or more of the following: analysis, explanation, response, or answer to the above questions.

[0242] For example, after obtaining the user's input question, the question can be fed into a pre-trained question-answering model. Through the relevant processing of the question-answering model, the textual answer to the question can be obtained from the output of the question-answering model.

[0243] In some embodiments, after obtaining the user's input question, the question may be corrected, optimized, and polished before generating a text answer based on the question-answering model.

[0244] S1802: Perform semantic segmentation on the text response to obtain multiple short sentences.

[0245] As an example, textual responses can be semantically segmented according to the sentence dimension to obtain multiple sentences.

[0246] S1803: Identify the first short sentence among multiple short sentences that requires additional information on other modalities.

[0247] As an example, a pre-trained language model can be used to determine whether multiple short sentences need to be supplemented with information from other modalities, in order to identify the first short sentence among the multiple short sentences that needs to be supplemented with information from other modalities.

[0248] As another example, based on preset strategies and rules, it can be determined whether multiple short sentences need to be supplemented with information from other modalities, so as to determine the first short sentence among multiple short sentences that needs to be supplemented with information from other modalities.

[0249] In some embodiments, a pre-trained language model or a preset strategy and rules can be used to determine the modality type that needs to be supplemented in the first short sentence, such as including but not limited to images, audio, video, vibration, indicator lights, AR / VR images, etc.

[0250] In some embodiments, the first short phrase that requires additional information on other modalities may also be marked, such as by using placeholders to mark the first short phrase.

[0251] It should be noted that in the embodiments of this application, there may be one or more first phrases that require supplementation with information of other modalities, and this is not limited. In the case where there are multiple first phrases that require supplementation with information of other modalities, the same processing method can be used to obtain the corresponding second modal information for each first phrase.

[0252] S1804: Obtain information about the second modality corresponding to the first short phrase.

[0253] As an example, a search engine can be invoked to search for modal information for the first short phrase based on placeholders and other markers, in order to obtain information about the second modality corresponding to the first short phrase.

[0254] As an example, the modality type to be supplemented can be determined first. Then, a search engine can be invoked to search for modality information based on placeholders and other markers, in order to obtain information about the second modality that matches the determined modality type. For instance, if the first sentence is "monkeys are a common name for some primates," and the determined modality type is an image modality, the first sentence can be input into the search engine, and the image channel can be requested to retrieve the most relevant image of a monkey. Similarly, if the first sentence is "giraffes are ruminant artiodactyls that live in Africa," and the determined modality type is an image modality, the server can input the first sentence into the search engine and request the image channel to retrieve the most relevant image of a giraffe.

[0255] It should be noted that the embodiments of this application do not limit the specific channels and processes for obtaining the information of the second modality corresponding to the first short sentence. For example, information of other modalities can be obtained through multi-channel search. With the advancement of technology, the information of the second modality corresponding to the first short sentence can also be generated through language models, depending on the specific situation.

[0256] Furthermore, the embodiments of this application do not limit the number of second modal information corresponding to a first phrase. For example, a first phrase may correspond to one second modal information or multiple second modal information. In the case where a first phrase corresponds to multiple second modal information, the modal types of the multiple second modal information may be the same (e.g., all are images, videos, audio, etc.) or different (e.g., respectively are images, audio, video, etc.), without specific limitation.

[0257] As described above, in the solutions provided by the actual embodiments of this application, the information of the first mode and the information of the second mode may be obtained by the server or by the electronic device.

[0258] As an example, please refer to Figure 19. Figure 19(a) and Figure 19(b) illustrate the acquisition and display process of two types of content provided in the embodiments of this application, with the first modality information and the second modality information being obtained by the server, and the first information being a question input by the user.

[0259] As shown in Figure 19(a), the electronic device can detect user actions, such as detecting user input of a question. After detecting the user's input question, the electronic device can send the user's input question to the server. After the server receives the user's input question from the electronic device, firstly, the server generates a text answer for the question based on a question-and-answer model; then, the server obtains the second modality information corresponding to the text answer and sends the text answer and the corresponding second modality information to the electronic device. Correspondingly, after receiving the text answer, the electronic device renders and displays the text answer, and while the user is browsing the text answer, based on the detected target text that the user is interested in, it obtains the second modality information corresponding to the target text from the second modality information corresponding to the text answer, and renders and displays the second modality information corresponding to the target text.

[0260] In the example shown in Figure 19(a), the server obtaining the second modal information corresponding to the text response may occur before the server sends the text response to the electronic device, or after the server sends the text response to the electronic device, or simultaneously with the server sending the text response to the electronic device; the electronic device displaying the text response may occur before the server sends the second modal information corresponding to the text response to the electronic device, or after the server sends the second modal information corresponding to the text response to the electronic device, or simultaneously with the server sending the second modal information corresponding to the text response to the electronic device; the electronic device detecting the target text that the user is interested in in the text response may occur before the server sends the second modal information corresponding to the text response to the electronic device, or after the server sends the second modal information corresponding to the text response to the electronic device, or simultaneously with the server sending the second modal information corresponding to the text response to the electronic device.

[0261] As shown in Figure 19(b), the electronic device can detect user operations, such as detecting user input of a question. After detecting the user's input question, the electronic device can send the user's input question to the server. After the server receives the user's input question from the electronic device, the server generates a text answer for the question based on a question-and-answer model and sends the text answer to the electronic device. Correspondingly, after receiving the text answer, the electronic device renders and displays the text answer. During the user's browsing of the text answer, based on the target text detected in the text answer that the user is interested in, the electronic device requests the server to obtain the second modality information corresponding to the target text. The server obtains the second modality information corresponding to the target text according to the electronic device's request and sends it to the electronic device. After receiving the second modality information corresponding to the target text, the electronic device displays the second modality information corresponding to the target text. As an example, please refer to Figure 20. Figures 20(a) and 20(b) illustrate the acquisition and display process of two other types of content provided in the embodiments of this application, using the acquisition of first modality information and second modality information by the electronic device as an example.

[0262] As shown in Figure 20(a), the electronic device can detect user actions, such as detecting user input of a question. After detecting the user's input question, firstly, the electronic device generates a text answer to the question based on a question-and-answer model; then, the electronic device obtains the second modality information corresponding to the text answer; finally, the electronic device renders and displays the text answer, and while the user is browsing the text answer, based on the target text that the user is interested in, it obtains the second modality information corresponding to the target text from the second modality information corresponding to the text answer, and completes the rendering and display of the second modality information corresponding to the target text.

[0263] In the example shown in Figure 20(a), the electronic device may acquire the information of the second modality corresponding to the text response before the electronic device displays the text response, after the electronic device displays the text response, or simultaneously with the electronic device displaying the text response, without limitation.

[0264] As shown in Figure 20(b), the electronic device can detect user actions, such as detecting user input of a question. After detecting the user's input question, the electronic device can generate and display a text answer based on a question-and-answer model. While the user is browsing the text answer, the electronic device can obtain and display the second modality information corresponding to the target text that the user is interested in from the detected text answer.

[0265] Of course, in some embodiments, the acquisition of information in the first modality and the acquisition of information in the second modality may also be accomplished collaboratively by the electronic device and the server. For example, the acquisition of information in the first modality may be the responsibility of the electronic device, and the acquisition of information in the second modality may be the responsibility of the server; or, the acquisition of information in the first modality may be the responsibility of the server, and the acquisition of information in the second modality may be the responsibility of the electronic device.

[0266] It is understood that, based on the content display process provided in the above embodiments of this application, the electronic device can accurately and promptly display second-modal information such as images, audio, video, vibration, indicator lights, and AR / VR images related to the target text, based on the user's browsing of text content, such as the target text that the user is interested in. This assists the user in accurately, fully, and promptly understanding the information in the content displayed by the electronic device, thereby improving the user experience. For example, the electronic device can identify the target text that the user is interested in based on the user's browsing position, click position, etc., and then accurately and promptly display second-modal information such as images, audio, video, vibration, indicator lights, and AR / VR images related to the target text. Therefore, the problem of insufficient correlation between the image and text as shown in Figure 1, which leads to viewing inconvenience for the user and is prone to user misunderstanding, will not occur. Furthermore, taking the image modality information as an example, based on the solution provided in this application, the target text and its corresponding second-modal information will always be displayed on the same interface. Therefore, the problem of the user being unable to understand the content by combining mathematical diagrams during text scrolling, as shown in Figure 2, which leads to viewing inconvenience and affects the user experience, will not occur.

[0267] In addition, in this embodiment, the information of the second modality displayed by the electronic device is obtained by sequentially semantically segmenting the answer to the user's input question into short sentences, and then matching the information of other modalities according to the short sentence dimension. Therefore, the information of the second modality is at the short sentence granularity and has a high degree of correlation with the text. Thus, it can accurately and fully assist the user in understanding the corresponding short sentences and contextual meaning, thereby improving the user experience.

[0268] It should be understood that the various solutions in the embodiments of this application can be used in a reasonable combination, and the explanations or descriptions of the various terms appearing in the embodiments can be referenced or explained to each other in the various embodiments, without limitation.

[0269] It should also be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0270] It is understood that, in order to implement the functions of any of the above embodiments, electronic devices or servers include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0271] This application embodiment can divide electronic devices or servers into functional modules. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0272] It should also be understood that the various modules in an electronic device or server can be implemented in software and / or hardware, without specific limitations. In other words, the electronic device or server is presented in the form of functional modules. Here, "module" can refer to application-specific integrated circuits (ASICs), circuits, processors and memory that execute one or more software or firmware programs, integrated logic circuits, and / or other devices that can provide the above functions.

[0273] In an alternative approach, when data transmission is implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are implemented. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disk (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0274] The steps of the methods or algorithms described in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disk, portable hard disk, CD-ROM, or any other form of storage medium well known in the art. One exemplary embodiment couples a storage medium to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Furthermore, the ASIC can reside in an electronic device or server. Alternatively, the processor and storage medium can exist as discrete components.

[0275] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

Claims

1. A method for displaying content, characterized in that, Applied to electronic devices, the method includes: Obtain information of the first modality, which is used to indicate the first content; The first content is displayed in the first modality; Identify the target content within the first set of content that is of interest to the user; Obtain information about the second modality corresponding to the target content, wherein the information about the second modality is used to indicate the second content; Perform a first operation to display the second content to the user in the second modality and / or the third modality.

2. The method according to claim 1, characterized in that, The execution of the first operation includes: The second content is displayed in the second mode via the display component of the electronic device; and / or, the second content is displayed in the second mode and / or the third mode via the display component of the cooperating device.

3. The method according to claim 2, characterized in that, The display component of the collaborative device displays the second content in the second mode and / or the third mode, including: Send the information of the second modality to the collaborative device so that the second content can be displayed in the second modality through the display component of the collaborative device; or, After converting the information of the second modality into information of the third modality, the information of the third modality is sent to the collaborative device so that the second content can be displayed in the third modality through the display component of the collaborative device; or... The information of the second modality is sent to the collaborative device so that the second content is displayed in the third modality through the display component of the collaborative device.

4. The method according to claim 3, characterized in that, The device information of the collaborative device is matched with the second mode and / or the third mode.

5. The method according to any one of claims 1-4, characterized in that, The information of the first modality is determined based on the first information input by the user.

6. The method according to any one of claims 1-5, characterized in that, The acquisition of information about the first mode includes: Generate, search, receive or read pre-stored information of the first modality from the server.

7. The method according to any one of claims 1-6, characterized in that, The target content that the user is interested in in the first content is determined according to at least one of the following methods: Detect the focusing position of the user's eyeballs; Receive user actions regarding the first content; Detects the position of the user's finger when swiping; Detect the content centered on the screen.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: acquiring information about a second modality related to the first content; The step of obtaining the information of the second modality corresponding to the target content includes: obtaining the information of the second modality corresponding to the target content from the information of the second modality related to the first content.

9. The method according to claim 8, characterized in that, The information about the second modality related to the first content is obtained through the following process: The first content is semantically segmented to obtain multiple sub-contents; Identify one or more first sub-contents among the plurality of sub-contents that require supplementation with information of other modalities; Obtain information about the second modality corresponding to the one or more first sub-contents; The target content overlaps with at least one character in one or more first sub-contents.

10. The method according to any one of claims 1-9, characterized in that, The display of the second content in the second modality includes: The second content is displayed in either a fixed display area or a non-fixed display area.

11. The method according to claim 10, characterized in that, The screen of the electronic device is a foldable screen, and displaying the second content in the second mode includes: When the screen is in a folded state, the second content is displayed in the non-fixed display area; When the screen is in the unfolded state, the second content is displayed in the fixed display area.

12. The method according to any one of claims 1-11, characterized in that, The display of the second content in the second modality includes: The second content is displayed in a floating mode, and the second content does not obscure the target content; or, The second content is displayed in a tiled manner, and the second content does not obscure the first content.

13. The method according to claim 12, characterized in that, The screen of the electronic device is a foldable screen, and displaying the second content in the second mode includes: When the screen is in a folded state, the second content is displayed in the floating display mode; When the screen is in the unfolded state, the second content is displayed in the tiled display mode.

14. The method according to any one of claims 1-13, characterized in that, The modal type of the first mode, the second mode, or the third mode includes any one or more of the following: image mode, video mode, audio mode, vibration mode, indicator light mode, and AR / VR image mode.

15. The method according to any one of claims 1-14, characterized in that, The method further includes: In response to a user's action on the second content, information in a fourth modality related to sub-content in the first content other than the target content is displayed, or sub-content in the first content at a first position related to the second content is displayed.

16. The method according to any one of claims 5-15, characterized in that, The first information includes any one or more of the following: images, videos, music, and keywords; the information of the first modality includes any one or more of the following related to the first information: introduction, analysis, and explanation. And / or, The first information includes a question, and the information of the first modality includes any one or more of the following in response to the question: analysis, explanation, answer, or response.

17. A method for displaying content, characterized in that, Applied to a server, the method includes: Receive the first information from the electronic device; Information of the first modality is obtained based on the first information, and the information of the first modality is used to indicate the first content; Information of the second modality is obtained based on information of the first modality, and information of the second modality is related to a portion of the content in the first content; Send the first modality information to the electronic device, and send the second modality information to the electronic device and / or the cooperating device.

18. The method according to claim 17, characterized in that, The step of obtaining information about the first modality based on the first information includes: Generate, search, and read the pre-stored information of the first modality based on the first information.

19. The method according to claim 17 or 18, characterized in that, The step of obtaining the information of the second mode based on the information of the first mode includes: Semantic segmentation is performed on the first content indicated by the information of the first modality to obtain multiple short sentences; Identify one or more first sentences from the plurality of short sentences that require additional information on other modalities; Generate, search, or retrieve pre-stored information about the second modality associated with the one or more first phrases.

20. The method according to any one of claims 17-19, characterized in that, The first information includes any one or more of the following: images, videos, music, and keywords; the information of the first modality includes any one or more of the following related to the first information: introduction, analysis, and explanation. And / or, The first information includes a question, and the information of the first modality includes any one or more of the following in response to the question: analysis, explanation, answer, or response.

21. The method according to any one of claims 17-20, characterized in that, The modal type of the first mode or the second mode includes any one or more of the following: image mode, video mode, audio mode, vibration mode, indicator light mode, and AR / VR image mode.

22. An electronic device, characterized in that, The electronic device includes: Memory is used to store computer program instructions; A processor for executing the computer program instructions to support the electronic device in implementing the method as described in any one of claims 1-16.

23. A server, characterized in that, The server includes: A transceiver is used to send and receive signals. Memory is used to store computer program instructions; A processor for executing the computer program instructions to support the server in implementing the method as described in any one of claims 17-21.

24. A content display system, characterized in that, The content display system includes: An electronic device for implementing the method as described in any one of claims 1-16; A server for implementing the method as described in any one of claims 17-21.

25. The content display system according to claim 24, characterized in that, The content display system also includes: One or more cooperating devices are used to display second content related to the first content during the process of the electronic device displaying the first content.

26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processing circuit, implement the method as described in any one of claims 1-16 or 17-21.

27. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-16 or 17-21.

Citation Information

Patent Citations

  • Method and apparatus for generating customized content based on user intent

    US20210209289A1

  • Methods and systems for suggesting an enhanced multimodal interaction

    US20230236858A1

  • Search method and electronic device

    WO2023029993A1

  • Information recommendation method and electronic device

    WO2024022258A1