Image display method, image generation method, and electronic device
By displaying target images and text related to users browsing multimedia content, the problem of insufficient diversity of lock screen wallpapers in the prior art is solved, improving user experience and providing traffic display opportunities.
Patent Information
- Application Number
- PCT/CN2024/116021
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-08-30
- Publication Date
- 2025-05-22
AI Technical Summary
The prior art cannot provide lock screen wallpapers that are highly correlated with users and are diverse, resulting in poor user experience.
By receiving the user's operations, the target image and target text related to the user's browsing multimedia content are displayed, and the target model is used to generate images and text based on the user's browsing data and display style.
It improves the diversity of lock screen, desktop and screen-out interfaces, enhances the relevance between users and interface content, improves the user experience, and provides traffic display opportunities.
Smart Images

Figure CN2024116021_22052025_PF_FP_ABST
Abstract
Description
Image display method, image generation method and electronic device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 15, 2023, with application number 202311523564.9 and application name “An image display method, image generation method and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of terminals, and in particular to an image display method, an image generation method, and an electronic device. Background Art
[0003] In the prior art, in terminals such as mobile phones and tablet computers, the wallpaper of the terminal lock screen interface is generally an image selected by the user, or an image randomly selected by the terminal from a preset wallpaper image library.
[0004] The method of using a user-selected image as the lock screen wallpaper lacks diversity because the lock screen wallpaper does not change after the user selects it. The method of using an image randomly selected by the terminal from a preset wallpaper image library as the lock screen wallpaper can continuously change the lock screen wallpaper to meet the diversity requirement, but these images used as lock screen wallpapers are often not highly relevant to the user, affecting the user experience.
[0005] It can be seen that in the existing method, the terminal cannot provide wallpapers that are highly relevant to the user and meet the diversity requirements.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide an image display method, an image generation method, and an electronic device, which solve the problems in the prior art of low correlation between wallpapers and users and low wallpaper diversity.
[0008] According to the first aspect of the embodiments of the present application, the present application provides an image display method, applied to a terminal, including: receiving a first operation of a user to open a first application; in response to the first operation, displaying a target image and a target text, wherein the target image is at least related to first data, and the first data is used to indicate the multimedia content browsed by the user using a second application, and the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, and the media data includes at least one of the text, audio, video and image of the multimedia content; the target text is at least related to second data, and the second data is used to indicate usage data when the user uses the second application to browse the multimedia content.
[0009] The technical solution provided in the first aspect above can display a target image related to multimedia content viewed by the user in other applications, as well as target text related to the user's usage data when browsing multimedia content in other applications, when a user opens a first application such as a desktop application, a lock screen application, an off-screen application, or a system application. In this way, target images and target text related to the multimedia content viewed by the user can be displayed on the interface of the lock screen application, desktop application, off-screen application, or system application. In other words, both the target image and target text are relevant to the user. If the user has a large amount of multimedia content browsed historically, corresponding target images and target text can be generated based on the multiple multimedia contents browsed by the user, thereby increasing the diversity of interfaces such as the lock screen, desktop, and off-screen. Such target images and target text avoid the monotony of template-generated images, thereby improving the user experience. In addition, the lock screen wallpaper, desktop wallpaper, and widgets on the off-screen interface can actually serve as traffic entry points to divert traffic to applications on the terminal, and such target images and target text also provide excellent traffic display opportunities.
[0010] In a possible implementation, the target image is further related to a target display style, and / or the target text is further related to a target display style.
[0011] In one possible implementation, the target display style is a display style selected by the user from a plurality of displayed alternative display styles; or, the target display style is a display style that best matches the user among preset display styles determined based on multimedia content browsed by the user.
[0012] In a possible implementation, the method further includes: displaying a second control; receiving a fourth operation of the user on the second control; and displaying multimedia content corresponding to the target image in the second application in response to the fourth operation.
[0013] That is, after displaying the target image and target text, a jump control can also be displayed. After the user performs the fourth operation on the jump control, the user can jump to the corresponding multimedia application and display the multimedia content corresponding to the target image. This plays a good role in attracting users and can increase user access to multimedia applications.
[0014] In a possible implementation, the method further includes: sending the target display style to the server.
[0015] In a possible implementation, before displaying the target image and target text, the method further includes: receiving the target image and target text from a server.
[0016] That is, the terminal may send the target display style to the server, and the server provides the target image and target text after receiving the target display style.
[0017] In one possible implementation, the method further includes: displaying a first control; receiving a second operation of the user on the first control; and sending first information to the server in response to the second operation, wherein the first information is used to instruct the server to delete the target image and target text in the image library corresponding to the user, wherein the image library stores multiple images and texts generated based on usage data when the user browses multimedia content and the multimedia content browsed by the user.
[0018] That is, if the target image and target text are provided by a server, the server also stores a user-specific image library, which contains multiple target images and target text generated using the above method. After displaying the target image and target text, the terminal can also display a delete control. When the user performs a second operation on the delete control, the terminal can send a delete message to the server, instructing the server to delete the displayed target image and target text from the user's corresponding image library. This improves the user experience by preventing content that the user is not interested in from being pushed to the user through interaction with the terminal.
[0019] In one possible implementation, the above method also includes: obtaining first data and second data; using the first data, second data and target display style as input, and generating a target image and target text through a target model, wherein the target model has the function of generating corresponding images and texts based on the input multimedia content and display style browsed by the user.
[0020] That is, the terminal can use the first data, the second data, and the target display style as input to generate the target image and target text through the target model. Generating the target image and target text through the target model can increase the diversity of the target images and target text, thereby improving the user experience. The target model is a model that can generate corresponding images and text based on the multimedia content and display style browsed by the user. In this way, the target image and / or target text that conforms to the target display style can be generated according to the user's needs, thereby making the pushed target image and target text more compatible with the user and improving the user experience.
[0021] In one possible implementation, the target model includes an image generation model and a text generation model. Using the first data, the second data, and the target display style as input, the target image and target text are generated using the target model. This includes: using the first data and the target display style as input, using the image generation model to generate the target image; and using the text generation model to generate the target text, using the second data as input. This can achieve better generation results, ensuring that the generated target image and target text better meet requirements.
[0022] In one possible implementation, when the first data includes at least metadata, the method further includes: obtaining a feature vector of the first data; generating a first descriptive word based on the metadata and the target display style, and generating a second descriptive word based on the second data, wherein the first descriptive word is used to describe the characteristics of the image to be generated, and the second descriptive word is used to describe the characteristics of the text to be generated; using the first data and the target display style as input, and generating a target image through an image generation model, including: using the feature vector of the first data and the first descriptive word as input, and generating the target image through the image generation model; using the second data as input, and generating the target text through a text generation model, including: using the second descriptive word as input, and generating the target text through the text generation model. This allows the target model to determine the characteristics of the image and text to be generated, thereby outputting a target image and target text that better meet the requirements.
[0023] In one possible implementation, the second descriptor can be generated based on the second data and metadata, or the second descriptor can be generated based on the second data, metadata, and medium data. In this way, the target text generated using the second descriptor can include more content, further improving the interpretability of the target image and enhancing the user experience.
[0024] In one possible implementation, the method further includes: displaying a first control; receiving a second operation of the user on the first control; in response to the second operation, deleting a target image and a target text in an image library corresponding to the user; the image library stores a plurality of images and texts generated based on usage data when the user browses multimedia content and the multimedia content browsed by the user.
[0025] In other words, if the target image and target text are provided by the terminal itself, the terminal also stores a user-specific image library, which contains multiple target images and target text generated using the above method. After displaying the target image and target text, the terminal can also display a delete control. After the user performs a second operation on the delete control, the terminal can delete the displayed target image and target text from the image library. This prevents users from being pushed content they are not interested in, improving the user experience.
[0026] In one possible implementation, the first application is a lock screen application; the displaying of the target image and target text includes: displaying the target image in the lock screen interface; receiving a third operation of the user in the lock screen interface, and displaying the target text.
[0027] That is, the first application opened by the user may be a lock screen application. Specifically, the target image and target text may be displayed by displaying the target image on the lock screen interface and then displaying the target text upon receiving a third user action (such as swiping up). In this way, the target text does not obscure the main body of the target image, thereby improving the user experience.
[0028] In one possible implementation, the above-mentioned display of the target image and target text includes: the first application is a desktop application, and the target image and target text are used as desktop wallpaper and displayed on the desktop; or, the first application is an off-screen application, and the target image and target text are used as widgets and displayed on the off-screen interface; or, the first application is a system application, and the target image and target text are used as startup images and displayed on the startup interface of the first application.
[0029] That is to say, the first application opened by the user can be a desktop application, in which case the target image and target text can be displayed as desktop wallpaper. The first application opened by the user can also be an off-screen application, in which case the target image and target text can be displayed as widgets on the off-screen interface. The first application opened by the user can be a system application, in which case the target image and target text can be displayed as startup images on the startup interface of the system application. The method of the embodiment of the present application can be applied to a variety of scenarios, improving the relevance of the images displayed in each scenario to the user and the diversity of the displayed content, thereby improving the user experience.
[0030] According to the second aspect of the embodiment of the present application, the present application also provides an image generation method, which is applied to a server, including: obtaining first data and second data of a terminal, wherein the first data is used to indicate the multimedia content browsed by a user corresponding to the terminal using a second application, wherein the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data including at least one of the text, audio, video and image of the multimedia content; the second data is used to indicate usage data when the user uses the second application to browse the multimedia content; based on the first data and the second data, generating a target image and target text; wherein the target image is at least related to the first data, and the target text is at least related to the second data; and sending the target image and target text to the terminal.
[0031] In other words, the server can generate target images and target text related to the multimedia content browsed by the user on the terminal, and send the generated target images and target text to the terminal. In this way, the server can provide the terminal with a large number of images that are strongly related to the user, improving the user experience of the terminal.
[0032] In a possible implementation, the target image is further related to a target display style, and / or the target text is further related to a target display style.
[0033] In a possible implementation, the method further includes: receiving a target display style sent by the terminal, where the target display style is a display style selected by the user from a plurality of candidate display styles displayed on the terminal.
[0034] In a possible implementation, the target display style is a display style that has the highest matching degree with the user among preset display styles determined based on multimedia content browsed by the user.
[0035] In one possible implementation, the above-mentioned generation of the target image and target text based on the first data and the second data includes: taking the first data, the second data and the target display style as input, and generating the target image and target text through a target model, wherein the target model has the function of generating corresponding images and texts based on the multimedia content and display style browsed by the input user.
[0036] In other words, the server can use the first data, the second data, and the target display style as input and generate the target image and target text using a target model. Generating target images and target text using a target model can increase the diversity of target images and target text, enhancing the user experience. A target model is a model that generates corresponding images and text based on the multimedia content and display style viewed by the user. This allows the generation of target images and / or target text that match the target display style based on user needs, ensuring a higher degree of match between the pushed target image and target text and the user, improving the user experience.
[0037] In one possible implementation, the target model includes an image generation model and a text generation model. Using the first data, the second data, and the target display style as input, the target image and target text are generated using the target model. This includes: using the first data and the target display style as input, using the image generation model to generate the target image; and using the text generation model to generate the target text, using the second data as input. This can achieve better generation results, ensuring that the target image and target text generated by the server better meet requirements.
[0038] In one possible implementation, when the first data includes at least metadata, the method further includes: obtaining a feature vector of the first data; generating a first descriptive word based on the metadata and the target display style, and generating a second descriptive word based on the second data, wherein the first descriptive word is used to describe the characteristics of the image to be generated, and the second descriptive word is used to describe the characteristics of the text to be generated; using the first data and the target display style as input, and generating a target image through an image generation model, including: using the feature vector of the first data and the first descriptive word as input, and generating a target image through an image generation model; using the second data as input, and generating a target text through a text generation model, including: using the second descriptive word as input, and generating a target text through a text generation model. This allows the target model on the server to determine the characteristics of the image and text to be generated, thereby outputting a target image and target text that better meet the requirements.
[0039] In one possible implementation, the method further includes: receiving a first message from a terminal; and in response to the first message, deleting a target image and target text from an image library corresponding to the user; the image library storing a plurality of images and text generated based on the user's usage data when browsing multimedia content and the multimedia content browsed by the user. This prevents content that the user is not interested in from being pushed to the user, thereby improving the user experience.
[0040] According to a third aspect of the embodiments of the present application, the present application further provides a device having the function of implementing the electronic device behavior in the method described in the first or second aspect above. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, for example, a processing module, a sending module, and a display module.
[0041] According to a fourth aspect of the embodiments of the present application, the present application also provides an electronic device, including a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program to implement the method described in the first aspect above.
[0042] According to a fifth aspect of an embodiment of the present application, the present application further provides a server comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the method of the second aspect above.
[0043] According to the sixth aspect of the embodiments of the present application, the present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to implement the method of the first aspect or the second aspect mentioned above.
[0044] According to the seventh aspect of the embodiments of the present application, the present application also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code is running in an electronic device, the processor in the electronic device executes the method of the first aspect or the second aspect above.
[0045] According to the eighth aspect of the embodiment of the present application, the present application also provides a chip, which includes a memory and a processor, the memory is used to store computer programs, and the processor is used to call and run computer programs from the memory to execute the method described in the first or second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] FIG1 is a schematic diagram of a lock screen interface provided by the related art;
[0047] FIG2 is a schematic diagram of another screen-off interface provided by the related art;
[0048] FIG3 is a simplified schematic diagram of a system architecture for applying a method according to an embodiment of the present application;
[0049] FIG4 is a schematic diagram of the composition of a terminal provided in an embodiment of the present application;
[0050] FIG5 is a schematic diagram of the composition of a server provided in an embodiment of the present application;
[0051] FIG6 is a flowchart of an image display method provided in an embodiment of the present application;
[0052] FIG7 is a flowchart of another image display method provided in an embodiment of the present application;
[0053] FIG8 is a schematic diagram of a terminal interface provided in an embodiment of the present application;
[0054] FIG9 is a schematic diagram of generating a target image and target text according to an embodiment of the present application;
[0055] FIG10 is a flowchart of an image generation method provided in an embodiment of the present application;
[0056] FIG11 is a schematic diagram of another terminal interface provided in an embodiment of the present application;
[0057] FIG12 is a schematic diagram of another terminal interface provided in an embodiment of the present application. DETAILED DESCRIPTION
[0058] Lock screen, home screen, and screen off functions are widely used on mobile phones, computers, tablets, car computers, TVs, and other devices, and are essential for normal terminal use. For example, when using a terminal, users must first enter the lock screen interface and then unlock the device to use other terminal functions. For another example, when using an application installed on the terminal, users generally need to first enter the desktop and open the application by clicking the application icon displayed on the desktop.
[0059] Based on the above, the information display of the lock screen interface, desktop and off screen interface is an important issue that needs to be considered. In the relevant technology, there are several ways to display information.
[0060] First, the user selects an image stored in the terminal as the lock screen wallpaper or desktop wallpaper to display image information on the lock screen interface and desktop; the user can also select a theme as the theme of the mobile phone, where the theme is designed and developed by a company or individual developer to personalize the terminal skin interface such as the lock screen, wallpaper, icons, notification bar, SMS, dialing, contacts, settings, etc.; the user also selects a widget from the terminal's preset widgets and displays the widget on the off-screen interface to transmit relevant information through the widget, such as the clock widget can transmit time information.
[0061] The lock screen wallpaper, desktop wallpaper, themes, and widgets on the off-screen interface do not change after the user selects them, which results in a lack of diversity in the lock screen interface, desktop, and off-screen interface, making it difficult to arouse user interest. However, the lock screen wallpaper, desktop wallpaper, and widgets on the off-screen interface can actually serve as traffic entrances to divert traffic to applications on the terminal, which wastes a good traffic display opportunity.
[0062] Second, lock screen wallpapers are random images or illustrated magazines pushed by the terminal's system or applications. This often results in lock screen content being less relevant to the user. Users are unfamiliar with the content, and even if they spend some time reading it, they often find it difficult to develop interest in the content displayed on the lock screen due to its low relevance. They often quickly skip over the content, wasting the traffic exposure opportunities brought by the lock screen wallpaper, which serves as a traffic entry point.
[0063] Moreover, the random images or illustrated magazines displayed on the lock screen interface often come from a single wallpaper provider, which makes the lock screen source single and cannot meet the personalized needs of users.
[0064] Third, in addition to displaying wallpaper on the lock screen, another lock screen display solution is to display the interface information of a certain application running in the background on the lock screen. For example, when a user is playing a song using music software on the terminal, as shown in Figure 1, the lock screen can display the song title, artist name, album name, album cover, lyrics, etc., and can also display controls for controlling song playback.
[0065] In this solution, the lock screen interface is generated by a template and lacks diversity.
[0066] It can be seen that in the related art, the display solutions for the lock screen interface, desktop and screen-off interface have problems such as templateization, monotony and lack of diversity, and the displayed content lacks relevance to the user and is difficult to arouse the user's interest.
[0067] Based on this, the present application proposes an image display method and an image generation method. When a user opens a first application such as a desktop application, a lock screen application, an off-screen application, or a system application, a target image related to the multimedia content browsed by the user on other applications can be displayed, as well as target text related to the usage data of the user when browsing multimedia content on other applications.
[0068] In other words, the image display method applied to the terminal includes: receiving a first operation of a user to open a first application; displaying a target image and a target text in response to the first operation; wherein the target image is at least related to the first data, the first data is used to indicate the multimedia content browsed by the user using the second application, and the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data includes at least one of the text, audio, video and image of the multimedia content; the target text is at least related to the second data, and the second data is used to indicate usage data when the user uses the second application to browse the multimedia content.
[0069] In this way, target images and target texts related to the multimedia content browsed by the user can be displayed on the interfaces of lock screen applications, desktop applications, off-screen applications, or system applications. In other words, both the target image and the target text are related to the user. Since users often have a lot of multimedia content browsed in the past, corresponding target images and target texts can be generated based on the multiple multimedia contents browsed by the user, which also improves the diversity of interfaces such as the lock screen, desktop, and off-screen. In addition, such target images and target texts avoid the monotony of images generated by templates, can increase the user's usage rate of interfaces such as the lock screen, desktop, and off-screen, and improve the user experience.
[0070] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in the technical solutions of this application are in compliance with relevant laws and regulations and do not violate public order and good morals. For example, in the technical solutions of this application, the processing of user personal information is performed with the user's authorization, which is explained here and will not be repeated below.
[0071] For the usage scenarios of the present application, the above-mentioned first application can be an application that can display an image when it is turned on. For example, the first application in this application can be a lock screen application, a desktop application, an off-screen application, a system application, and the like. That is, the present application can be applied to wallpaper display scenarios of interfaces such as the lock screen, desktop, off-screen, and application startup. For example, when the first application is a lock screen application, the target image and target text can be displayed as a lock screen wallpaper on the lock screen interface; when the first application is a desktop application, the target image and target text can be combined and displayed as a desktop wallpaper; when the first application is any application of the desktop application, the lock screen application, or the off-screen application, a corresponding theme can be generated based on the target image and target text, and the target image and target text can be displayed in the above three applications; when the first application is a system application, the target image and target text can be displayed on the startup interface of the system application, that is, the target image and target text can be combined as the wallpaper of the application startup interface. In addition, when there is cooperation between the terminal manufacturer and the third-party application installed on the terminal, the target image and target text can also be displayed on the startup interface of the third-party application other than the system application.
[0072] System applications refer to applications developed by terminal manufacturers, such as applications that are pre-installed when the terminal is purchased.
[0073] The off-screen interface is the display interface for system applications. It's the interface that a terminal may enter before entering the lock screen when in the locked state. This interface typically supports widgets. Widgets can also be referred to as widgets, small widgets, small widgets, and widgets. As shown in Figure 2, 201 is a widget that displays the time. In this embodiment of the present application, a target image and target text can be displayed as widgets on the off-screen interface.
[0074] The first operation can also be an operation that triggers the opening of the first application. For example, in a lock screen scenario, the first operation can be an operation that can wake up the screen, such as pressing the lock screen button of the terminal, lifting the terminal, etc. In a desktop scenario, the first operation can be an operation to return to the desktop, such as closing an application, returning to the desktop, etc. In an off-screen scenario, the first operation can be an operation that can wake up the off-screen interface, such as locking the screen or lifting the terminal, etc. In the case where the first application is a system application, the first operation can be an operation on the icon of the system application, an operation to switch from the background interface to the system application, etc. Of course, as described above, the first application can also be other applications that can display images when opened, such as third-party applications installed on the terminal and released by application developers.
[0075] The second application may be an application that includes multimedia content. For ease of understanding, the second application will be referred to as a multimedia application below. The multimedia application mentioned in the embodiments of the present application may be a multimedia application developed by a terminal manufacturer, such as a video, music, reading, podcast, and game application pre-installed on the terminal. In this embodiment, a target image may be generated based on the user's usage record of the multimedia application developed by the terminal manufacturer, which may increase the traffic of these multimedia applications and increase the user's frequency of use of these multimedia applications.
[0076] When there is cooperation between the terminal manufacturer and the application developer, and the user agrees to use the usage data of the application developed by the application developer to generate wallpapers or widgets for the lock screen, desktop, and off-screen interfaces, the multimedia application can also be a multimedia application developed by the application developer, such as a third-party video application installed on the terminal.
[0077] In this application, users can browse multimedia content through the multimedia application in the terminal. Multimedia applications are applications that contain multimedia content, and multimedia content can include audio, video, text, and so on. For example, when the multimedia application is a video application and the multimedia content is video, users can browse videos through the video application on the terminal. For example, when the multimedia application is a music application and the multimedia content is music, users can listen to music through the music application on the terminal. For example, when the multimedia application is a reading application and the multimedia content is the text in a book, users can read books through the reading application on the terminal. For example, when the multimedia application is a game application and the multimedia content is screenshots and introductions corresponding to the game, users can browse game introductions, install games, enter game applications to play, and so on through the game application on the terminal. It should be noted that the game application here does not refer to the application corresponding to the game itself, but to the application that manages the games installed on the terminal. In some terminals, this game application is called a "game center."
[0078] As for the relationship between the first application and the second application, in the above example, in the lock screen, desktop and screen off scenarios, the first application and the second application can be different applications, for example, the first application is a lock screen application and the second application can be a video application. In addition, in the scenario where the first application is a system application, the first application and the second application can be the same application or different applications. For example, when the first application is a clock application and the second application is a multimedia application, then the first application and the second application are different applications. When the first application is a multimedia application, the second application can be the same multimedia application as the first application or a multimedia application different from the first application. For example, when the first application is a video application developed by the terminal manufacturer, the second application can be a music application developed by the terminal manufacturer, the second application can also be the video application, or the second application can be a third-party video application installed on the terminal.
[0079] For the displayed target image and target text, the target image is related to the multimedia content that the user has browsed, that is, the target image has the same theme as the multimedia content, or is related to the specific content of the multimedia content. This application does not limit the specific form of the target image; any image related to the multimedia content can be used as the target image of this application.
[0080] For example, if the multimedia content browsed by the user is a video about the scenery of a certain area, the target image may be an image of the scenery of the area, for example, the target image may be a video frame in the video, the target image may also be a classic image mentioned in the subtitles of the video, the target image may also be a publicly available image of the scenery of the area retrieved from the Internet, or an image related to the content in the video generated based on one or more of the video frames, audio, and subtitles in the video.
[0081] For example, if the multimedia content a user is browsing is a song, the target image can be an image related to the song, such as the song's album cover. If the song has a corresponding music video (MV), the target image can be a video frame from the music video. For example, if the song is about youth, the target image can be a usable image related to youth found online. The target image can also be an image related to the song generated based on one or more of the song's audio, lyrics, and MV video frames.
[0082] For example, if the multimedia content a user is browsing is a book, the target image can be an image related to the book, such as the book cover. The target image can also be an image related to the book retrieved from the Internet. The target image can also be an image related to the book generated based on one or more of the book cover or the text in the book.
[0083] For another example, if the multimedia content browsed by the user is a certain game, the target image may be an image related to the game. Similar to the above example, the target image may be a promotional video or promotional image of the game. The target image may also be an image related to the game retrieved from the Internet. The target image may also be an image related to the game generated based on one or more of the video frames of the promotional video of the game, the audio corresponding to the promotional video, and the game introduction.
[0084] For the example where the target image is generated based on multimedia content, the specific method for generating the image will be described in detail below and will not be described here. In addition, the above examples are only for the purpose of helping to understand the present solution and do not represent a limitation on the embodiments of the present application.
[0085] Target text refers to text related to the user's usage data when browsing multimedia content. Usage data, or data describing a user's browsing behavior, can include information such as when the user browsed the multimedia content, the user's browsing progress, whether the user has favorited or liked the multimedia content, whether the user has commented on the multimedia content, and the identifier of the multimedia content corresponding to the user's behavior. The target text can be related to the usage data by including specific usage data and the title of the multimedia content corresponding to the usage data. The title of the multimedia content corresponding to the usage data can be determined based on the identifier of the multimedia content included in the usage data. For example, the target text could be "You browsed 'Challenger' on January 1, 2023" or "You browsed 'Challenger' on January 1, 2023. Come and reminisce." For another example, if the target image is related to the content the user was watching before locking the screen, the target text could be "You are watching 'Challenger'." Furthermore, to facilitate user recall of related content, the target text can also include data describing the multimedia content, such as metadata and media data of the multimedia content. The metadata and media data of multimedia content will be described in detail below and will not be elaborated here.
[0086] By displaying the target text, users can be reminded that they have browsed the multimedia content corresponding to the target image, increasing the interpretability of the target image. The target text can make users aware of the content of the target image and let them know that the target image is related to the multimedia content they are browsing, thereby improving the user experience.
[0087] In addition, in addition to being related to the above-mentioned data, in some embodiments of the present application, the target image and target text may also be related to the target display style, and / or the target text may also be related to the target display style.
[0088] Regarding the multimedia content related to the displayed target image and target text, the displayed target image and target text may be related to the user's history of browsing multimedia content over the past period of time. The specific method will be described in detail later and will not be described here. In addition, the displayed target image and target text may also be related to the multimedia content that the user was browsing before locking the screen, closing the application, or returning to the desktop. For example, if it is determined that the user has browsed new multimedia content, a target image may be generated based on the multimedia content, and a target text may be generated based on the usage data of the multimedia content. After opening the lock screen interface, the off-screen interface, the application, or the desktop, the target image and target text are displayed.
[0089] For the users in this application, in some embodiments, the user can correspond to the terminal device itself. Based on this, the target image and target text can be generated only based on the local browsing history of the second application on the one terminal, that is, the browsing history of the second application on the terminal is used as data related to the user to generate the target image and target text. In some other embodiments, the user can correspond to the user account, the terminal can log in to the user account, and different terminals can log in to the same user account. Based on this, the target image and target text can also be generated based on the browsing history of the second application on the terminal logged in to the same user account, that is, the browsing history of the second application on multiple terminals logged in to the same account can be used as data related to the user to generate the target image and target text.
[0090] For the application device of the method of the present application, the method of the present application can be executed by the terminal alone, which can protect user privacy. The solution of the present application can also be implemented through the interaction between the terminal and the server. Implementing the solution of the present application through the interaction between the terminal and the server can reduce the processing burden of the terminal and save terminal computing power.
[0091] Next, the embodiment of the present application will be explained by taking the interaction between the terminal and the server and applying the method of the embodiment of the present application to the lock screen interface as an example.
[0092] As shown in Figure 3, Figure 3 is a simplified schematic diagram of a system architecture provided by an embodiment of the present application to which the above method can be applied. As shown in Figure 3, the system includes at least a terminal 31 and a server 32.
[0093] In a specific implementation, the terminal 31 may be a mobile phone, a tablet computer, a handheld computer, a personal computer (PC), a cellular phone, a personal digital assistant (PDA), a wearable device (such as a smartwatch), a smart home device (such as a television), an in-vehicle computer, a game console, and an augmented reality (AR) or virtual reality (VR) device. This embodiment does not impose any special restrictions on the specific device form of the terminal 31. The terminal 31 may include one or more applications. The application may be a system application, such as a lock screen application, a desktop application, an off-screen application, and a video application, reading application, game application, and music application pre-configured on the mobile phone by the terminal manufacturer. The application may also be a third-party application.
[0094] In addition, the system may include multiple terminals, and the server 32 may interact with the multiple terminals at the same time to provide target images and target texts to the multiple terminals.
[0095] Please refer to Figure 4, which is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The methods in the following embodiments can be implemented in a terminal having the above hardware structure.
[0096] As shown in FIG4 , the terminal 31 may include a processor 110, an external memory interface 120, an internal memory 121, a USB interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a radio frequency module 150, a communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a SIM card interface 195. The sensor module may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and the like.
[0097] The touch sensor 180K, microphone 170C, antenna 1, antenna 2, RF module 150, and communication module 160 may serve as input devices of terminal 31 for receiving information input by a user or other devices. The speaker 170A, receiver 170B, and display screen 194 may serve as output devices of terminal 31 for outputting information input by a user or provided to a user, as well as various menus of terminal 31.
[0098] The illustrated structure of the embodiment of the present invention does not limit the terminal 31. It may include more or fewer components than shown, or some components may be combined or separated, or arranged differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0099] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into the same processor.
[0100] The controller is the decision-maker that directs the various components of terminal 31 to coordinate operations according to instructions. It serves as the nerve center and command center of terminal 31. Based on instruction opcodes and timing signals, the controller generates operational control signals to control instruction fetching and execution.
[0101] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This can store instructions or data that have just been used or are being recycled by the processor. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids duplicate accesses, reduces processor latency, and thus improves system efficiency.
[0102] In the embodiment of the present application, the processor 110 can be used to execute the steps of the embodiment of the present application, such as controlling the display screen 194 to display the target image and target text, and after receiving a first user operation, processing the first operation to trigger the process of controlling the display screen 194 to display the target image and target text. In addition, the processor 110 can also implement functions such as obtaining a target style through the steps of the following embodiments, which are described in detail below.
[0103] In some embodiments, the processor 110 may include an interface, which may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0104] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor may include multiple I2C bus lines. The processor can be coupled to a touch sensor, a charger, a flash, a camera, etc. via different I2C bus interfaces. For example, the processor can be coupled to a touch sensor via an I2C interface, enabling communication between the processor and the touch sensor via the I2C bus interface, thereby implementing the touch function of terminal 31.
[0105] The I2S interface can be used for audio communication. In some embodiments, the processor can include multiple I2S buses. The processor can be coupled to the audio module via the I2S bus to enable communication between the processor and the audio module. In some embodiments, the audio module can transmit audio signals to the communication module via the I2S interface, enabling the function of answering calls through a Bluetooth headset.
[0106] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module and the communication module can be coupled via a PCM bus interface. In some embodiments, the audio module can also transmit audio signals to the communication module via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication, and the sampling rates of the two interfaces are different.
[0107] The UART interface is a universal serial data bus used for asynchronous communication. This bus is a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is generally used to connect the processor and the communication module 160. For example, the processor communicates with the Bluetooth module via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module can transmit audio signals to the communication module via the UART interface, enabling the function of playing music through Bluetooth headphones.
[0108] The MIPI interface can be used to connect the processor to peripheral devices such as displays and cameras. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor and camera communicate via the CSI interface to enable the camera function of terminal 31. The processor and display communicate via the DSI interface to enable the display function of terminal 31.
[0109] The GPIO interface is software-configurable. It can be configured as either a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor to a camera, display, communication module, audio module, sensor, etc. The GPIO interface can also be configured as an I2C interface, I2S interface, UART interface, MIPI interface, etc.
[0110] The USB interface 130 can be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface can be used to connect a charger to charge the terminal 31, or to transfer data between the terminal 31 and peripheral devices. It can also be used to connect headphones to play audio. It can also be used to connect other electronic devices, such as AR devices.
[0111] The interface connection relationship between the modules shown in the embodiment of the present invention is for illustrative purposes only and does not limit the structure of the terminal 31. The terminal 31 may adopt different interface connection modes or a combination of multiple interface connection modes in the embodiment of the present invention.
[0112] The charging management module 140 is used to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module can receive charging input from the wired charger via a USB port. In some wireless charging embodiments, the charging management module can receive wireless charging input via the wireless charging coil of the terminal 31. While the charging management module is charging the battery, it can also provide power to the terminal device through the power management module 141.
[0113] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module receives input from the battery and / or the charging management module and provides power to the processor, internal memory, external memory, display, camera, and communication module. The power management module can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some embodiments, the power management module 141 can also be provided in the processor 110. In some embodiments, the power management module 141 and the charging management module can also be provided in the same device.
[0114] The wireless communication function of the terminal 31 can be implemented through antenna 1, antenna 2, radio frequency module 150, communication module 160, modem and baseband processor.
[0115] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in Terminal 31 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, a cellular network antenna can be reused as a wireless local area network diversity antenna. In some embodiments, the antenna can be used in conjunction with a tuning switch.
[0116] The RF module 150 can provide a communication processing module for wireless communication solutions including 2G / 3G / 4G / 5G applied on the terminal 31. It can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The RF module receives electromagnetic waves from the antenna 1, and filters, amplifies and processes the received electromagnetic waves, and transmits them to the modem for demodulation. The RF module can also amplify the signal modulated by the modem and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some functional modules of the RF module 150 can be set in the processor 150. In some embodiments, at least some functional modules of the RF module 150 can be set in the same device as at least some modules of the processor 110.
[0117] The modem may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker, a receiver, etc.) or displays an image or video through a display screen. In some embodiments, the modem may be an independent device. In some embodiments, the modem may be independent of the processor and be set in the same device as the radio frequency module or other functional modules.
[0118] The communication module 160 can provide a communication processing module for wireless communication solutions including wireless local area networks (WLAN), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied on the terminal 31. The communication module 160 can be one or more devices integrating at least one communication processing module. The communication module receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signal, and sends the processed signal to the processor. The communication module 160 can also receive the signal to be sent from the processor, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0119] In some embodiments, the antenna 1 of the terminal 31 is coupled to the radio frequency module, and the antenna 2 is coupled to the communication module. This allows the terminal 31 to communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS) and / or satellite based augmentation system (SBAS).
[0120] The terminal 31 can communicate with the server 32 through the communication function of the terminal, so that the target image and target text can be obtained from the server 32 through the above module. In some embodiments, the mobile phone can also send the target display style, user usage data of the multimedia content in the second application, etc. to the server through the above module.
[0121] Terminal 31 implements display functions through a GPU, display screen 194, and an application processor. For example, these components can be used to display target images and target text. In some embodiments, the mobile phone can also use the above modules to display controls, alternative display styles, etc., thereby completing the interaction between the mobile phone and the user. The GPU is a microprocessor for image processing that connects the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.
[0122] Display screen 194 is used to display images, videos, and the like. The display screen includes a display panel. The display panel can be an LCD (liquid crystal display), an OLED (organic light-emitting diode), an active-matrix organic light-emitting diode (AMOLED), a MiniLED, a MicroLED, a Micro-oLED, or a quantum dot light-emitting diode (QLED). In some embodiments, terminal 31 can include one or N display screens, where N is a positive integer greater than one.
[0123] Still as shown in FIG4 , the terminal 31 can implement the shooting function through the ISP, camera 193 , video codec, GPU, display screen and application processor.
[0124] The ISP processes data fed back by the camera. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located in camera 193.
[0125] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal 31 may include 1 or N cameras, where N is a positive integer greater than 1.
[0126] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal 31 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0127] Video codecs are used to compress or decompress digital video. Terminal 31 may support one or more codecs. This allows Terminal 31 to play or record videos in various encoding formats, such as MPEG1, MPEG2, MPEG3, and MPEG4.
[0128] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications such as image recognition, face recognition, speech recognition, and text comprehension on Terminal 31.
[0129] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 31. The external memory card communicates with the processor via the external memory interface to implement data storage functions. For example, files such as music and videos can be stored in the external memory card.
[0130] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the mobile phone by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a lock screen function, a desktop function, a screen off function), etc. The data storage area can store data created during the use of the mobile phone (such as a target image and target text to be displayed), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0131] In the embodiment of the present application, the internal memory 121 may specifically include RAM (random access memory) and ROM (read-only memory). RAM is an internal memory that directly exchanges data with the processor 110, also called main memory (or internal memory). It can be read and written at any time, and the speed is very fast. It is usually used as a temporary data storage medium for the operating system or other running programs. The data stored in ROM can be easily read out, unlike RAM, which can be quickly and conveniently rewritten. However, the data stored in ROM is relatively stable, and the stored data will not change after power failure; its structure is relatively simple and it is easy to read, so it is often used to store various fixed programs and data.
[0132] RAM and ROM may include one or more partitions.
[0133] Taking ROM as an example, as shown in Figure 2, ROM can include system partitions (e.g., System partition and Recovery partition), program partitions (e.g., Data partition), and storage partitions (e.g., SDCard partition). The system partition can be used to store the operating system (e.g., Android system), restore and backup systems, swap space, hardware underlying space, and other resources; the program partition is used to store third-party apps installed on the terminal. For each app, the terminal creates a corresponding Data directory within the Data partition. For example, if an app's package name is weixin.com, a directory named weixin.com can be created within the Data partition. Application data generated by the app, such as chat logs and transferred files, is stored in the weixin.com directory. The app can only access data in this directory and cannot access directories of other apps. The storage partition is equivalent to the "portable hard drive" recognized by the phone when connected to a PC. This space is at the user's disposal and can store data such as large game data packages, music, pictures, and videos.
[0134] Terminal 31 can also dynamically create new partitions in RAM or ROM. For example, if 2GB of ROM is unoccupied, Terminal 31 can create a partition named "aaa" in this 2GB. Of course, Terminal 31 can also dynamically destroy a previously created partition, and the data within the partition will also be destroyed when the new partition is destroyed.
[0135] In the embodiment of the present application, when the terminal 31 is restored to factory settings, one or more partitions in the ROM of the terminal 31 are generally formatted. For example, the terminal 31 may format the Data partition in the ROM, deleting all files and folders in the Data partition, thereby clearing all applications installed in the Data partition and the data in each application, thereby restoring the terminal 31 to the state when it was sold by the factory.
[0136] The terminal 31 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0137] The audio module is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module can also be used to encode and decode audio signals. In some embodiments, the audio module can be provided in the processor 110, or some functional modules of the audio module can be provided in the processor 110.
[0138] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The terminal 31 can listen to music or listen to hands-free calls through the speaker.
[0139] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the terminal 31 receives a call or voice message, the voice can be heard by placing the receiver close to the ear.
[0140] Microphone 170C, also known as a "microphone" or "microphone," is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can put their mouth close to the microphone and speak, inputting the sound signal into the microphone. Terminal 31 can be provided with at least one microphone. In some embodiments, Terminal 31 can be provided with two microphones, which, in addition to collecting sound signals, can also implement noise reduction. In some embodiments, Terminal 31 can also be provided with three, four, or more microphones to collect sound signals, reduce noise, identify the sound source, implement directional recording functions, and the like.
[0141] The headphone jack 170D is used to connect a wired headphone. The headphone jack can be a USB interface, or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0142] The pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor can be set on the display screen. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can be composed of at least two parallel plates made of conductive material. When a force acts on the pressure sensor, the capacitance between the electrodes changes. The terminal 31 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display screen, the terminal 31 detects the intensity of the touch operation based on the pressure sensor. The terminal 31 can also calculate the position of the touch based on the detection signal of the pressure sensor. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than a first pressure threshold acts on a short message application icon, an instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on a short message application icon, an instruction to create a new short message is executed.
[0143] The gyroscope sensor 180B can be used to determine the motion posture of the terminal 31. In some embodiments, the angular velocity of the terminal 31 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor. The gyroscope sensor can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor detects the angle of the terminal 31 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the terminal 31 through reverse movement to achieve anti-shake. The gyroscope sensor can also be used for navigation and somatosensory game scenes.
[0144] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the terminal 31 calculates the altitude based on the air pressure value measured by the air pressure sensor to assist in positioning and navigation.
[0145] The magnetic sensor 180D includes a Hall effect sensor. Terminal 31 can utilize the magnetic sensor to detect the opening and closing of a flip case. In some embodiments, when Terminal 31 is a flip phone, Terminal 31 can utilize the magnetic sensor to detect the opening and closing of the flip cover. Furthermore, based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.
[0146] Accelerometer 180E can detect the magnitude of the acceleration of terminal 31 in all directions (generally three axes). When terminal 31 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify terminal posture, enabling applications such as switching between landscape and portrait modes and pedometers.
[0147] The distance sensor 180F is used to measure distance. The terminal 31 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the terminal 31 can use the distance sensor to measure distance to achieve fast focus.
[0148] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. Infrared light is emitted outward through the light emitting diode. A photodiode is used to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 31. When insufficient reflected light is detected, it can be determined that there is no object near the terminal 31. The terminal 31 can use the proximity light sensor to detect when the user holds the terminal 31 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.
[0149] Ambient light sensor 180L is used to sense ambient light brightness. Terminal 31 can adaptively adjust the display brightness based on the perceived ambient light. The ambient light sensor can also be used to automatically adjust the white balance when taking photos. The ambient light sensor can also work in conjunction with the proximity sensor to detect whether Terminal 31 is in a pocket to prevent accidental touches.
[0150] Fingerprint sensor 180H is used to collect fingerprints. Terminal 31 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc. In lock screen applications, the fingerprint sensor can be used to authenticate the user during the unlocking process. In this embodiment of the application, the mobile phone can use the fingerprint sensor to complete the unlocking function on the lock screen interface.
[0151] Temperature sensor 180J is used to detect temperature. In some embodiments, terminal 31 uses the temperature detected by the temperature sensor to implement a temperature management strategy. For example, when the temperature reported by the temperature sensor exceeds a threshold, terminal 31 may reduce the performance of a processor located near the temperature sensor to reduce power consumption and implement thermal protection.
[0152] The touch sensor 180K, also known as a "touch panel," can be placed on a display screen to detect touches applied to or near it. The detected touches can be communicated to an application processor to determine the type of touch event and provide corresponding visual output through the display screen.
[0153] The touch sensor can identify the user's operations on the phone, such as the user opening the lock screen and other applications, the user selecting and clicking a certain control, etc.
[0154] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor can acquire vibration signals from the vibrating bones of the human body. The bone conduction sensor can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor can also be set in headphones. The audio module 170 can parse the vibration signals of the vibrating bones of the human body acquired by the bone conduction sensor into voice signals to implement voice functions. The application processor can parse heart rate information based on the blood pressure signals acquired by the bone conduction sensor to implement heart rate detection functions.
[0155] The buttons 190 include a power button, a volume button, etc. The buttons can be mechanical buttons or touch buttons. The terminal 31 receives the button input and generates key signal input related to the user settings and function control of the terminal 31.
[0156] Motor 191 can generate vibration prompts. The motor can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. Touch operations acting on different areas of the display screen can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0157] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.
[0158] The SIM card interface 195 is used to connect a subscriber identity module (SIM). The SIM card can be connected to or disconnected from the terminal 31 by inserting it into or removing it from the SIM card interface. The terminal 31 can support one or N SIM card interfaces, where N is a positive integer greater than one. The SIM card interface can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface at the same time. The types of the multiple cards can be the same or different. The SIM card interface can also be compatible with different types of SIM cards. The SIM card interface can also be compatible with external memory cards. The terminal 31 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the terminal 31 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 31 and cannot be separated from the terminal 31.
[0159] The software system of the terminal 31 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the Android system with a layered architecture is used as an example to illustrate the software structure of the terminal 31.
[0160] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0161] In this embodiment, please refer to Figure 5, which is a schematic diagram of the structure of a server provided in this embodiment of the application. As shown in Figure 5, the server 32 may include a processing module 501, a storage module 502, and a communication module 503.
[0162] The processing module 301 is the control center of the system. For example, the processing module 501 can be any one or a combination of CPU, GPU, field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC).
[0163] The storage module 502 is used to store data, such as user usage data and data corresponding to multimedia content. The storage module 502 can also store computer instructions. The processing module 501 reads and executes the computer instructions stored in the storage module 502 to implement the server function.
[0164] The communication module 503 is used for communication between the server and the terminal.
[0165] The method of this embodiment will be briefly described below with reference to Figures 3 and 6. In Figure 3, the solid-line boxes represent modules in the server or terminal or actions performed by the modules, and the dotted-line boxes represent data stored in the terminal or server.
[0166] In the system of the embodiment of the present application, as shown in FIG3 , in step 702 marked in FIG3 , the terminal 31 is responsible for sending the usage data generated by the user browsing the multimedia application to the server 32, and is responsible for assisting the user in selecting the display style. The specific process can be found in the description of step 702 of FIG7 below. Corresponding to FIG6 , a window for the user to turn on the “customized lock screen” function can be provided on the terminal. The window can also display multiple preset display styles, such as “simple” style, “two-dimensional” style, “fresh” style and “national style” style. The preset display style can also be other display styles, such as “cartoon” style, etc., which are not specifically limited in this embodiment. The user can select a display style from multiple preset display styles, and the selected display style is the style of at least one of the generated target image and target text.
[0167] In FIG3 , as shown in step 703 marked in FIG3 , the server 32 may take the user-selected display style and user data as input. As shown in step 704 , the target image and target text may be generated based on the above input data through an artificial intelligence model. The target image and target text together constitute the lock screen image and text. The user data here may include the user's usage data and the data of the multimedia content browsed by the user. The terminal may only send the user's usage data to the server. The artificial intelligence model may include an image-text-speech three-tower model for feature extraction, an image generation model for generating images, and a text generation model for generating text. The user data here includes the user's usage data of multimedia content, as well as the metadata and / or media data corresponding to the multimedia content. The specific implementation process of the above steps 703 and 704 can be found in the description of FIG7 below. Although two multimedia applications are shown in FIG3 , it should be noted that multimedia applications may also include game applications, reading applications, etc. Corresponding to FIG6 , the server may generate a lock screen image and text with the user-selected display style based on the user data and the user-selected display style.
[0168] As indicated in step 705 in Figure 3 , server 32 in Figure 3 can send a generated lock screen image to terminal 31. For the specific implementation process, see the description of step 705 in Figure 7 below. After receiving the lock screen image from server 32, terminal 31 can display the received lock screen image on the lock screen interface. Corresponding to Figure 6 , the terminal can display the lock screen image generated based on the display style selected by the user. Furthermore, the user can also operate on the lock screen image, which can redirect the terminal to the multimedia content interface corresponding to the lock screen image in the multimedia application. For example, if the lock screen image is generated based on the movie "Challenger" that the user has viewed and the selected "minimalist" display style, in Figure 6 , the terminal can display a "minimalist" lock screen image related to "Challenger" as the lock screen wallpaper. For example, if "Challenger" is a movie about users pushing their limits and includes a video frame of the user jumping over a cliff, the minimalist-style image shown in Figure 6 can be generated. Furthermore, the user can also operate on the lock screen interface to trigger the terminal to directly redirect from the lock screen to the details interface of "Challenger" in the video application. It should be noted that "Challenger" here is just an example given for the convenience of description and does not correspond to the actual movie.
[0169] It should be noted that the embodiments shown in Figures 6 and 3 do not limit the present application. For example, in other embodiments of the present application, the terminal 31 may not assist the user in selecting a display style. In this case, the server 32 may only use the user data as input and generate lock screen images and texts through an artificial intelligence model. For another example, the user may not manually turn on the "customized lock screen" function, and the function may be turned on by default; the user may not select a display style, and the display style may be a default style, or the display style may be a display style that the user may tend to use based on the analysis of the multimedia content that the user has browsed historically. The server 32 may also not generate target images and target texts through an artificial intelligence model.
[0170] Next, the embodiment of the present application will be described in detail with reference to FIG7 . In the scenario where the method of the present application is implemented through interaction between a terminal and a server, among the steps shown in FIG7 , steps 701 and 702 shown by the dotted lines are optional, that is, steps 701 and 702 may not be performed. FIG7 is a flow chart of an embodiment of the present application, including:
[0171] Step 701: The terminal displays multiple candidate display styles and determines a target display style.
[0172] In this embodiment of the present application, the target display style is the display style selected by the user from multiple displayed alternative display styles. The display style is also the style of the displayed image and / or text. For example, the display style can be a "minimalist" style, a "two-dimensional" style, etc. The target display style is also the style of the target image and / or target text to be generated in subsequent steps.
[0173] As a possible implementation, step 701 may specifically include: the terminal displays multiple candidate display styles; and receiving a user's selection operation of a target display style from the multiple candidate display styles. The terminal may determine the display style selected by the user as the target display style.
[0174] The selection operation may specifically be a user triggering operation on a control corresponding to a target display style in the interface, for example, the user may specifically click on a control corresponding to a target display style in a display screen.
[0175] As an example, take the target display style as the style of the target image to be generated and the terminal as a mobile phone. The user can turn on the customized lock screen function in the corresponding settings interface. In response to the user's operation of turning on the customized lock screen function, the mobile phone displays an interface as shown in Figure 8, which includes multiple alternative display styles. Figure 8 shows the situation of 4 alternative display styles, and each alternative display style corresponds to a control. Taking the target display style as the "simple" style as an example, the above selection operation can be a user's click operation on the control 801. In response to the above click operation on the control 801, the mobile phone determines that the target display style is the simple style.
[0176] This description uses the target display style as an example of the target image to be generated. It is understood that the target display style can be selected for both the target image and the target text. Alternatively, after determining the target image display style, the above process can be repeated for the target text to determine the style corresponding to the target text. In other words, the target image and target file to be generated can have the same or different styles.
[0177] Step 702: The terminal sends the target display style to the server.
[0178] After the terminal sends the target display style to the server, the server can receive the target display style from the terminal. In this embodiment, the target image and target text are generated by the server. Therefore, after the terminal determines the target display style, it can send the target display style to the server so that the server can determine whether the user has enabled the customized lock screen function. Based on the received information, the server can analyze and determine the target display style selected by the user, and then generate a target image and / or target text with the target display style.
[0179] In addition to executing step 702 to send the target display style to the server, the terminal may also send the user's usage data to the server. That is, when a user uses a multimedia application on the terminal, the terminal often needs to send the user's usage data to the server. For example, for a video application, if a user watches a video using the video application, the video application may send usage data to the server based on the user's usage history to indicate that the user watched the video during that time period. That is, the usage data may include the time the user watched the video and the identifier of the video being watched. Furthermore, the usage data may also include the user's viewing progress, the content and time of the user's comments and barrages, and so on.
[0180] Step 701 and step 702 are a process in which a user turns on a customized lock screen by selecting a display style. By customizing the display style, a personalized customized lock screen can be generated according to user needs. It is easy to understand that the target display style can increase the personalization and diversification of the lock screen, and steps 701 and 702 can also be defaulted. For example, in some embodiments, the target display style can be determined, and the customized lock screen can be generated only according to the first data and the second data below; or in other embodiments, the target image and target text are generated according to the first data, the second data and the default display style. For example, after the user turns on the customized lock screen function, if the display style is not selected, the display style corresponding to the user is determined to be the default display style; for example, the terminal can turn on the customized lock screen function by default, and use the default display style to generate the target image and target text by default.
[0181] In addition to the methods of steps 701 and 702, the target display style can also be determined by other means. For example, the target display style can be the display style that best matches the user among the preset display styles determined by the terminal based on the multimedia content viewed by the user, and the target display style can be sent to the server in step 702. Alternatively, the target display style can be the display style that best matches the user among the preset display styles determined by the server based on the multimedia content viewed by the user. In other words, the target display style is the display style that best matches the user among the preset display styles determined based on the multimedia content viewed by the user, and the subject of determining the target display style can be the terminal or the server.
[0182] A specific determination method may be to input the names of multimedia data viewed by the user into a classification model and use the classification model to determine the display style that best matches the user. Alternatively, the target display style may be determined by clustering the feature vectors of videos viewed by the user and determining the display style that best matches the feature vector at the cluster center as the target display style. The above two specific implementation methods are merely examples and do not limit the present application.
[0183] After executing steps 701 and 702 or the other steps for determining the target display style, the server can determine that the customized lock screen function is enabled on the terminal and needs to provide the target image and target text to the terminal. Then, steps 703-705 can be executed to generate and send the target image and target text to the terminal.
[0184] It should be noted that in some embodiments, steps 701-702 can be performed only once, and the subsequent steps 703-706 can be performed multiple times, that is, the user can only enable the customized lock screen function once, and the server can generate multiple target images and / or target texts that conform to the target display style, and each time the user enables the lock screen application, a different target image and target text can be displayed, thereby increasing the diversity of the lock screen interface. In other embodiments, steps 701-706 can be performed sequentially, with each step only performed once. In other embodiments, steps 701-702 can be performed multiple times, and after steps 701-702 are performed multiple times, steps 703-706 are continued to be performed. In this way, by performing steps 701-702 multiple times, the user can set a variety of different target display styles. When steps 703-706 are subsequently performed multiple times, target images and target texts with multiple display styles can be generated, thereby increasing the diversity of the lock screen interface.
[0185] Step 703: The server obtains the first data and the second data of the terminal.
[0186] Step 704: The server generates a target image and target text according to the first data and the second data.
[0187] Among them, the target image is at least related to the first data, and the target text is at least related to the second data; the first data is used to indicate the multimedia content browsed by the user corresponding to the terminal using the multimedia application, and the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, and the media data includes at least one of the text, audio, video and image of the multimedia content; the second data is used to indicate the usage data when the user uses the multimedia application to browse the multimedia content. The specific meaning of the usage data has been explained above and will not be repeated here.
[0188] That is, the target image generated is related to the multimedia content browsed by the user, and the target text generated is related to the user's usage data. The specific description of the two can be found above and will not be repeated here. In some embodiments, the generated target image and target text can also be related to the target display style of steps 701-702. Here, taking the lock screen scenario as an example, the target text and the target image with the target display style are generated by artificial intelligence generated content (AI generated content, AIGC) to explain steps 703 and 704 in detail.
[0189] Specifically, step 704 may include: obtaining first data and second data; using the first data, second data and target display style as input, and generating a target image and target text through a target model; the target model has the function of generating corresponding images and texts based on the input multimedia content and display style browsed by the user.
[0190] As described above, the terminal sends the user's usage data to the server. After receiving the user's usage data, the server may store the usage data. Furthermore, the server generally stores multimedia data of multimedia applications, such as a multimedia database for multimedia applications. Acquiring the first data and the second data here may involve acquiring the second data from the user's usage data stored on the server, determining the identifier of the multimedia content viewed by the user, and acquiring the first data that matches the determined identifier of the multimedia content from the multimedia database for the multimedia application stored on the server.
[0191] Specifically, the target model may be a model that inputs the first data, the second data, and the target display style, and outputs a target image including the target text. Specifically, the output image includes the text, and the text content does not obscure the main portion of the image. For example, as shown in FIG9 , the target image including the target text is 901, the target text is 902, the main portion of the image is a person, and the target text does not obscure the person.
[0192] Furthermore, the font and color of the text can also be consistent with the display style of the target image or the overall display effect of the image. For example, if the main part of the image is red and occupies the largest area, the text color can also be red. For another example, if the target image is in the traditional Chinese style, the text can use fonts such as cursive script or regular script.
[0193] The target model may further include two models, namely, an image generation model and a text generation model. Then, step 704 specifically includes: taking the first data and the target display style as input, generating a target image through the image generation model; and taking the second data as input, generating a target text through the text generation model.
[0194] The image generation model here can be a generative adversarial network (GAN) or a diffusion model (DM). The text generation model can be a generative pre-trained transformer (GPT) or other similar models.
[0195] In the embodiments of the present application, a pre-trained image generation model and text generation model can be used to directly generate a target image and target text. Furthermore, in the embodiments of the present application, sample data labeled with display styles can be obtained for the pre-trained image generation model, specifically for the scenario of the embodiments of the present application, and the image generation model can be further trained using the sample data. This allows the image generation model to better learn the characteristics of different display styles, thereby enabling the generation of a target image that better matches the target display style in step 704, based on the target display style.
[0196] Furthermore, before generating the target image and target text using the target model, the method may further include extracting a feature vector corresponding to the first data (including metadata and media data) of the multimedia content and using the feature vector as input to the image generation model or to the image generation model and the text generation model. Two descriptive words may also be generated to indicate the content requirements of the model output. This allows the model to better generate target images and target text that meet the requirements based on the input.
[0197] In other words, when the first data includes at least metadata, before step 704, it also includes: obtaining a feature vector of the first data; generating a first description word based on the metadata and the target display style, and generating a second description word based on the second data, the first description word is used to describe the characteristics of the image to be generated, and the second description word is used to describe the characteristics of the text to be generated.
[0198] In the above case, step 704 includes: taking the feature vector of the first data and the first description word as input, generating a target image through the image generation model; taking the second description word as input, outputting a target text through the text generation model.
[0199] The above generation process is detailed in Figure 10. Next, we will describe this generation process in detail, using Figure 10 as an example, using a video application as the multimedia application and the movie "Challenger" as the multimedia content being viewed by the user. In Figure 10, the solid-line boxes represent the specific model, and the dashed-line boxes represent the input or output data.
[0200] For "Challenger", metadata may include basic information about the movie, such as the title, director, screenwriter, etc.; media data is the specific content of the movie, such as the movie poster, trailer, plot introduction, complete video, complete audio, and complete subtitles.
[0201] First, a feature vector of the first data can be obtained. The obtained feature vector can include any one or more of an image feature vector, an audio feature vector, and a text feature vector. Feature vectors corresponding to different types of data can be extracted using different models. Next, the process of obtaining feature vectors will be specifically described using the image feature vector, audio feature vector, and text feature vector of "Challenger" as an example, combined with the first, second, and third steps below.
[0202] First, the video and / or image data of "Challenger" is input into the computer vision model to extract the image feature vector of the movie.
[0203] Specifically, the computer vision model may be a contrastive language-image pre-training (CLIP) model. The video of "Challenger" may be a complete video of the movie and / or a promotional video of the movie. The image of "Challenger" may include a poster image and / or a screenshot of the movie video. The screenshot of the movie video may be an official screenshot of the movie video released for promotion, or a barrage or a video frame with a large number of views. In the case where the user's usage data includes data on the user's barrage, the screenshot of the movie video may also be a screenshot corresponding to the video progress of the user's barrage, etc.
[0204] Second, the complete audio of "Challenger" is input into the audio signal processing model to obtain the audio feature vector of the movie.
[0205] Third, the text data of "Challenger" is input into the natural language processing model to obtain the text feature vector of the movie.
[0206] The text data here may include the title, director, screenwriter, plot description, and complete subtitles of "The Challenger." In one possible implementation, the text input into the natural language processing model may include text from the media data of "The Challenger." In another possible implementation, the text input into the natural language processing model may also include metadata of "The Challenger." In yet another possible implementation, the text input into the natural language processing model may include at least one of metadata of "The Challenger" and text from the media data.
[0207] Among them, the computer vision model in the first, the audio signal processing model in the second, and the natural language processing model in the third together constitute the image-text-speech three-tower model mentioned in Figure 3.
[0208] Any one or more of the three steps above may be omitted. Specifically, one or more of the three steps above may be selectively performed to complete feature vector extraction for different types of multimedia content and based on the characteristics of the multimedia content. For example, in this embodiment, the target image and target text are generated for the multimedia content "Challenger." For video application scenarios, where image and video data are abundant, the server can generate the target image based solely on image frames in the video. Therefore, in this embodiment, the second and third steps may also be omitted.
[0209] Fourth, a first description word is generated according to the first data and the target display style, and a second description word is generated according to the second data.
[0210] A specific method for generating descriptive words may be through a natural language processing model. For example, in this embodiment, metadata / media data of "Challenger", target display style, and user usage data of "Challenger" may be input into a natural language processing model to output a first descriptive word and a second descriptive word.
[0211] Specifically, the first descriptor controls target image generation and is generated based on the target display style and metadata / media data for "Challenger." The following example illustrates generating the first descriptor based on the target display style and the title of "Challenger." For a "minimalist" target display style, the first descriptor might be "Generate a minimalist image of "Challenger." The target display style and the title describe the target image, while the first descriptor describes the characteristics of the target image to be generated. Based on this first descriptor, the image generation model can generate an image related to "Challenger" and in the target display style.
[0212] In addition to generating the first descriptive word based on the target display style and the title of the multimedia content, descriptive words can also be generated based on other data. For example, descriptive words can be generated based on the target display style, the title of the multimedia content, and the introduction of the multimedia content. For example, "Generate a minimalist image of a person pushing their limits, 'Challenger'." The above examples are provided for ease of understanding only and do not limit the embodiments of the present application.
[0213] The second descriptor is used to control the generation of the target text. The second descriptor is generated based at least on the user's usage data for "Challenger." This user usage data may include the identifiers of the content the user browses and a description of their specific usage behavior. For example, the user's usage data may include the time the user watches "Challenger" and the progress of watching "Challenger." For example, the second descriptor may specifically be "Generate a lock screen description to remind the user that they watched the movie pictured on XX month XX day."
[0214] In addition to describing the specific usage behavior, the second descriptor can also include the title of the multimedia content the user browsed. In other words, the user's specific usage behavior and the title of the multimedia content the user browsed can be determined based on the user's usage data, and the second descriptor can be generated based on these two factors. For example, the second descriptor could be "Generate a lock screen description to remind the user that they watched 'The Challenger' on XX month XX day."
[0215] In addition, for the natural language processing model, during the training phase, some samples including descriptive words can be obtained for the natural language processing model that has completed pre-training, and the natural language processing model can be further trained for the usage scenarios in the embodiments of this application, so that the natural language processing model can generate descriptive words that meet the requirements of the embodiments of this application.
[0216] In the above example, the first and second descriptive words are generated by a natural language processing model, which can increase the diversity of the generated descriptive words. It is easy to understand that the first and second descriptive words can also be generated not according to the artificial intelligence model, for example, they can be generated according to a template. For example, the template of the first descriptive word can be "generate an image of Y in style X", where X can be the label of the target display style, such as "simple", "second dimension" or "national style", etc., and Y is the title of the video. The second descriptive word is similar to the first descriptive word and will not be repeated here.
[0217] Fifth, obtain the unified eigenvector.
[0218] In some embodiments, at least two of the first, second, and third steps may be performed. In this case, at least two feature vectors among the image feature vector, audio feature vector, and text feature vector can be output. Then, the two or three feature vectors output by the above process are input into the information fusion module to align and fuse the features of different modalities and generate a unified feature vector.
[0219] Sixth, the feature vector and the first description word are input into the image generation model to obtain the target image; the second description word is input into the text generation model to obtain the target text.
[0220] Among them, the feature vector of the input image generation model can be a unified feature vector obtained in the fifth step. When only any one of the first, second and third steps is performed, the above feature vector can be any one of the image feature vector, audio feature vector and text feature vector.
[0221] Still taking the target display style as "simple" and the multimedia content as "Challenger" as an example, the above steps can output a minimalist-style image of "Challenger" as shown in (a) of Figure 11. When the usage data is the time the user watched "Challenger", the following target text can be generated: "You watched "Challenger" on XX month XX day, come and reminisce." When the usage data is the progress of the user watching "Challenger", the following target text can be generated: "You have watched 50% of the content of "Challenger" on the video application, come and continue watching." The above target text can enhance the interpretability of the image shown in (a) of Figure 11, so that users can understand the image content and the purpose of pushing the image, thereby improving the user experience. At the same time, the target text can encourage users to use the video application more, increasing traffic for the video application.
[0222] The above examples are explained using video applications as an example. It is easy to understand that other multimedia application processing methods are similar to the above process, except that the media data input in the first, second and third steps are different.
[0223] For example, in a music application, in the first step, you can input the song's MV and / or the song's album cover, in the second step, you can input the song's audio content, and in the third step, you can input the song's lyrics. In a reading application, in the first step, you can input the book's cover and / or illustrations, and in the third step, you can input the book's text and / or a brief introduction. In a game application, in the first step, you can input game promotional screenshots and / or promotional videos, in the second step, you can input the game's promotional video's audio and / or the game's soundtrack, and in the third step, you can input the game's introduction. The above examples are merely examples and do not limit this application.
[0224] Furthermore, in the aforementioned scenarios, media data can include images, videos, text, and audio. However, in other scenarios, media data may only include one or two of these. For example, in a reading app, a book may only contain text, so the media data may only include text data. Alternatively, in a music app, a song may only contain audio and lyrics, so the media data may only include audio and text data.
[0225] In this case, the process of extracting some feature vectors can be omitted, and the vectors input to the image generation model also need to be adjusted accordingly. For example, in the scenario of a reading application, when the medium data only includes text data (i.e., the text of the book being read), the first, second, and fifth steps above can be omitted, and the image generation model inputs the text feature vector. For example, in the scenario of a music application, when the medium data includes text and audio, the first step can be omitted, and the image generation model inputs the fusion vector of the text feature vector and the audio feature vector.
[0226] In addition, even if the multimedia application includes four types of data, namely, images, videos, texts and audios, in the first, second and third steps, feature vectors can be extracted based on only any one, two or three of the four types of data, and the target image can be further generated based on the extracted feature vectors.
[0227] The above example is explained by taking the generation of a target image with a target display style as an example. If you need to generate a target text with a target display style, the generation method is similar to the method mentioned above. For example, you can first perform any one or more of the first, second and third steps. In the fourth step, a second descriptive word is generated based on the second data and the target display style. The target display style used to generate the second descriptive word here can be the same as the target display style used to generate the first descriptive word, or it can be different from the target display style used to generate the first descriptive word. Then continue to perform the subsequent fifth step. In the sixth step, the feature vector and the second descriptive word can be input into the text generation model to obtain the target text with the target display style.
[0228] Step 705: The server sends the target image and target text to the terminal.
[0229] Correspondingly, the terminal can receive the target image and target text from the server.
[0230] Regarding the execution timing of step 703, step 704, and step 705, step 703 and step 704 can be executed periodically, that is, the target image and target text are generated periodically using the user's historical usage data and the data of the multimedia content corresponding to the historical usage data. In other words, the target image and target text are generated based on the multimedia content that the user has browsed in the past. In addition, an image library is maintained for each user, and the image library stores at least two target images generated by the above method and their corresponding target texts. Step 705 can also be executed periodically, that is, the server periodically pushes the target image and target text to the terminal through step 705. Specifically, the server can periodically randomly select an image from the above image library and send it to the terminal. Among them, the execution cycles of step 704 and step 705 can be different or the same.
[0231] Furthermore, steps 703 and 704 may also be executed when new usage data is detected. That is, when a user uses a multimedia application, a target image and target text corresponding to the multimedia content the user is viewing using the multimedia application may be generated in real time. For example, if the user is watching "The Challenger," a target image for "The Challenger" (as shown in FIG9 ) may be generated, along with target text to remind the user to continue watching "The Challenger."
[0232] Step 706: The terminal receives a first operation of the user to open a lock screen application; in response to the first operation, the terminal displays a target image and a target text.
[0233] That is, in the lock screen scene, when the user opens the lock screen interface, the target image and target text are displayed. The description of the first operation is detailed in the previous text and will not be repeated here.
[0234] In some embodiments, step 705 may be executed periodically. After the server sends the target image and target text to the terminal in step 705, the terminal may first cache the received target image and target text on the terminal. When the user opens the lock screen, desktop, off screen, or system application, the target image and target text are retrieved from the cached content for display. If there are multiple cached target images and target texts, any pair of target images and target texts may be selected from them for display.
[0235] In other embodiments, step 705 may also be executed after receiving a trigger from the terminal. For example, step 706 may be executed first, where the terminal receives the first operation of the user to open the lock screen application. In response to the first operation, step 705 is triggered, and the target image and target text are further displayed in step 706.
[0236] It should be noted that the above two examples are only two examples for illustrating the execution order of step 705 and step 706 in this application, and do not represent limitations on the embodiments of this application.
[0237] The method for displaying the target image and target text may be to display the target image and target text simultaneously in response to a first operation. Alternatively, the method may be to display the target image on the lock screen interface and then display the target text upon receiving a third operation from the user on the lock screen interface.
[0238] The third operation here can be a triggering operation on a control in the lock screen interface, such as a click or swipe up on control 1101 in (a) of Figure 11, or the third operation can be a swipe up operation on the lock screen interface that is different from the unlock screen. In this way, the terminal displays the target text after receiving the third operation, which can increase the interpretability of the target image through the target text without obscuring the main body of the target image, and can also enhance the interaction between the terminal and the user, thereby improving the user experience.
[0239] Taking the third operation as an upward swipe operation on control 1101 as an example, the above process is described in detail. As shown in Figure 11, after the user opens the lock screen interface of the lock screen application through the first operation, the terminal can display the interface including the target image shown in (a) in Figure 11 in response. The user swipes upward on control 1101 in the lock screen interface shown in (a) in Figure 11. After the terminal receives the operation, in response, the terminal can display an animation of control 1101 rising from the bottom of the lock screen interface, and finally display the interface shown in (b) in Figure 11. Among them, in (b) in Figure 11, the text included in control 1101 is the target text.
[0240] As shown in (b) of FIG11 , in the interface displaying the target image and target text, in addition to control 1101 , the lock screen interface may also have controls 1102 and 1103 , and the functions of these two controls will be described in detail below.
[0241] Control 1102 is also called jump control, also referred to as second control in this application. In the embodiment of this application, by triggering the jump control by the user, the user can jump from the lock screen interface to the interface of the multimedia content corresponding to the image on the lock screen interface.
[0242] In addition, if the user sets an unlock password or other unlock verification, the user identity needs to be verified before jumping to the multimedia content interface. The method of verifying the user identity can refer to the method of related technology and will not be repeated here.
[0243] In other words, the method of the embodiment of the present application also includes: displaying a second control; receiving a fourth operation of the user on the second control; and displaying multimedia content corresponding to the target image in the second application in response to the fourth operation.
[0244] The fourth operation is a user triggering operation on the control 1102 shown in (b) of Figure 11. For example, the fourth operation may be an operation in which the user clicks the control 1102. The multimedia content is displayed in the second application, that is, jumping to an interface for browsing the multimedia content, such as the interface shown in Figure 12.
[0245] Next, the above process will be described using the fourth operation, which is the user clicking the control 1102. When the lock screen displays the image shown in FIG11(b), the user clicks the control 1102, and the terminal jumps to the "Challenger" playback interface shown in FIG12.
[0246] The control 1103 is also called the delete control, also referred to as the first control in this application. The delete control is used to stop displaying the current target image and target text after receiving a trigger.
[0247] Specifically, the method of the embodiment of the present application further includes: displaying the first control; receiving a second operation of the user on the first control; and sending the first information to the server in response to the second operation.
[0248] The first information is used to instruct the server to delete the target image and target text in the image library corresponding to the user; the image library stores multiple images and texts generated based on the usage data of the user browsing multimedia content and the multimedia content browsed by the user.
[0249] The second operation may be a user triggering control 1103, such as a user clicking on control 1103. For example, after the terminal displays the content shown in FIG11(b), the user may click on control 1103. Upon receiving the user's click on control 1103, the terminal may generate first information and send the first information to the server. The server receives the first information from the terminal and, in response to the first information, deletes the corresponding image and text from the image library, which is the image library consisting of the multiple target images and target text generated in step 705 above.
[0250] In the above embodiment, the lock screen application is used as an example for description. In addition, the first application can also be a desktop application, an off-screen application, or a system application. In the case where the first application is within another application, the overall implementation method is similar to steps 701-706, except that the specific implementation method of step 706 is different.
[0251] Specifically, displaying the target image and target text in step 706 includes: if the first application is a desktop application, at least the target image and target text are displayed as desktop wallpaper. Alternatively, if the first application is an off-screen application, the target image and target text are displayed as a widget on the off-screen interface. For example, widget 201 shown in FIG. 2 can be replaced with the target image and target text.
[0252] In addition, when the first application is any application of a desktop application, a lock screen application or an off-screen application, a theme of the terminal can be generated according to the target image and target text, and the target text and target image can be displayed through the theme in the three applications. Among them, a theme is a personalized design of a terminal skin interface, and a theme can be designed and developed by a company or an individual developer. The content in the theme that can be personalized and replaced by the target image and target text may specifically include wallpaper, icons and backgrounds in the interface covered by the theme. The interface covered by the theme may include notification bar, text messages, dialing, contacts and settings. Among them, the wallpaper or background can be replaced by the target image and target text, such as displaying the image shown in Figure 9 in the background of the wallpaper or interface. The icon can be replaced by the target image or the theme content in the target image. For example, the application icon on the desktop can be the target image 901 as shown in Figure 9, and the application icon on the desktop can also be the main content of the target image, such as the person shown in Figure 9.
[0253] Alternatively, the first application is a system application, and the target image and target text are used as the startup image and displayed on the startup interface of the first application. Applications generally need to be loaded during the startup process, and the application logo or advertising image will be displayed on the startup interface during the loading process. In this application, the above-mentioned target image and target text can also be displayed on the startup interface of the system application. In addition, if there is cooperation between the developer of the third-party application and the terminal manufacturer, or the terminal operating system has the authority to control the startup interface of the third-party application, the above-mentioned target image and target text can also be displayed on the third-party application.
[0254] In the above embodiment, the implementation method of the embodiment of the present application is described in detail using the interaction between the terminal and the server as an example. As mentioned above, the embodiment of the present application can also be executed by the terminal alone. When executed by the terminal alone, the specific implementation method is similar to steps 701-706 above. Unlike steps 701-706, when the terminal alone executes the above method, steps 702 and 705 may not be executed. Steps 703 and 704 are executed by the terminal.
[0255] In addition, when the terminal is executed alone, the target display style may be determined in step 701 or may be a display style that best matches the user among preset display styles determined based on multimedia content browsed by the user.
[0256] Furthermore, if the aforementioned delete control exists, upon receiving the second operation, the terminal may delete the target image and target text from the image library corresponding to the user in response to the second operation; the image library stores multiple images and text generated based on the user's usage data when browsing multimedia content and the multimedia content browsed by the user. That is, the terminal stores the aforementioned image library, and in response to the second operation, the terminal deletes the corresponding image and text from the image library.
[0257] Some other embodiments of the present application also provide a device having the function of implementing the behavior of the electronic device (such as a terminal or server) in the above-mentioned embodiment. The function can be implemented by hardware, or it can be implemented by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions, for example, a processing module, a sending module, and a display module. As an example, the processing module can execute the above-mentioned step 701 or step 703; the sending module can execute the above-mentioned step 702 or step 705; and the display module can execute the above-mentioned step 706.
[0258] The present application also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above method.
[0259] The present application also provides a server, comprising a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to execute the computer programs to implement the above method.
[0260] The present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to implement the method.
[0261] The present application also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.
[0262] The present application also provides a chip, which includes a memory and a processor, the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory to execute the above method.
[0263] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An image display method, characterized in that: Applied to a terminal, the method comprises: Receiving a first operation of a user opening a first application; In response to the first operation, displaying a target image and a target text; Among them, the target image is at least related to first data, the first data is used to indicate the multimedia content browsed by the user using the second application, the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data includes at least one of the text, audio, video and image of the multimedia content; the target text is at least related to second data, the second data is used to indicate usage data of the user when browsing the multimedia content using the second application.
2. The method according to claim 1, characterized in that The target image is further related to a target display style, and / or the target text is further related to the target display style.
3. The method according to claim 2, characterized in that The target display style is a display style selected by the user from a plurality of displayed candidate display styles; or, the target display style is a display style that has the highest matching degree with the user among preset display styles determined based on multimedia content browsed by the user.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Display a second control; receiving a fourth operation of the user on the second control; In response to the fourth operation, the multimedia content corresponding to the target image is displayed in the second application.
5. The method according to any one of claims 2 to 4, characterized in that: The method further comprises: The target display style is sent to the server.
6. The method according to any one of claims 1 to 5, characterized in that: Before displaying the target image and the target text, the method further includes: The target image and the target text are received from the server.
7. The method according to claim 5 or 6, characterized in that: The method further comprises: Display a first control; receiving a second operation of the user on the first control; In response to the second operation, first information is sent to the server, wherein the first information is used to instruct the server to delete the target image and the target text in an image library corresponding to the user, wherein the image library stores multiple images and texts generated based on usage data of the user when browsing multimedia content and the multimedia content browsed by the user.
8. The method according to any one of claims 2 to 4, characterized in that: The method further comprises: Acquire the first data and the second data; The first data, the second data and the target display style are used as input, and the target image and the target text are generated through a target model, wherein the target model has the function of generating corresponding images and texts based on the multimedia content and display style browsed by the input user.
9. The method according to claim 8, characterized in that The target model includes an image generation model and a text generation model; the step of taking the first data, the second data and the target display style as input and generating the target image and the target text through the target model includes: Taking the first data and the target display style as input, generating the target image through the image generation model; The second data is taken as input and the target text is generated through the text generation model.
10. The method according to claim 9, characterized in that In the case where the first data at least includes the metadata, the method further includes: Obtaining a feature vector of the first data; Generate a first description word according to the metadata and the target display style, and generate a second description word according to the second data, wherein the first description word is used to describe the characteristics of the image to be generated, and the second description word is used to describe the characteristics of the text to be generated; The step of taking the first data and the target display style as input and generating the target image through the image generation model includes: The feature vector of the first data and the first description word are used as input, and the image generation model is used to generate the target image; The step of taking the second data as input and generating the target text through the text generation model includes: The second description word is used as input, and the target text is generated through the text generation model.
11. The method according to any one of claims 8 to 10, characterized in that: The method further comprises: Displaying a first control; receiving a second operation of the user on the first control; In response to the second operation, the target image and the target text are deleted in an image library corresponding to the user; the image library stores a plurality of images and texts generated based on usage data of the user when browsing multimedia content and the multimedia content browsed by the user.
12. The method according to any one of claims 1 to 11, characterized in that: The first application is a lock screen application; and the displaying of the target image and the target text includes: Displaying the target image in the lock screen interface; A third operation of the user in the lock screen interface is received, and the target text is displayed.
13. The method according to any one of claims 1 to 11, characterized in that: The displaying of the target image and the target text comprises: The first application is a desktop application, and the target image and the target text are used as desktop wallpaper and displayed on the desktop; or, The first application is a screen-off application, and the target image and the target text are displayed as widgets on a screen-off interface; or, The first application is a system application, and the target image and the target text are used as startup images and displayed on a startup interface of the first application.
14. An image generation method, characterized in that: Applied to a server, the method comprises: Acquire first data and second data of a terminal, wherein the first data is used to indicate multimedia content browsed by a user corresponding to the terminal using a second application, wherein the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data including at least one of text, audio, video and image of the multimedia content; and the second data is used to indicate usage data when the user uses the second application to browse the multimedia content; Generate a target image and a target text according to the first data and the second data; wherein the target image is at least related to the first data, and the target text is at least related to the second data; The target image and the target text are sent to the terminal.
15. The method according to claim 14, characterized in that The target image is further related to a target display style, and / or the target text is further related to the target display style.
16. The method according to claim 15, characterized in that The method further comprises: The target display style sent by the terminal is received, where the target display style is a display style selected by the user from a plurality of candidate display styles displayed by the terminal.
17. The method according to claim 15, characterized in that The target display style is a display style that has the highest matching degree with the user among preset display styles determined according to the multimedia content browsed by the user.
18. The method according to any one of claims 15 to 17, characterized in that: The step of generating the target image and the target text according to the first data and the second data includes: The first data, the second data and the target display style are used as input, and the target image and the target text are generated through a target model, wherein the target model has the function of generating corresponding images and texts based on the multimedia content and display style browsed by the input user.
19. The method according to claim 18, characterized in that The target model includes an image generation model and a text generation model; the step of taking the first data, the second data and the target display style as input and generating the target image and the target text through the target model includes: Taking the first data and the target display style as input, generating the target image through the image generation model; The second data is taken as input and the target text is generated through the text generation model.
20. The method according to claim 19, characterized in that In the case where the first data at least includes the metadata, the method further includes: Obtaining a feature vector of the first data; Generate a first description word according to the metadata and the target display style, and generate a second description word according to the second data, wherein the first description word is used to describe the characteristics of the image to be generated, and the second description word is used to describe the characteristics of the text to be generated; The step of taking the first data and the target display style as input and generating the target image through the image generation model includes: Taking the feature vector of the first data and the first description word as input, generating the target image through the image generation model; The step of taking the second data as input and generating the target text through the text generation model includes: The second description word is used as input, and the target text is generated through the text generation model.
21. The method according to any one of claims 14 to 20, characterized in that: The method further comprises: receiving first information from the terminal; In response to the first information, the target image and the target text are deleted in an image library corresponding to the user; the image library stores a plurality of images and texts generated based on the usage data of the user when browsing multimedia content and the multimedia content browsed by the user.
22. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the method according to any one of claims 1 to 13.
23. A server, characterized in that: include: a memory for storing instructions executed by one or more processors of the server; A processor, when the processor executes the instructions in the memory, can cause the server to execute the method described in any one of claims 14 to 21.
24. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the method according to any one of claims 1 to 21.
25. A chip, characterized in that: The chip includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory to execute the method according to any one of claims 1-21.
Citation Information
Patent Citations
Image display method, image generation method and electronic equipment
CN120010967A
Screen locking wallpaper recommendation method and device and electronic equipment
CN110688578A
Display method, mobile terminal and device with storage function
CN111176499A
Information reminding method and server
CN113449185A
Desktop display method and device, electronic equipment and medium
CN114237801A