Image display method, image generation method and electronic equipment

By displaying target images and text related to users browsing multimedia content on the terminal, the problems of insufficient diversity of lock screen wallpaper and low user relevance are solved, improving user experience and providing traffic display opportunities.

CN120010967APending Publication Date: 2025-05-16PETAL CLOUD TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311523564.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, terminal lock screen wallpapers are insufficient in diversity, and randomly selected wallpapers have low correlation with users, which affects the user experience.

Method used

By receiving user operations, display target images and target text related to the user browsing multimedia content, use the target model to generate images and text that meet user needs, and improve the diversity of wallpaper and user relevance.

Benefits of technology

It realizes the display of user-related content on the lock screen, desktop and screen-out interfaces, improves the user experience and provides traffic display opportunities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010967A_ABST
    Figure CN120010967A_ABST
Patent Text Reader

Abstract

The invention provides an image display method, an image generation method and electronic equipment, relates to the field of terminals, and solves the problems of relatively low association degree between wallpaper and a user and relatively low wallpaper diversity. According to the specific scheme, under the condition that a user opens a first application such as a desktop application, a screen locking application, a screen turning-off application or a system application, a target image related to multimedia content browsed by the user on other applications can be displayed, and a target text related to use data when the user browses the multimedia content on other applications can be displayed. The target image and the target text related to the multimedia content browsed by the user are displayed on the interface of the screen locking application, the desktop application, the screen turning-off application or the system application, and the relevance between the displayed content and the user is improved. And a plurality of target images and target texts can be generated according to various browsing histories of the user, so that the diversity of the displayed content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of terminals, and in particular to an image display method, an image generation method and an electronic device. Background Art

[0002] In the prior art, in terminals such as mobile phones and tablet computers, the wallpaper of the terminal lock screen interface is generally an image selected by the user, or an image randomly selected by the terminal from a preset wallpaper image library.

[0003] For the method of using a certain image selected by the user as the lock screen wallpaper, since the lock screen wallpaper will not change after the user selects it, the diversity of the lock screen wallpaper is insufficient. For the method of using an image randomly selected by the terminal from a preset wallpaper image library as the lock screen wallpaper, although this method can make the lock screen wallpaper constantly change to meet the diversity requirements, these images used as lock screen wallpapers are often not highly relevant to the user, affecting the user experience.

[0004] It can be seen that in the prior art methods, the terminal cannot provide wallpapers that are highly relevant to the user and meet the diversity requirements. Summary of the invention

[0005] The embodiments of the present application provide an image display method, an image generation method and an electronic device, which solve the problems in the prior art of low correlation between wallpapers and users and low diversity of wallpapers.

[0006] According to a first aspect of an embodiment of the present application, the present application provides an image display method, applied to a terminal, comprising: receiving a first operation of a user to open a first application; in response to the first operation, displaying a target image and a target text, wherein the target image is at least related to first data, the first data is used to indicate multimedia content browsed by the user using a second application, the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data includes at least one of the text, audio, video and image of the multimedia content; the target text is at least related to second data, the second data is used to indicate usage data when the user uses the second application to browse the multimedia content.

[0007] The technical solution provided in the first aspect above can display target images related to multimedia content browsed by the user on other applications, and display target text related to the usage data when the user browses multimedia content on other applications, when the user opens a first application such as a desktop application, a lock screen application, an off-screen application or a system application. In this way, target images and target texts related to multimedia content browsed by the user can be displayed on the interface of the lock screen application, the desktop application, the off-screen application or the system application. In other words, both the target image and the target text are related to the user. If the user has browsed a lot of multimedia content in history, then the corresponding target images and target texts can be generated respectively according to the multiple multimedia content browsed by the user, thereby improving the diversity of the lock screen, desktop, off-screen and other interfaces; such target images and target texts avoid the monotony of the images generated by the template, and can improve the user's experience; in addition, the lock screen wallpaper, desktop wallpaper and the widgets of the off-screen interface can actually be used as traffic entrances to divert traffic to the applications on the terminal, and such target images and target texts can also provide good traffic display opportunities.

[0008] In a possible implementation, the target image is also related to the target display style, and / or the target text is also related to the target display style.

[0009] In a possible implementation, the target display style is a display style selected by the user from a plurality of displayed candidate display styles; or, the target display style is a display style that best matches the user among preset display styles determined based on multimedia content browsed by the user.

[0010] In a possible implementation, the method further includes: displaying a second control; receiving a fourth operation of the user on the second control; and displaying multimedia content corresponding to the target image in the second application in response to the fourth operation.

[0011] That is, after displaying the target image and target text, a jump control can also be displayed. After the user performs the fourth operation on the jump control, the user can jump to the corresponding multimedia application and display the multimedia content corresponding to the target image. This plays a good role in attracting traffic and can increase user access to multimedia applications.

[0012] In a possible implementation manner, the method further includes: sending the target display style to the server.

[0013] In a possible implementation manner, before displaying the target image and the target text, the method further includes: receiving the target image and the target text from a server.

[0014] That is, the terminal can send the target display style to the server, and the server provides the target image and the target text after receiving the target display style.

[0015] In a possible implementation, the method further includes: displaying a first control; receiving a second operation of the user on the first control; and sending first information to the server in response to the second operation, wherein the first information is used to instruct the server to delete a target image and a target text in an image library corresponding to the user, wherein the image library stores a plurality of images and texts generated based on usage data when the user browses multimedia content and the multimedia content browsed by the user.

[0016] That is to say, when the target image and the target text are provided by the server, the server also stores an image library corresponding to the user, and the image library stores multiple target images and target texts generated by the above method. After the terminal displays the target image and the target text, it can also display a delete control. After the user performs a second operation on the delete control, the terminal can send a delete message to the server, and the delete message is used to instruct the server to delete the displayed target image and target text in the image library corresponding to the user. In this way, through the interaction between the user and the terminal, the user is no longer pushed content that the user is not interested in, thereby improving the user experience.

[0017] In a possible implementation, the method further includes: obtaining first data and second data; using the first data, the second data and the target display style as input, and generating a target image and a target text through a target model, wherein the target model has the function of generating corresponding images and texts based on the input multimedia content and display style browsed by the user.

[0018] That is to say, the terminal can use the first data, the second data and the target display style as input, and generate the target image and the target text through the target model. In this way, by generating the target image and the target text through the target model, the diversity of the target image and the target text can be increased, and the user experience can be improved. The target model is a model that can generate corresponding images and texts according to the multimedia content and display style browsed by the input user. In this way, according to user needs, a target image and / or target text that meets the target display style can be generated, so that the pushed target image and target text are more matched with the user, and the user experience is improved.

[0019] In a possible implementation, the target model includes an image generation model and a text generation model; taking the first data, the second data and the target display style as input, and generating the target image and the target text through the target model, includes: taking the first data and the target display style as input, and generating the target image through the image generation model; taking the second data as input, and generating the target text through the text generation model. In this way, a better generation effect can be achieved, so that the generated target image and target text are more in line with the requirements.

[0020] In a possible implementation, when the first data includes at least metadata, the method further includes: obtaining a feature vector of the first data; generating a first descriptive word according to the metadata and the target display style, and generating a second descriptive word according to the second data, wherein the first descriptive word is used to describe the characteristics of the image to be generated, and the second descriptive word is used to describe the characteristics of the text to be generated; taking the first data and the target display style as input, and generating a target image through an image generation model, including: taking the feature vector of the first data and the first descriptive word as input, and generating a target image through an image generation model; taking the second data as input, and generating a target text through a text generation model, including: taking the second descriptive word as input, and generating a target text through a text generation model. In this way, the target model can determine the characteristics of the image and text to be generated, thereby outputting a target image and target text that better meet the requirements.

[0021] In a possible implementation, the second descriptor may also be generated based on the second data and metadata, or the second descriptor may be generated based on the second data, metadata, and medium data. In this way, the target text generated by the second descriptor may include more content, further improving the interpretability of the target image and improving the user experience.

[0022] In a possible implementation, the method further includes: displaying a first control; receiving a second operation of the user on the first control; in response to the second operation, deleting a target image and a target text in an image library corresponding to the user; the image library stores a plurality of images and texts generated based on usage data when the user browses multimedia content and the multimedia content browsed by the user.

[0023] That is to say, when the target image and target text are provided by the terminal itself, the terminal also stores an image library corresponding to the user, and the image library stores multiple target images and target texts generated by the above method. After displaying the target image and target text, the terminal can also display a delete control. After the user performs a second operation on the delete control, the terminal can delete the displayed target image and target text from the above image library. In this way, content that the user is not interested in is no longer pushed to the user, thereby improving the user experience.

[0024] In a possible implementation, the first application is a lock screen application; the displaying of the target image and the target text includes: displaying the target image in the lock screen interface; receiving a third operation of the user in the lock screen interface, and displaying the target text.

[0025] That is, the first application opened by the user may be a lock screen application, and the specific method of displaying the target image and the target text may be to display the target image on the lock screen interface, and then display the target text when receiving a third operation (such as swiping up) from the user. In this way, the target text does not block the main body of the target image, thereby improving the user experience.

[0026] In one possible implementation, the above-mentioned display of the target image and the target text includes: the first application is a desktop application, and the target image and the target text are used as desktop wallpaper and displayed on the desktop; or, the first application is an off-screen application, and the target image and the target text are used as widgets and displayed on the off-screen interface; or, the first application is a system application, and the target image and the target text are used as startup images and displayed on the startup interface of the first application.

[0027] That is to say, the first application opened by the user can be a desktop application, in which case the target image and target text can be displayed as desktop wallpaper. The first application opened by the user can also be an off-screen application, in which case the target image and target text can be used as widgets and displayed on the off-screen interface. The first application opened by the user can be a system application, in which case the target image and target text can be used as startup images and displayed on the startup interface of the system application. The method of the embodiment of the present application can be applied to a variety of scenarios, to enhance the relevance of the images displayed in each scenario to the user, and to enhance the diversity of the displayed content, thereby enhancing the user experience.

[0028] According to the second aspect of the embodiment of the present application, the present application also provides an image generation method, applied to a server, comprising: obtaining first data and second data of a terminal, wherein the first data is used to indicate the multimedia content browsed by a user corresponding to the terminal using a second application, wherein the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data including at least one of the text, audio, video and image of the multimedia content; the second data is used to indicate usage data when the user uses the second application to browse the multimedia content; based on the first data and the second data, generating a target image and a target text; wherein the target image is at least related to the first data, and the target text is at least related to the second data; and sending the target image and the target text to the terminal.

[0029] That is, the server can generate target images and target texts related to the multimedia content browsed by the user in the terminal according to the multimedia content browsed by the user, and send the generated target images and target texts to the terminal. In this way, the server can provide the terminal with a large number of images that are strongly related to the user, thereby improving the user experience of the terminal.

[0030] In a possible implementation, the target image is also related to the target display style, and / or the target text is also related to the target display style.

[0031] In a possible implementation manner, the method further includes: receiving a target display style sent by the terminal, where the target display style is a display style selected by the user from a plurality of candidate display styles displayed by the terminal.

[0032] In a possible implementation, the target display style is a display style that has the highest matching degree with the user among preset display styles determined according to multimedia content browsed by the user.

[0033] In one possible implementation, the above-mentioned generating a target image and a target text based on the first data and the second data includes: taking the first data, the second data and the target display style as input, and generating the target image and the target text through a target model, wherein the target model has the function of generating corresponding images and texts based on the input multimedia content and display style browsed by the user.

[0034] That is to say, the server can use the first data, the second data and the target display style as input, and generate the target image and the target text through the target model. In this way, by generating the target image and the target text through the target model, the diversity of the target image and the target text can be increased, and the user experience can be improved. The target model is a model that can generate corresponding images and texts according to the multimedia content and display style browsed by the input user. In this way, according to the user's needs, the target image and / or target text that conforms to the target display style can be generated, so that the pushed target image and target text are more compatible with the user, and the user experience is improved.

[0035] In a possible implementation, the target model includes an image generation model and a text generation model; taking the first data, the second data and the target display style as input, and generating the target image and the target text through the target model, includes: taking the first data and the target display style as input, and generating the target image through the image generation model; taking the second data as input, and generating the target text through the text generation model. In this way, a better generation effect can be achieved, so that the target image and target text generated by the server are more in line with the requirements.

[0036] In a possible implementation, when the first data includes at least metadata, the method further includes: obtaining a feature vector of the first data; generating a first descriptive word according to the metadata and the target display style, and generating a second descriptive word according to the second data, wherein the first descriptive word is used to describe the characteristics of the image to be generated, and the second descriptive word is used to describe the characteristics of the text to be generated; using the first data and the target display style as input, and generating a target image through an image generation model, including: using the feature vector of the first data and the first descriptive word as input, and generating a target image through an image generation model; using the second data as input, and generating a target text through a text generation model, including: using the second descriptive word as input, and generating a target text through a text generation model. In this way, the target model on the server can determine the characteristics of the image and text to be generated, thereby outputting a target image and target text that better meet the requirements.

[0037] In a possible implementation, the method further includes: receiving first information from a terminal; in response to the first information, deleting a target image and a target text in an image library corresponding to the user; the image library stores a plurality of images and texts generated based on the usage data of the user when browsing multimedia content and the multimedia content browsed by the user. In this way, content that the user is not interested in is no longer pushed to the user, thereby improving the user experience.

[0038] According to a third aspect of the embodiments of the present application, the present application further provides a device having the function of implementing the electronic device behavior in the method of the first aspect or the second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, for example, a processing module, a sending module, and a display module.

[0039] According to a fourth aspect of an embodiment of the present application, the present application also provides an electronic device, comprising a memory and a processor, the memory being used to store a computer program, and the processor being used to execute the computer program to implement the method of the first aspect above.

[0040] According to a fifth aspect of an embodiment of the present application, the present application further provides a server, comprising a memory and a processor, the memory being used to store a computer program, and the processor being used to execute the computer program, so as to implement the method of the second aspect mentioned above.

[0041] According to the sixth aspect of the embodiments of the present application, the present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to implement the method of the first aspect or the second aspect mentioned above.

[0042] According to the seventh aspect of the embodiments of the present application, the present application also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying a computer-readable code. When the computer-readable code is running in an electronic device, the processor in the electronic device executes the method of the first aspect or the second aspect above.

[0043] According to the eighth aspect of the embodiments of the present application, the present application also provides a chip, which includes a memory and a processor, the memory is used to store computer programs, and the processor is used to call and run computer programs from the memory to execute the method described in the first aspect or the second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of a lock screen interface provided for related technologies;

[0045] Figure 2 A schematic diagram of another screen-off interface provided for related technologies;

[0046] Figure 3 A simplified schematic diagram of a system architecture for applying a method of an embodiment of the present application provided in an embodiment of the present application;

[0047] Figure 4 A schematic diagram of the composition of a terminal provided in an embodiment of the present application;

[0048] Figure 5 A schematic diagram of the composition of a server provided in an embodiment of the present application;

[0049] Figure 6 A flowchart of an image display method provided in an embodiment of the present application;

[0050] Figure 7 A flowchart of another image display method provided in an embodiment of the present application;

[0051] Figure 8 A schematic diagram of a terminal interface provided in an embodiment of the present application;

[0052] Fig. 9 A schematic diagram of generating a target image and a target text provided in an embodiment of the present application;

[0053] Fig.10 A flowchart of an image generation method provided in an embodiment of the present application;

[0054] Fig.11 A schematic diagram of another terminal interface provided in an embodiment of the present application;

[0055] Fig.12 A schematic diagram of another terminal interface is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0056] The lock screen function, desktop function and screen off function are widely used in mobile phones, computers, tablets, car computers, TVs and other terminals. The normal use of the terminal is inseparable from these functions. For example, when using a terminal, the user needs to enter the lock screen interface first, and then unlock it to use other functions of the terminal. For another example, when using an application installed in the terminal, it is generally necessary to enter the desktop first and open the application by clicking the application icon displayed on the desktop.

[0057] Based on the above, the information display of the lock screen interface, desktop and off-screen interface is an important issue that needs to be considered. In the related art, there are several information display methods as follows.

[0058] First, the user selects an image stored in the terminal as the lock screen wallpaper or desktop wallpaper to display image information on the lock screen interface and desktop; the user can also select a theme as the theme of the mobile phone, where the theme is designed and developed by a company or individual developer to personalize the terminal skin interface such as the lock screen, wallpaper, icon, notification bar, text messages, dialing, contacts, settings, etc.; the user also selects a widget from the preset widgets of the terminal and displays the widget on the off-screen interface to transmit relevant information through the widget, such as the clock widget can transmit time information, etc.

[0059] The lock screen wallpaper, desktop wallpaper, themes and widgets on the screen-off interface will not change after the user selects them, which results in a lack of diversity in the lock screen interface, desktop and screen-off interface, making it difficult to arouse user interest. However, the lock screen wallpaper, desktop wallpaper and widgets on the screen-off interface can actually be used as traffic entrances to divert traffic to applications on the terminal, which wastes a good traffic display opportunity.

[0060] Second, the lock screen wallpaper is a random image or graphic magazine pushed by the terminal system or the application in the terminal. This makes the lock screen content often less relevant to the user. The user is unfamiliar with the lock screen content. Even if they spend a certain amount of time to read the lock screen content, due to the low relevance of the lock screen content to the user, the user is often difficult to be interested in the content displayed on the lock screen interface, and will quickly skip the lock screen content, wasting the traffic display opportunity brought by the lock screen wallpaper as a traffic entrance.

[0061] Moreover, the random images or illustrated magazines displayed on the lock screen interface often come from a single wallpaper provider, which makes the lock screen source single and cannot meet the personalized needs of users.

[0062] Third, in addition to displaying wallpaper on the lock screen, there is another lock screen display scheme, that is, displaying the interface information of a certain application running in the background in the lock screen scene. For example, when a user uses a music software on the terminal to play a song, Figure 1 As shown, the lock screen interface displays the song title, singer name, album name, album cover, lyrics, etc. of the song currently being played on the terminal, and can also display controls for controlling song playback.

[0063] In this solution, the lock screen interface is generated by a template and lacks diversity.

[0064] It can be seen that in the related art, the display solutions for the lock screen interface, desktop and screen-off interface have problems such as templateization, monotony and lack of diversity, and the displayed content lacks relevance to the user and is difficult to arouse the user's interest.

[0065] Based on this, the present application proposes an image display method and an image generation method. When a user opens a first application such as a desktop application, a lock screen application, an off-screen application or a system application, a target image related to multimedia content browsed by the user on other applications can be displayed, as well as a target text related to usage data when the user browses multimedia content on other applications.

[0066] In other words, the image display method applied to the terminal includes: receiving a first operation of a user to open a first application; in response to the first operation, displaying a target image and a target text; wherein the target image is at least related to first data, the first data is used to indicate multimedia content browsed by the user using a second application, and the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data includes at least one of the text, audio, video and image of the multimedia content; the target text is at least related to second data, and the second data is used to indicate usage data when the user uses the second application to browse the multimedia content.

[0067] In this way, the target image and target text related to the multimedia content browsed by the user can be displayed in the interface of the lock screen application, desktop application, off-screen application or system application. In other words, the target image and target text are related to the user. And the multimedia content browsed by the user in history is often large, so the corresponding target image and target text can be generated according to the multiple multimedia contents browsed by the user, which also improves the diversity of the lock screen, desktop, off-screen and other interfaces. In addition, such target images and target texts avoid the monotony of the images generated by the template, can increase the user's usage rate of the lock screen, desktop, off-screen and other interfaces, and improve the user's experience.

[0068] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the relevant laws and regulations and do not violate public order and good morals. For example, in the technical solution of this application, the processing of user personal information is carried out with the authorization of the user, which is uniformly explained here and will not be repeated below.

[0069] For the use scenario of the present application, the above-mentioned first application can be an application that can display an image when it is turned on. For example, the first application in the present application can be a lock screen application, a desktop application, an off-screen application, a system application, and the like, that is, the present application can be applied to wallpaper display scenarios of the lock screen, desktop, off-screen, and application startup interfaces. For example, when the first application is a lock screen application, the target image and the target text can be displayed as a lock screen wallpaper on the lock screen interface; when the first application is a desktop application, the target image and the target text can be combined as a desktop wallpaper for display; when the first application is any application of a desktop application, a lock screen application, or an off-screen application, it can also be based on the target image and the target text. Generate a corresponding theme, and display the target image and the target text in the above three applications; when the first application is a system application, the target image and the target text can be displayed on the startup interface of the system application, that is, the target image and the target text can be combined as the wallpaper of the application startup interface. In addition, in the case where there is cooperation between the terminal manufacturer and the third-party application installed on the terminal, the target image and the target text can also be displayed on the startup interface of the third-party application other than the system application.

[0070] System applications refer to applications developed by terminal manufacturers, such as applications that are pre-installed when the terminal is purchased.

[0071] Among them, the off-screen interface is the display interface of the system application. It is an interface that the terminal may enter before entering the lock screen interface when it is locked. Generally, widgets can be added to this interface. Widgets can also be called widgets, window widgets, window widgets, widgets, etc. Figure 2 As shown, Figure 2 201 in the figure is a widget, which is used to display time. In the embodiment of the present application, the target image and the target text can be used as widgets and displayed on the off-screen interface.

[0072] The first operation can also be an operation that triggers the opening of the first application. For example, in a lock screen scenario, the first operation can be an operation that can wake up the screen, such as pressing the lock screen button of the terminal, lifting the terminal, and so on. In a desktop scenario, the first operation can be an operation to return to the desktop, such as closing an application, returning to the desktop, and so on. In an off-screen scenario, the first operation can be an operation that can wake up the off-screen interface, such as locking the screen or lifting the terminal. In the case where the first application is a system application, the first operation can be an operation on the icon of the system application, an operation to switch from the background interface to the system application, and so on. Of course, as described above, the first application can also be other applications that can display images when turned on, such as third-party applications installed on the terminal and released by application developers.

[0073] The second application may be an application including multimedia content. For ease of understanding, the second application is referred to as a multimedia application below. The multimedia application mentioned in the embodiment of the present application may be a multimedia application developed by a terminal manufacturer, such as a video, music, reading, podcast, and game application pre-installed on the terminal. In this embodiment, a target image may be generated based on a user's usage record of the multimedia application developed by the terminal manufacturer, which may increase the traffic of these multimedia applications and increase the user's frequency of use of these multimedia applications.

[0074] When there is cooperation between the terminal manufacturer and the application developer, and the user agrees to use the usage data of the application developed by the application developer to generate wallpapers or widgets for the lock screen, desktop, and off screen interfaces, the multimedia application can also be a multimedia application developed by the application developer, such as a third-party video application installed on the terminal.

[0075] In this application, users can browse multimedia content through multimedia applications in the terminal. Multimedia applications are applications that contain multimedia content, and multimedia content can include audio, video, text, and the like. For example, when the multimedia application is a video application and the multimedia content is video, users can browse videos through the video application on the terminal. For example, when the multimedia application is a music application and the multimedia content is music, users can listen to music through the music application on the terminal. For example, when the multimedia application is a reading application and the multimedia content is the text in a book, users can read books through the reading application on the terminal. For example, when the multimedia application is a game application and the multimedia content is the screenshots and introduction of the game, users can browse the game introduction, install the game, enter the game application to play, and so on through the game application on the terminal. It should be noted that the game application here does not refer to the application corresponding to the game itself, but to the application that manages the games installed on the terminal. In some terminals, the game application is called "game center".

[0076] As for the relationship between the first application and the second application, in the above example, in the lock screen, desktop and screen off scenarios, the first application and the second application can be different applications, for example, the first application is a lock screen application, and the second application can be a video application. In addition, in the scenario where the first application is a system application, the first application and the second application can be the same application or different applications. For example, in the case where the first application is a clock application and the second application is a multimedia application, then the first application and the second application are different applications. In the case where the first application is a multimedia application, the second application can be the same multimedia application as the first application, or it can be a multimedia application different from the first application. For example, in the case where the first application is a video application developed by a terminal manufacturer, the second application can be a music application developed by the terminal manufacturer, the second application can also be the video application, or the second application can be a third-party video application installed on the terminal.

[0077] For the displayed target image and target text, the target image is related to the multimedia content browsed by the user, that is, the target image has the same theme as the multimedia content, or is related to the specific content of the multimedia content. This application does not limit the specific form of the target image, and any image related to the multimedia content can be used as the target image of this application.

[0078] For example, if the multimedia content browsed by the user is a video about the scenery of a certain area, the target image may be an image of the scenery of the area, for example, the target image may be a video frame in the video, the target image may also be a classic image mentioned in the subtitles of the video, the target image may also be a publicly available image of the scenery of the area retrieved from the Internet, or an image related to the content in the video generated based on one or more of the video frames, audio and subtitles in the video.

[0079] For another example, if the multimedia content browsed by the user is a song, the target image may be an image related to the song, such as the album cover of the song. If the song has a corresponding music video (MV), the target image may be the video frame of the music video. For example, if the song is related to the theme of youth, the target image may be an image related to the theme of youth retrieved on the Internet. The target image may also be an image related to the song generated based on one or more of the audio, lyrics and video frames of the MV of the song.

[0080] For another example, if the multimedia content browsed by the user is a book, the target image may be an image related to the book, such as the cover of the book, or an image related to the book retrieved from the Internet. The target image may also be an image related to the book generated based on one or more of the cover of the book or the text in the book.

[0081] For another example, if the multimedia content browsed by the user is a certain game, the target image may be an image about the game. Similar to the above example, the target image may be a promotional video or promotional image of the game. The target image may also be an image related to the game retrieved from the Internet. The target image may also be an image related to the game generated based on one or more of the video frames of the promotional video of the game, the audio corresponding to the promotional video, and the game introduction.

[0082] For the example where the target image is generated based on multimedia content, the specific method for generating the image will be described in detail below and will not be described here. In addition, the above example is only to help understand the present solution and does not represent a limitation on the embodiments of the present application.

[0083] The target text is a text about the usage data of the user browsing multimedia content. The usage data is data describing the user's browsing behavior of multimedia content, which may include when the user browsed the multimedia content, the progress of the user browsing the multimedia content, whether the user has collected or liked the multimedia content, whether the user has commented on the multimedia content, and the identification of the multimedia content corresponding to the user's behavior, etc. The target text and the usage data may be related to the target text including specific usage data and the title of the multimedia content corresponding to the usage data, wherein the title of the multimedia content corresponding to the usage data can be determined according to the identification of the multimedia content included in the usage data. For example, the target text may be "You browsed "Challenger" on January 1, 2023", "You browsed "Challenger" on January 1, 2023, come and reminisce", etc. For another example, if the target image is related to the content that the user is watching before locking the screen, the target text may be "You are watching Challenger". In addition, in order to facilitate the user to recall the relevant content, the target text may also include data used to describe the multimedia content, such as the target text may include metadata, media data, etc. of the multimedia content. The metadata and media data of multimedia content will be described in detail below and will not be elaborated here.

[0084] By displaying the target text, the user can be reminded that they have browsed the multimedia content corresponding to the target image, which increases the interpretability of the target image. The target text can make the user aware of the content of the target image and make the user aware that the target image is related to the multimedia content they browsed, thereby improving the user experience.

[0085] In addition, in addition to being related to the above data, in some embodiments of the present application, the target image and the target text may also be related to the target display style, and / or the target text may also be related to the target display style.

[0086] As for which multimedia content the displayed target image and target text are related to, the displayed target image and target text may be related to the user's history of browsing multimedia content in the past period of time. The specific method will be described in detail later and will not be repeated here. In addition, the displayed target image and target text may also be related to the multimedia content that the user was browsing before locking the screen, closing the application, or returning to the desktop. For example, when it is determined that the user has browsed new multimedia content, a target image may be generated based on the multimedia content, and a target text may be generated based on the usage data of the multimedia content. After opening the lock screen interface, screen-off interface, application, or desktop, the target image and target text may be displayed.

[0087] For the users in this application, in some embodiments, the user can correspond to the terminal device itself. Based on this, the target image and target text can be generated only according to the local browsing history of the second application on the one terminal, that is, the browsing history of the second application on the terminal is used as data related to the user to generate the target image and target text. In some other embodiments, the user can correspond to the user account, the terminal can log in to the user account, and different terminals can log in to the same user account. Based on this, the target image and target text can also be generated according to the browsing history of the second application on the terminal logged in to the same user account, that is, the browsing history of the second application on multiple terminals logged in to the same account can be used as data related to the user to generate the target image and target text.

[0088] For the application device of the method of the present application, the method of the present application can be executed by the terminal alone, so as to protect the privacy of the user. The solution of the present application can also be implemented by the interaction between the terminal and the server. The solution of the present application can be implemented by the interaction between the terminal and the server, which can reduce the processing burden of the terminal and save the computing power of the terminal.

[0089] Next, the embodiment of the present application will be explained by taking the interaction between the terminal and the server and applying the method of the embodiment of the present application to the lock screen interface as an example.

[0090] like Figure 3 As shown, Figure 3 is a simplified schematic diagram of a system architecture to which the above method can be applied provided by an embodiment of the present application. Figure 3 As shown, the system at least includes a terminal 31 and a server 32 .

[0091] In specific implementation, the terminal 31 may be a mobile phone, a tablet computer, a handheld computer, a personal computer (PC), a cellular phone, a personal digital assistant (PDA), a wearable device (such as a smart watch), a smart home device (such as a TV), a car computer, a game console, and an augmented reality (AR) / virtual reality (VR) device, etc. This embodiment does not impose any special restrictions on the specific device form of the terminal 31. Among them, the terminal 31 may include one or more applications. The application may be a system application, such as a lock screen application, a desktop application, a screen-off application, and a video application, a reading application, a game application, and a music application pre-configured on the mobile phone by the terminal manufacturer. The application may also be a third-party application.

[0092] In addition, the system may include multiple terminals, and the server 32 may interact with the multiple terminals at the same time to provide target images and target texts for the multiple terminals.

[0093] Please refer to Figure 4 , is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The methods in the following embodiments can be implemented in a terminal having the above hardware structure.

[0094] like Figure 4 As shown, the terminal 31 may include a processor 110, an external memory interface 120, an internal memory 121, a USB interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a radio frequency module 150, a communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a SIM card interface 195, etc. The sensor module may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor, etc.

[0095] The touch sensor 180K, microphone 170C, antenna 1, antenna 2, radio frequency module 150, and communication module 160 can be used as input devices of the terminal 31 to receive information input by a user or other devices. The speaker 170A, receiver 170B, and display screen 194 can be used as output devices of the terminal 31 to output information input by a user or information provided to a user and various menus of the terminal 31.

[0096] The structure shown in the embodiment of the present invention does not constitute a limitation on the terminal 31. It may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0097] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processor (NPU), etc. Different processing units may be independent devices or integrated in the same processor.

[0098] The controller can be the decision maker that directs the various components of the terminal 31 to work in coordination according to the instructions. It is the nerve center and command center of the terminal 31. The controller generates an operation control signal according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0099] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. Instructions or data that have just been used or are used cyclically by the processor may be stored. If the processor needs to use the instruction or data again, it may be directly called from the memory. Repeated access is avoided, the waiting time of the processor is reduced, and the efficiency of the system is improved.

[0100] In the embodiment of the present application, the processor 110 can be used to execute the steps of the embodiment of the present application, for example, it can be used to control the display screen 194 to display the target image and the target text, and after receiving the first operation of the user, the first operation is processed to trigger the process of controlling the display screen 194 to display the target image and the target text, etc. In addition, the processor 110 can also complete the functions of obtaining the target style and the like through the steps of the following embodiments, and the specific description is detailed below.

[0101] In some embodiments, the processor 110 may include an interface, wherein the interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0102] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor may include multiple sets of I2C buses. The processor may be coupled to a touch sensor, a charger, a flash, a camera, etc. through different I2C bus interfaces. For example, the processor may be coupled to a touch sensor through an I2C interface, so that the processor and the touch sensor communicate through the I2C bus interface to realize the touch function of the terminal 31.

[0103] The I2S interface can be used for audio communication. In some embodiments, the processor can include multiple I2S buses. The processor can be coupled to the audio module via the I2S bus to achieve communication between the processor and the audio module. In some embodiments, the audio module can transmit audio signals to the communication module via the I2S interface to achieve the function of answering calls through a Bluetooth headset.

[0104] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module and the communication module can be coupled via a PCM bus interface. In some embodiments, the audio module can also transmit audio signals to the communication module via the PCM interface to realize the function of answering calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication, and the sampling rates of the two interfaces are different.

[0105] The UART interface is a universal serial data bus for asynchronous communication. The bus is a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is generally used to connect the processor and the communication module 160. For example, the processor communicates with the Bluetooth module through the UART interface to realize the Bluetooth function. In some embodiments, the audio module can transmit an audio signal to the communication module through the UART interface to realize the function of playing music through a Bluetooth headset.

[0106] The MIPI interface can be used to connect the processor to peripheral devices such as display screens and cameras. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor and the camera communicate via the CSI interface to implement the shooting function of the terminal 31. The processor and the display screen communicate via the DSI interface to implement the display function of the terminal 31.

[0107] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor to a camera, a display, a communication module, an audio module, a sensor, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0108] The USB interface 130 may be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface may be used to connect a charger to charge the terminal 31, or may be used to transmit data between the terminal 31 and a peripheral device. It may also be used to connect headphones to play audio through the headphones. It may also be used to connect other electronic devices, such as AR devices, etc.

[0109] The interface connection relationship between the modules shown in the embodiment of the present invention is only for illustrative purposes and does not constitute a structural limitation on the terminal 31. The terminal 31 may use different interface connection modes in the embodiment of the present invention, or a combination of multiple interface connection modes.

[0110] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module can receive charging input from a wired charger through a USB interface. In some wireless charging embodiments, the charging management module can receive wireless charging input through a wireless charging coil of the terminal 31. While the charging management module is charging the battery, it can also power the terminal device through the power management module 141.

[0111] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module receives input from the battery and / or the charging management module to power the processor, internal memory, external memory, display screen, camera, and communication module. The power management module can also be used to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), etc. In some embodiments, the power management module 141 can also be set in the processor 110. In some embodiments, the power management module 141 and the charging management module can also be set in the same device.

[0112] The wireless communication function of the terminal 31 can be realized through the antenna 1, the antenna 2, the radio frequency module 150, the communication module 160, the modem and the baseband processor.

[0113] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 31 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of the antennas. For example, a cellular network antenna can be reused as a wireless local area network diversity antenna. In some embodiments, the antenna can be used in combination with a tuning switch.

[0114] The RF module 150 can provide a communication processing module for wireless communication solutions including 2G / 3G / 4G / 5G, etc., applied to the terminal 31. It can include at least one filter, a switch, a power amplifier, a low noise amplifier (LowNoise Amplifier, LNA), etc. The RF module receives electromagnetic waves from the antenna 1, and filters, amplifies, and processes the received electromagnetic waves, and transmits them to the modem for demodulation. The RF module can also amplify the signal modulated by the modem, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the RF module 150 can be set in the processor 150. In some embodiments, at least some of the functional modules of the RF module 150 can be set in the same device as at least some of the modules of the processor 110.

[0115] The modem may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After the low-frequency baseband signal is processed by the baseband processor, it is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker, a receiver, etc.), or displays an image or video through a display screen. In some embodiments, the modem may be an independent device. In some embodiments, the modem may be independent of the processor and be set in the same device as the RF module or other functional modules.

[0116] The communication module 160 can provide a communication processing module for wireless communication solutions including wireless local area networks (WLAN), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), nearfield communication (NFC), infrared (IR), etc., which are applied to the terminal 31. The communication module 160 can be one or more devices integrating at least one communication processing module. The communication module receives electromagnetic waves via antenna 2, modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor. The communication module 160 can also receive the signal to be sent from the processor, modulate the frequency, amplify it, and convert it into electromagnetic waves for radiation through antenna 2.

[0117] In some embodiments, the antenna 1 of the terminal 31 is coupled to the radio frequency module, and the antenna 2 is coupled to the communication module. This allows the terminal 31 to communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long-term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS)) and / or satellite based augmentation system (SBAS).

[0118] Through the communication function of the terminal, the terminal 31 can communicate with the server 32, so that the target image and target text can be obtained from the server 32 through the above module. In some embodiments, the mobile phone can also send the target display style, the user's use data of the multimedia content in the second application to the server through the above module.

[0119] The terminal 31 implements the display function through the GPU, the display screen 194, and the application processor, etc. For example, the target image and the target text can be displayed through the above components. In some embodiments, the mobile phone can also display the controls, alternative display styles, etc. through the above modules, so as to complete the interaction between the mobile phone and the user. The GPU is a microprocessor for image processing, which connects the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0120] The display screen 194 is used to display images, videos, etc. The display screen includes a display panel. The display panel can be LCD (liquid crystal display), OLED (organic light-emitting diode), active-matrix organic light-emitting diode or active-matrix organic light-emitting diode (AMOLED), Miniled, MicroLed, Micro-oLed, quantum dot light emitting diodes (QLED), etc. In some embodiments, the terminal 31 may include 1 or N display screens, where N is a positive integer greater than 1.

[0121] Still Figure 1 As shown, the terminal 31 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen and application processor.

[0122] The ISP is used to process the data fed back by the camera. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0123] Camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal 31 may include 1 or N cameras, where N is a positive integer greater than 1.

[0124] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the terminal 31 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0125] The video codec is used to compress or decompress digital video. Terminal 31 can support one or more codecs. In this way, terminal 31 can play or record videos in multiple coding formats, such as MPEG1, MPEG2, MPEG3, MPEG4, etc.

[0126] NPU is a neural network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and can also continuously self-learn. Through NPU, applications such as intelligent cognition of terminal 31 can be realized, such as image recognition, face recognition, voice recognition, text understanding, etc.

[0127] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 31. The external memory card communicates with the processor via the external memory interface to implement a data storage function, such as storing music, video and other files in the external memory card.

[0128] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the mobile phone by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a lock screen function, a desktop function, a screen off function), etc. The data storage area can store data created during the use of the mobile phone (such as a target image and target text to be displayed), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0129] In the embodiment of the present application, the internal memory 121 may specifically include RAM (random access memory) and ROM (read-only memory). RAM is an internal memory that directly exchanges data with the processor 110, also called main memory (or internal memory). It can be read and written at any time, and the speed is very fast. It is usually used as a temporary data storage medium for operating systems or other running programs. The data stored in ROM can be easily read out, unlike RAM, which can be quickly and conveniently rewritten. However, the data stored in ROM is relatively stable, and the stored data will not change after power failure; its structure is relatively simple and it is convenient to read out, so it is often used to store various fixed programs and data.

[0130] RAM and ROM may include one or more partitions.

[0131] Take ROM as an example, Figure 2 As shown, the ROM may include system partitions (such as System partitions, Recovery partitions), program partitions (such as Data partitions), and storage partitions (such as SDCard partitions). Among them, the system partition can be used to store operating systems (such as Android systems), restore backup systems, swap space, hardware underlying space and other resources; the program partition is used to store third-party APPs installed on the terminal. For each APP, the terminal can create a corresponding Data directory in the Data partition. For example, if there is an APP package named weixin.com, a directory named weixin.com can be created in the Data partition. The application data generated by the operation of the APP, such as chat records, transfer files, etc., can all be stored in a directory named weixin.com, and the APP can only operate the data in this directory, and cannot operate the directories of other APPs; the storage partition is equivalent to the "mobile hard disk" identified after the mobile phone is connected to the PC. This part of the space can be freely used by the user and can store data packets, music, pictures, videos and other data of large games.

[0132] The terminal 31 can also dynamically create a new partition in the RAM or ROM. For example, if there is 2G of unoccupied space in the 4G ROM, the terminal 31 can create a partition named "aaa" in the 2G space. Of course, the terminal 31 can also dynamically destroy the newly created partition, and the data in the partition will be destroyed when the new partition is destroyed.

[0133] In the embodiment of the present application, when the terminal 31 is restored to factory settings, one or more partitions in the ROM of the terminal 31 are generally formatted. For example, the terminal 31 can format the Data partition in the ROM, delete all files and folders in the Data partition, so that all applications installed in the Data partition and the data in each application are cleared, thereby restoring the terminal 31 to the state when it was sold from the factory.

[0134] The terminal 31 can implement audio functions such as music playing and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor.

[0135] The audio module is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signal. The audio module can also be used to encode and decode audio signals. In some embodiments, the audio module can be arranged in the processor 110, or some functional modules of the audio module can be arranged in the processor 110.

[0136] The speaker 170A, also called a "speaker", is used to convert an audio electrical signal into a sound signal. The terminal 31 can listen to music or listen to a hands-free call through the speaker.

[0137] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the terminal 31 receives a call or voice message, the voice can be received by placing the receiver close to the ear.

[0138] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can put the sound signal into the microphone by putting the mouth close to the microphone. Terminal 31 can be provided with at least one microphone. In some embodiments, terminal 31 can be provided with two microphones, which can not only collect sound signals but also realize noise reduction function. In some embodiments, terminal 31 can also be provided with three, four or more microphones to realize sound signal collection, noise reduction, and can also identify the sound source, realize directional recording function, etc.

[0139] The earphone interface 170D is used to connect a wired earphone, and the earphone interface may be a USB interface, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0140] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor can be set on the display screen. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can be a parallel plate including at least two conductive materials. When a force acts on the pressure sensor, the capacitance between the electrodes changes. The terminal 31 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen, the terminal 31 detects the touch operation intensity according to the pressure sensor. The terminal 31 can also calculate the touch position according to the detection signal of the pressure sensor. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, an instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, an instruction to create a new short message is executed.

[0141] The gyroscope sensor 180B can be used to determine the motion posture of the terminal 31. In some embodiments, the angular velocity of the terminal 31 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor. The gyroscope sensor can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor detects the angle of the terminal 31 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the terminal 31 through reverse movement to achieve anti-shake. The gyroscope sensor can also be used for navigation and somatosensory game scenes.

[0142] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the terminal 31 calculates the altitude through the air pressure value measured by the air pressure sensor to assist positioning and navigation.

[0143] The magnetic sensor 180D includes a Hall sensor. The terminal 31 can use the magnetic sensor to detect the opening and closing of the flip leather case. In some embodiments, when the terminal 31 is a flip phone, the terminal 31 can detect the opening and closing of the flip cover according to the magnetic sensor. Then, according to the detected opening and closing state of the leather case or the opening and closing state of the flip cover, the flip cover automatic unlocking and other features are set.

[0144] The acceleration sensor 180E can detect the magnitude of the acceleration of the terminal 31 in various directions (generally three axes). When the terminal 31 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the terminal posture and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0145] The distance sensor 180F is used to measure the distance. The terminal 31 can measure the distance by infrared or laser. In some embodiments, when shooting a scene, the terminal 31 can use the distance sensor to measure the distance to achieve fast focusing.

[0146] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. Infrared light is emitted outward through the light emitting diode. The photodiode is used to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 31. When insufficient reflected light is detected, it can be determined that there is no object near the terminal 31. The terminal 31 can use the proximity light sensor to detect that the user holds the terminal 31 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.

[0147] The ambient light sensor 180L is used to sense the ambient light brightness. The terminal 31 can adaptively adjust the display brightness according to the perceived ambient light brightness. The ambient light sensor can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor can also cooperate with the proximity light sensor to detect whether the terminal 31 is in a pocket to prevent accidental touch.

[0148] The fingerprint sensor 180H is used to collect fingerprints. The terminal 31 can use the collected fingerprint characteristics to realize fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering calls, etc. In the lock screen application, the fingerprint sensor can be used to authenticate the identity during the unlocking process. In the embodiment of the present application, the mobile phone can complete the function of unlocking on the lock screen interface through the fingerprint sensor.

[0149] The temperature sensor 180J is used to detect temperature. In some embodiments, the terminal 31 uses the temperature detected by the temperature sensor to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor exceeds a threshold, the terminal 31 executes to reduce the performance of the processor located near the temperature sensor to reduce power consumption and implement thermal protection.

[0150] The touch sensor 180K, also known as a "touch panel", can be disposed on a display screen and used to detect a touch operation applied thereto or near the touch sensor. The detected touch operation can be transmitted to an application processor to determine the type of touch event and provide a corresponding visual output through the display screen.

[0151] The touch sensor can be used to identify the user's operations on the phone, such as the user's operation of opening the lock screen and other applications, the user's selection and click on a certain control, etc.

[0152] The bone conduction sensor 180M can obtain a vibration signal. In some embodiments, the bone conduction sensor can obtain a vibration signal of a vibrating bone block of the human body. The bone conduction sensor can also contact the human pulse to receive a blood pressure beat signal. In some embodiments, the bone conduction sensor can also be set in an earphone. The audio module 170 can parse out a voice signal based on the vibration signal of the vibrating bone block of the human body obtained by the bone conduction sensor to realize a voice function. The application processor can parse the heart rate information based on the blood pressure beat signal obtained by the bone conduction sensor to realize a heart rate detection function.

[0153] The key 190 includes a power key, a volume key, etc. The key may be a mechanical key or a touch key. The terminal 31 receives the key input and generates a key signal input related to the user settings and function control of the terminal 31.

[0154] Motor 191 can generate vibration prompts. The motor can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. Touch operations acting on different areas of the display screen can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0155] Indicator 192 may be an indicator light, which may be used to indicate charging status, power changes, messages, missed calls, notifications, etc.

[0156] The SIM card interface 195 is used to connect a subscriber identity module (SIM). The SIM card can be connected to and separated from the terminal 31 by inserting it into the SIM card interface or pulling it out from the SIM card interface. The terminal 31 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface at the same time. The types of the multiple cards can be the same or different. The SIM card interface can also be compatible with different types of SIM cards. The SIM card interface can also be compatible with external memory cards. The terminal 31 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the terminal 31 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 31 and cannot be separated from the terminal 31.

[0157] The software system of the terminal 31 can adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present invention takes the Android system of the layered architecture as an example to exemplify the software structure of the terminal 31.

[0158] The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system library, and the kernel layer.

[0159] In this embodiment, please refer to Figure 5 , is a schematic diagram of the structure of a server provided in an embodiment of the present application. Figure 5 As shown, the server 32 may include a processing module 501, a storage module 502, a communication module 503, and the like.

[0160] The processing module 301 is the control center of the system. For example, the processing module 501 can be any one or a combination of CPU, GPU, field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC).

[0161] The storage module 502 is used to store data, for example, the storage module 502 can be used to store user usage data and data corresponding to multimedia content. The storage module 502 can also store computer instructions, and the processing module 501 reads and executes the computer instructions stored in the storage module 502 to implement the server function.

[0162] The communication module 503 is used for communication between the server and the terminal.

[0163] Next, we will combine Figure 3 and Figure 6 The method of this embodiment is briefly described below. Figure 3 The solid-line boxes represent modules in the server or terminal or actions performed by the modules, and the dotted-line boxes represent data stored in the terminal or server.

[0164] In the system of the embodiment of the present application, Figure 3 As shown, Figure 3 In step 702, the terminal 31 is responsible for sending the usage data generated by the user browsing the multimedia application to the server 32, and is responsible for assisting the user in selecting the display style. The specific process can be seen in the following description. Figure 7 The description of step 702 corresponds to Figure 6 , the terminal can provide the user with a window for turning on the "customized lock screen" function, and the window can also display multiple preset display styles, such as "simple" style, "secondary" style, "fresh" style and "national style" style. The preset display style can also be other display styles, such as "cartoon" style, etc., which are not specifically limited in this embodiment. The user can select a display style from multiple preset display styles, and the selected display style is the style of at least one of the generated target image and target text.

[0165] Figure 3 In Figure 3 As shown in the marked step 703, the server 32 can take the display style and user data selected by the user as input, and as shown in step 704, the target image and target text can be generated according to the above input data through the artificial intelligence model, and the target image and target text together constitute the lock screen image and text. The user data here may include the user's usage data and the data of the multimedia content browsed by the user, and the terminal may send only the user's usage data to the server. Among them, the artificial intelligence model may include an image-text-speech three-tower model for extracting features, as well as an image generation model for generating images and a text generation model for generating text. The user data here includes the user's usage data of multimedia content, as well as the metadata and / or media data corresponding to the multimedia content. The specific implementation process of the above steps 703 and 704 can be found in the following description. Figure 7 Although Figure 3 Two multimedia applications are shown in FIG. 1 , but it should be noted that multimedia applications may also include game applications, reading applications, etc. Figure 6 , the server can generate a lock screen image and text with the display style selected by the user according to the user data and the display style selected by the user.

[0166] like Figure 3 Step 705 is marked, Figure 3 The server 32 can send the generated lock screen image to the terminal 31. The specific implementation process can be seen in the following description. Figure 7 After receiving the lock screen image and text sent by the server 32, the terminal 31 can display the received lock screen image and text on the lock screen interface. Figure 6 , the terminal can display the lock screen graphic generated according to the display style selected by the user. Furthermore, the user can also operate the lock screen graphic, and the terminal can jump to the interface of the multimedia content corresponding to the lock screen graphic in the multimedia application. For example, the lock screen graphic is generated based on the movie "Challenger" that the user has browsed and the display style of "simple" style selected. Figure 6In the example, the terminal can display a lock screen image and text related to the movie "Challenger" in a "simple" style as the lock screen wallpaper. For example, "Challenger" is a movie related to users challenging their limits, and the movie includes a video frame of a user jumping over a cliff. Then the following can be generated: Figure 6 In addition, the user can also operate the lock screen interface to trigger the terminal to directly jump from the lock screen interface to the details interface of the video application "Challenger". It should be noted that "Challenger" here is just an example for the convenience of description and does not correspond to the actual movie.

[0167] It should be noted that Figure 6 and Figure 3 The illustrated embodiments do not limit the present application. For example, in other embodiments of the present application, the terminal 31 may not assist the user in selecting a display style. In this case, the server 32 may only use the user data as input and generate lock screen images and texts through an artificial intelligence model. For another example, the user may not manually turn on the "customized lock screen" function, which may be turned on by default; the user may not select a display style, which may be a default style, or the display style may be a display style that the user may tend to use based on the analysis of the multimedia content that the user has browsed historically. The server 32 may also not generate target images and target texts through an artificial intelligence model.

[0168] Next, we will combine Figure 7 , to explain the embodiments of the present application in detail. In the scenario where the method of the present application is implemented through interaction between a terminal and a server, Figure 7 Among the steps shown in FIG. 7 , step 701 and step 702 shown by dotted lines are optional, that is, step 701 and step 702 may not be performed. Figure 7 It is a flow chart of an embodiment of the present application, including:

[0169] Step 701: The terminal displays multiple candidate display styles and determines a target display style.

[0170] In the embodiment of the present application, the target display style is the display style selected by the user from multiple displayed alternative display styles. The display style is the style of the displayed image and / or text, for example, the display style can be a "simple" style, a "two-dimensional" style, etc. The target display style is the style of the target image and / or target text that needs to be generated in the subsequent steps.

[0171] As a possible implementation, step 701 may specifically include: the terminal displays multiple candidate display styles; and receives a user's selection operation of a target display style from the multiple candidate display styles. The terminal may determine the display style selected by the user as the target display style.

[0172] The selection operation may specifically be a triggering operation of a user on a control corresponding to a target display style in an interface, for example, the user may specifically click on a control corresponding to a target display style in a display screen.

[0173] As an example, the target display style is the style of the target image to be generated, and the terminal is a mobile phone. The user can enable the customized lock screen function in the corresponding setting interface. In response to the user's operation of enabling the customized lock screen function, the mobile phone displays the following Figure 8 The interface shown includes multiple alternative display styles. Figure 8 4 alternative display styles are shown in FIG. 8 , and each alternative display style corresponds to a control. Taking the target display style as the "simple" style as an example, the above selection operation can be a user's click operation on the control 801. In response to the above click operation on the control 801, the mobile phone determines that the target display style is the simple style.

[0174] Here, the target display style is used as an example to illustrate the style of the target image to be generated. It can be understood that the selected target display style can be the style of both the target image and the target text. After determining the display style of the target image, the above process can be repeated for the target text to determine the style corresponding to the target text alone. In other words, the styles of the target image and the target file to be generated can be the same or different.

[0175] Step 702: The terminal sends the target display style to the server.

[0176] After the terminal sends the target display style to the server, the server can receive the target display style from the terminal. In this embodiment, the target image and the target text are generated by the server, so after the terminal determines the target display style, the target display style can be sent to the server, so that the server can determine whether the user has turned on the customized lock screen function, and analyze and determine the target display style selected by the user based on the received information, so as to subsequently generate the target image and / or target text with the target display style.

[0177] In addition, in addition to executing step 702 to send the target display style to the server, the terminal can also send the user's usage data to the server. That is, when the user uses the terminal's multimedia application, the terminal often needs to send the user's usage data to the server. For example, for a video application, when a user uses a video application to watch a certain video, the video application can send usage data to the server to indicate that the user watched the video in this time period based on the user's usage history, that is, the usage data may include the time the user watched the video and the identifier of the video watched. In addition, the usage data may also include the user's viewing progress, the content and time of the user sending barrages and comments, and so on.

[0178] Step 701 and step 702 are a process in which a user turns on a customized lock screen by selecting a display style. Through user-defined display styles, a personalized customized lock screen can be generated according to user needs. It is easy to understand that the target display style can increase the personalization and diversification of the lock screen, and steps 701 and 702 can also be defaulted. For example, in some embodiments, the target display style may not be determined, and a customized lock screen is generated only according to the first data and the second data below; or in other embodiments, a target image and a target text are generated according to the first data, the second data, and the default display style. For example, after the user turns on the customized lock screen function, if the display style is not selected, the display style corresponding to the user is determined to be the default display style; for another example, the terminal may turn on the customized lock screen function by default, and use the default display style to generate the target image and the target text by default.

[0179] In addition, in addition to the methods of step 701 and step 702, the target display style can also be determined by other methods. For example, the target display style can also be the display style with the highest degree of matching with the user among the preset display styles determined by the terminal based on the multimedia content browsed by the user, and the target display style is sent to the server through step 702. It can also be the display style with the highest degree of matching with the user among the preset display styles determined by the server based on the multimedia content browsed by the user. That is, the target display style is the display style with the highest degree of matching with the user among the preset display styles determined based on the multimedia content browsed by the user, and the subject of determining the target display style here can be the terminal or the server.

[0180] A specific determination method may be to input the names of multimedia data browsed by the user into a classification model, and determine the display style with the highest matching degree with the user through the classification model, or to cluster the feature vectors of the videos browsed by the user, and determine the display style with the highest matching degree with the feature vector of the cluster center as the target display style. The above two specific implementation methods are only examples and do not limit the present application.

[0181] After executing step 701 and step 702 or the other steps for determining the target display style, the server can determine that the customized lock screen function is enabled on the terminal and needs to provide the target image and target text to the terminal. Then, steps 703-705 can be executed to generate and send the target image and target text to the terminal.

[0182] It should be noted that, in some embodiments, steps 701-702 can be performed only once, and the subsequent steps 703-step 706 can be performed multiple times, that is, the user can only turn on the customized lock screen function once, and the server can generate multiple target images and / or target texts that meet the target display style, and each time the user turns on the lock screen application, different target images and target texts can be displayed, increasing the diversity of the lock screen interface. In other embodiments, steps 701-706 can be performed sequentially, and each step is only performed once. In other embodiments, steps 701-702 can also be performed multiple times. After steps 701-step 702 are performed multiple times, steps 703-step 706 are continued to be performed. In this way, by performing steps 701-702 multiple times, the user can set a variety of different target display styles. When steps 703-step 706 are performed multiple times in the subsequent multiple times, target images and target texts with multiple display styles can be generated, thereby increasing the diversity of the lock screen interface.

[0183] Step 703: The server obtains the first data and the second data of the terminal.

[0184] Step 704: The server generates a target image and a target text according to the first data and the second data.

[0185] Among them, the target image is at least related to the first data, and the target text is at least related to the second data; the first data is used to indicate the multimedia content browsed by the user corresponding to the terminal using the multimedia application, and the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, and the media data includes at least one of the text, audio, video and image of the multimedia content; the second data is used to indicate the usage data when the user uses the multimedia application to browse the multimedia content, and the specific meaning of the usage data has been explained above and will not be repeated here.

[0186] That is, the generated target image is related to the multimedia content browsed by the user, and the generated target text is related to the user's usage data. The specific description of the two can be found above and will not be repeated here. In some embodiments, the generated target image and target text may also be related to the target display style of steps 701-702. Here, taking the generation of target text and target image with target display style by artificial intelligence generated content (AI generated content, AIGC) in the lock screen scenario as an example, steps 703 and 704 will be described in detail.

[0187] Specifically, step 704 may include: obtaining first data and second data; using the first data, second data and target display style as input, and generating a target image and target text through a target model; the target model has the function of generating corresponding images and texts based on the input multimedia content and display style browsed by the user.

[0188] As described above, the terminal will send the user's usage data to the server, and the server can store the usage data after receiving the user's usage data. In addition, the server generally stores multimedia data of multimedia applications, such as a multimedia database of multimedia applications. Acquiring the first data and the second data here can be acquiring the second data from the user's usage data stored in the server, determining the identifier of the multimedia content browsed by the user, and acquiring the first data matching the determined identifier of the multimedia content from the multimedia database of the multimedia application stored in the server.

[0189] The target model may specifically be a model that inputs the first data, the second data and the target display style, and outputs a target image including the target text. That is, the output image includes the text, and the text content does not block the main part of the image, for example Fig. 9 As shown, the target image including the target text is 901, the target text is 902, the main part of the image is "person", and the target text does not cover the "person".

[0190] In addition, the font and color of the text can also be consistent with the display style of the target image or the overall display effect of the image. For example, if the main part of the image is red, the text can also be in red. For example, if the target image is in traditional Chinese style, the text can use fonts such as cursive script and regular script.

[0191] The target model may further include two models, namely, an image generation model and a text generation model. Then step 704 specifically includes: taking the first data and the target display style as input, generating a target image through the image generation model; taking the second data as input, generating a target text through the text generation model.

[0192] The image generation model here can be a generative adversarial network (GAN) or a diffusion model (DM). The text generation model can be a generative pre-trained transformer network (GPT) or other similar models.

[0193] In the embodiment of the present application, the pre-trained image generation model and text generation model can be used to directly generate the target image and target text. In addition, in the embodiment of the present application, the pre-trained image generation model can be used to obtain sample data marked with display style for the scene of the embodiment of the present application, and the image generation model can be further trained through the sample data, so that the image generation model can better learn the characteristics of different display styles, so that in step 704, according to the target display style, a target image that better conforms to the target display style can be generated.

[0194] In addition, before the target image and target text are generated by the target model, the feature vector corresponding to the first data (including metadata and medium data) of the multimedia content can be extracted, and the feature vector can be used as the input of the image generation model or the image generation model and the text generation model. Two descriptive words can also be generated, and the two descriptive words are used to prompt the requirements of the content output by the model. In this way, the model can better generate the target image and target text that meet the requirements according to the input.

[0195] In other words, in the case where the first data includes at least metadata, before step 704, it also includes: obtaining a feature vector of the first data; generating a first descriptive word based on the metadata and the target display style, and generating a second descriptive word based on the second data, the first descriptive word is used to describe the characteristics of the image to be generated, and the second descriptive word is used to describe the characteristics of the text to be generated.

[0196] In the above case, step 704 includes: taking the feature vector of the first data and the first description word as input, generating a target image through an image generation model; taking the second description word as input, generating a target text through a text generation model.

[0197] The above generation process can be found in Fig.10 Next, we will combine Fig.10 , taking the multimedia application as a video application and the multimedia content browsed by the user as the movie "Challenger" as an example, the above generation process is described in detail. Fig.10 The solid line box represents the specific model, and the dotted line box represents the input or output data.

[0198] For "Challenger", metadata may include basic information about the movie, such as the title, director, screenwriter, etc.; media data is the specific content of the movie, such as the movie's poster, trailer, plot introduction, complete video, complete audio, and complete subtitles.

[0199] First, a feature vector of the first data may be obtained. The obtained feature vector may include any one or more of an image feature vector, an audio feature vector, and a text feature vector. Feature vectors corresponding to different types of data may be extracted through different models. Next, the process of obtaining feature vectors will be specifically described by taking the image feature vector, audio feature vector, and text feature vector of "Challenger" as an example, and combining the first, second, and third steps below.

[0200] First, the video and / or image data of "Challenger" is input into a computer vision model to extract the image feature vector of the movie.

[0201] Among them, the computer vision model can specifically be a contrastive language-image pre-training (CLIP) model. The video of "Challenger" can be a complete video of the movie and / or a promotional video of the movie. The image of "Challenger" can include a poster image and / or a movie video screenshot. Among them, the movie video screenshot can be an official movie video screenshot released for promotion, or it can be a video frame with a large number of bullet comments or playbacks. In the case where the user's usage data includes data on the user's posting of bullet comments, the movie video screenshot can also be a screenshot corresponding to the video progress of the user posting the bullet comments, etc.

[0202] Second, the complete audio of "Challenger" is input into the audio signal processing model to obtain the audio feature vector of the movie.

[0203] Third, the text data of "Challenger" is input into the natural language processing model to obtain the text feature vector of the movie.

[0204] The text data here may include the name, director, screenwriter, plot introduction and complete subtitles of "Challenger". In one possible implementation, the text input into the natural language processing model may include the text in the medium data of "Challenger". In another possible implementation, the text input into the natural language processing model may also include metadata of "Challenger". In yet another possible implementation, the text input into the natural language processing model may include at least one of the metadata of "Challenger" and the text in the medium data.

[0205] Among them, the computer vision model in the first part, the audio signal processing model in the second part, and the natural language processing model in the third part together constitute Figure 3 The image-text-speech three-tower model mentioned in.

[0206] Any one or more of the above three steps may be omitted. Specifically, one or more of the above three steps may be selectively executed to complete the extraction of feature vectors for different types of multimedia content and for the characteristics of the multimedia content. For example, in this embodiment, the target image and target text are generated for the multimedia content "Challenger". For video application scenarios, there are abundant image and video data in this scenario, and the server can generate the target image only based on the image frames in the video. Therefore, in this embodiment, the second and third steps may also be omitted.

[0207] Fourth, a first descriptive word is generated according to the first data and the target display style, and a second descriptive word is generated according to the second data.

[0208] A specific method for generating the descriptive words may be through a natural language processing model. For example, in the present embodiment, metadata / media data of "Challenger", target display style and user usage data of "Challenger" may be input into a natural language processing model to output a first descriptive word and a second descriptive word.

[0209] Specifically, the first descriptor is used to control the generation of the target image, and is generated according to the target display style and the metadata / media data of "Challenger". Next, we will take the generation of the first descriptor according to the target display style and the title of "Challenger" as an example. When the target display style is "simple", the first descriptor can be "generate a simple-style "Challenger" image". The target display style and the title of "Challenger" are descriptions of the target image, and the first descriptor is a descriptor that describes the characteristics of the target image to be generated. Then, the image generation model can generate an image related to "Challenger" with the target display style according to the first descriptor.

[0210] In addition, in addition to generating the first description word according to the target display style and the title of the multimedia content, the description word may also be generated according to other data, for example, the description word may also be generated according to the target display style, the title of the multimedia content, and the introduction of the multimedia content, and "generate a simple style image of "Challenger" describing a person challenging the limit" may be generated. The above examples are only given for ease of understanding and do not limit the embodiments of the present application.

[0211] The second description word is used to control the generation of the target text, and the second description word is generated at least based on the user's usage data of "Challenger". The user's usage data may include the identification of the content browsed by the user and the description of the specific usage behavior. For example, the user's usage data may include the time when the user watched "Challenger", the progress of watching "Challenger", etc. For example, the second description word may specifically be "generate a lock screen description to remind the user that the movie in the picture was watched on XX month XX day".

[0212] In addition to the description of the specific usage behavior, the second description word may also include the title of the multimedia content browsed by the user. That is, here, the specific usage behavior of the user and the title of the multimedia content browsed by the user can be determined according to the user's usage data, and the second description word is generated based on the two. For example, the second description word may be "generate a lock screen description to remind the user that he watched "The Challenger" on XX month XX day".

[0213] In addition, for the natural language processing model, during the training phase, it is also possible to obtain some samples including descriptive words for the natural language processing model that has completed pre-training, and further train the natural language processing model for the usage scenarios in the embodiments of the present application, so that the natural language processing model can generate descriptive words that meet the requirements of the embodiments of the present application.

[0214] In the above example, the first descriptive word and the second descriptive word are generated by a natural language processing model, which can increase the diversity of the generated descriptive words. It is easy to understand that the first descriptive word and the second descriptive word may not be generated according to the artificial intelligence model, for example, they may be generated according to a template. For example, the template of the first descriptive word may be "generate an image of Y in style X", where X can be the label of the target display style, such as "simple", "secondary" or "national style", etc., and Y is the title of the video. The second descriptive word is similar to the first descriptive word, and will not be repeated here.

[0215] Fifth, obtain a unified eigenvector.

[0216] In some embodiments, at least two of the first, second and third steps may be performed. In this case, at least two feature vectors among the image feature vector, audio feature vector and text feature vector can be output. Then, the two or three feature vectors output by the above process are input into the information fusion module to align and fuse the features of different modalities and generate a unified feature vector.

[0217] Sixth, the feature vector and the first description word are input into the image generation model to obtain the target image; the second description word is input into the text generation model to obtain the target text.

[0218] Among them, the feature vector of the input image generation model can be a unified feature vector obtained in the fifth step. When only any one of the first, second and third steps is executed, the feature vector can be any one of an image feature vector, an audio feature vector and a text feature vector.

[0219] Still taking the target display style as "simple" and the multimedia content as "Challenger" as an example, the above steps can output the following: Fig.11The image of The Challenger in a simple style as shown in (a) in the figure. When the usage data is the time when the user watched The Challenger, the following target text can be generated: "You watched The Challenger on XX / XX / XX, come and reminisce." When the usage data is the progress of the user watching The Challenger, the following target text can be generated: "You have watched 50% of The Challenger in the video app, come and continue watching." The above target text can enhance Fig.11 The interpretability of the image shown in (a) enables users to know the image content and the purpose of pushing the image, improving the user experience. At the same time, through the target text, users can use the video application more and increase the traffic for the video application.

[0220] The above examples are explained by taking video applications as an example. It is easy to understand that other multimedia application processing methods are similar to the above process, except that the media data input in the first, second and third steps are different.

[0221] For example, in a music application, in the first step, you may input the MV of a song and / or the album cover of a song, in the second step, you may input the audio content of the song, and in the third step, you may input the lyrics of the song. In a reading application, in the first step, you may input the cover of a book and / or illustrations in the book, and in the third step, you may input the text and / or a brief introduction of the book. In a game application, in the first step, you may input a screenshot and / or promotional video of a game, in the second step, you may input the audio of the game promotional video and / or the game soundtrack, and in the third step, you may input the introduction of the game. The above examples are merely illustrative and do not constitute limitations on this application.

[0222] In addition, in the above scenarios, media data may include four types of data: images, videos, text, and audio. In some scenarios, media data may only include one or two of the following: images, videos, text, and audio. For example, in a reading application, the book you are reading may only have text content, so media data may only include text data. In another example, in a music application, a song may only have audio and lyrics, so media data may only include audio data and text data.

[0223] In this case, the process of extracting some feature vectors can be omitted, and the vectors input to the image generation model also need to be adjusted accordingly. For example, in the scenario of reading applications, when the medium data only includes text data (i.e., the text of the book being read), the first, second, and fifth steps above can be omitted, and the image generation model inputs the text feature vector. For example, in the scenario of music applications, when the medium data includes text and audio, the first step can be omitted, and the image generation model inputs the fusion vector of the text feature vector and the audio feature vector.

[0224] In addition, even if the multimedia application includes four types of data, namely, images, videos, texts and audios, in the first, second and third steps, feature vectors can be extracted based on only any one, two or three of the four types of data, and the target image can be further generated based on the extracted feature vectors.

[0225] The above example is explained by taking the target image generating the target display style as an example. If you need to generate a target text with the target display style, the generation method is similar to the method in the previous article. For example, you can first perform any one or more of the first, second and third steps. In the fourth step, generate a second descriptive word based on the second data and the target display style. The target display style used to generate the second descriptive word here can be the same as the target display style used to generate the first descriptive word, or it can be different from the target display style used to generate the first descriptive word. Then continue to execute the subsequent fifth step. In the sixth step, the feature vector and the second descriptive word can be input into the text generation model to obtain the target text with the target display style.

[0226] Step 705: The server sends the target image and target text to the terminal.

[0227] Correspondingly, the terminal can receive the target image and target text from the server.

[0228] As for the execution timing of step 703, step 704 and step 705, step 703 and step 704 can be executed periodically, that is, the target image and target text are generated periodically using the user's historical usage data and the data of the multimedia content corresponding to the historical usage data. That is, the target image and target text are generated based on the multimedia content that the user has browsed in the past. And an image library is maintained for each user, and the image library stores at least two target images generated by the above method and their corresponding target texts. Step 705 can also be executed periodically, that is, the server periodically pushes the target image and target text to the terminal through step 705. Specifically, the server can periodically randomly select an image from the above image library and send it to the terminal. Among them, the execution cycles of step 704 and step 705 can be different or the same.

[0229] In addition, step 703 and step 704 can also be executed when new usage data is detected, that is, when a user uses a certain multimedia application, a target image and target text corresponding to the multimedia content browsed by the user using the multimedia application can be generated in real time. For example, if the user is watching "The Challenger", a target image of "The Challenger" can be generated (such as Fig. 9 ), and generates a target text for reminding the user to continue watching "The Challenger".

[0230] Step 706: The terminal receives a first operation of the user to open a lock screen application; in response to the first operation, a target image and a target text are displayed.

[0231] That is, in the lock screen scene, when the user opens the lock screen interface, the target image and the target text are displayed. The description of the first operation is detailed in the previous text and will not be repeated here.

[0232] In some embodiments, step 705 may be executed regularly. After the server sends the target image and target text to the terminal in step 705, the terminal may first cache the received target image and target text in the terminal. When the user opens the lock screen, desktop, screen off or system application, the target image and target text are obtained from the cached content for display. When there are multiple cached target images and target texts, a pair of target images and target texts may be selected arbitrarily for display.

[0233] In other embodiments, step 705 may also be executed after receiving a trigger from the terminal. For example, step 706 may be executed first, in which the terminal receives a first operation from the user to open a lock screen application, and in response to the first operation, step 705 is triggered to execute, and the target image and target text are further displayed in step 706.

[0234] It should be noted that the above two examples are only two examples for illustrating the execution order of step 705 and step 706 in the present application, and do not represent limitations on the embodiments of the present application.

[0235] As for the method for displaying the target image and the target text, the target image and the target text may be displayed simultaneously in response to the first operation. Alternatively, the target image may be displayed in the lock screen interface; and the target text may be displayed upon receiving the third operation of the user in the lock screen interface.

[0236] The third operation here can be a trigger operation on the control in the lock screen interface, for example, it can be for Fig.11 In (a), the third operation may be a click or swipe operation on the control 1101, or a swipe operation on the lock screen interface that is different from the unlock screen. In this way, the terminal displays the target text after receiving the third operation, which can increase the interpretability of the target image through the target text without blocking the main body of the target image, and can also enhance the interaction between the terminal and the user, thereby improving the user experience.

[0237] Taking the third operation as an upward swipe operation on the control 1101 as an example, the above process is described in detail. Fig.11 As shown, after the user opens the lock screen interface of the lock screen application through the first operation, the terminal can display Fig.11 The interface including the target image is shown in (a) of FIG. Fig.11 In the lock screen interface shown in (a) of FIG. 1 , a swipe operation is performed on the control 1101. After the terminal receives the operation, in response, the terminal can display an animation of the control 1101 rising from the bottom of the lock screen interface, and finally display the following Fig.11 The interface shown in (b) in FIG. Fig.11 In (b), the text included in the control 1101 is the target text.

[0238] like Fig.11 As shown in (b), in the interface displaying the target image and the target text, in addition to the control 1101, the lock screen interface may also have controls 1102 and 1103. The functions of these two controls will be described in detail below.

[0239] Control 1102 is also called a jump control, also referred to as a second control in this application. In the embodiment of this application, through the user's triggering operation on the jump control, the user can jump from the lock screen interface to the interface of the multimedia content corresponding to the image on the lock screen interface.

[0240] In addition, when the user sets an unlocking password or other unlocking verification, the user identity needs to be verified before jumping to the interface of the multimedia content. The method of verifying the user identity can refer to the method of the related art, which will not be repeated here.

[0241] In other words, the method of the embodiment of the present application also includes: displaying a second control; receiving a fourth operation of the user on the second control; and displaying multimedia content corresponding to the target image in the second application in response to the fourth operation.

[0242] The fourth operation is the user's Fig.11 The trigger operation of the control 1102 shown in (b) of FIG. 1102, for example, the fourth operation may be an operation in which the user clicks the control 1102. Displaying the multimedia content in the second application, that is, jumping to an interface where the multimedia content can be browsed, such as Fig.12 The interface shown.

[0243] Next, the above process will be described by taking the fourth operation, that is, the user clicking the control 1102, as an example. Fig.11 In the case of the image shown in (b) in FIG. 1 , by the user clicking on the control 1102, the terminal can jump to the Fig.12 The playback interface of "Challenger" is shown.

[0244] The control 1103 is also called the delete control, also referred to as the first control in this application. The delete control is used to stop displaying the current target image and target text after receiving a trigger.

[0245] Specifically, the method of the embodiment of the present application further includes: displaying the first control; receiving a second operation of the user on the first control; and sending the first information to the server in response to the second operation.

[0246] The first information is used to instruct the server to delete the target image and target text in the image library corresponding to the user; the image library stores multiple images and texts generated based on the usage data of the user browsing multimedia content and the multimedia content browsed by the user.

[0247] The second operation may be a user triggering operation on the control 1103, such as a user clicking the control 1103. For example, when the terminal displays Fig.11 After the content shown in (b) in the figure is displayed, the user can click on the control 1103. After receiving the user's click operation on the control 1103, the terminal can generate first information and send the first information to the server. The server receives the first information from the terminal and deletes the corresponding image and text from the image library in response to the first information. The image library is the image library composed of multiple target images and target texts generated in step 705 above.

[0248] In the above embodiment, the lock screen application is used as an example for explanation. In addition, the first application may also be a desktop application, a screen-off application, or a system application, etc. When the first application is in other applications, the overall implementation method is similar to steps 701-706, except that the specific implementation method of step 706 is different.

[0249] Specifically, the display of the target image and target text in step 706 includes: the first application is a desktop application, at least the target image and target text are displayed as desktop wallpaper on the desktop. Or, the first application is an off-screen application, and the target image and target text are displayed as widgets on the off-screen interface. For example, Figure 2 The widget 201 is shown replaced with a target image and target text.

[0250] In addition, when the first application is any application in the desktop application, lock screen application or screen-off application, the terminal theme can also be generated according to the target image and target text, and the target text and target image can be displayed through the theme in the three applications. Among them, the theme is a personalized design of the terminal skin interface, and the theme can be designed and developed by a company or an individual developer. The content that can be personalized and replaced by the target image and target text in the theme can specifically include wallpapers, icons, and backgrounds in the interface covered by the theme. The interfaces covered by the theme may include notification bars, text messages, dialing, contacts, and settings. The wallpaper or background can be replaced by the target image and target text, such as displaying it in the background of the wallpaper or interface. Fig. 9The icon can be replaced by the target image or the subject content in the target image. For example, the application icon on the desktop can be as follows Fig. 9 As shown in the target image 901, the application icon in the desktop can also be the main content of the target image, for example, it can be as follows Fig. 9 The person shown.

[0251] Or, the first application is a system application, and the target image and target text are used as the startup image and displayed on the startup interface of the first application. The application generally needs to be loaded during the startup process, and the application logo or advertising image, etc. will be displayed on the startup interface during the loading process. In this application, the above target image and target text can also be displayed on the startup interface of the system application. In addition, if there is cooperation between the developer of the third-party application and the terminal manufacturer, or the terminal operating system has the authority to control the startup interface of the third-party application, the above target image and target text can also be displayed on the third-party application.

[0252] In the above embodiment, the implementation method of the embodiment of the present application is described in detail by taking the interaction between the terminal and the server as an example. As described above, the embodiment of the present application can also be executed by the terminal alone. When it is executed by the terminal alone, the specific implementation method is similar to the above steps 701-706. Unlike steps 701-706, when the terminal executes the above method alone, steps 702 and 705 may not be executed, and steps 703 and 704 are executed by the terminal.

[0253] In addition, when the terminal is executed alone, the target display style may be determined in step 701 or may be a display style that has the highest matching degree with the user among preset display styles determined based on the multimedia content browsed by the user.

[0254] In addition, in the case where the above-mentioned deletion control exists, after receiving the second operation, the terminal can delete the target image and target text in the image library corresponding to the user in response to the second operation; the image library stores multiple images and texts generated based on the usage data when the user browses the multimedia content and the multimedia content browsed by the user. That is, the terminal stores the above-mentioned image library, and the terminal deletes the corresponding image and text from the image library in response to the second operation.

[0255] Some other embodiments of the present application also provide a device having the function of implementing the behavior of the electronic device (such as a terminal or a server) in the above-mentioned embodiment. The function can be implemented by hardware, or it can be implemented by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions, for example, a processing module, a sending module, and a display module. As an example, the processing module can execute the above-mentioned step 701 or step 703; the sending module can execute the above-mentioned step 702 or step 705; the display module can execute the above-mentioned step 706.

[0256] The present application also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above method.

[0257] The present application also provides a server, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above method.

[0258] The present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to implement the method.

[0259] The present application also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0260] The present application also provides a chip, which includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory to execute the above method.

[0261] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An image display method, characterized in that: Applied to a terminal, the method comprises: Receiving a first operation of a user opening a first application; In response to the first operation, displaying a target image and a target text; Among them, the target image is at least related to first data, the first data is used to indicate the multimedia content browsed by the user using the second application, the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data includes at least one of the text, audio, video and image of the multimedia content; the target text is at least related to second data, the second data is used to indicate usage data of the user when browsing the multimedia content using the second application.

2. The method according to claim 1, characterized in that The target image is further related to a target display style, and / or the target text is further related to the target display style.

3. The method according to claim 2, characterized in that The target display style is a display style selected by the user from a plurality of displayed candidate display styles; or, the target display style is a display style that has the highest matching degree with the user among preset display styles determined based on multimedia content browsed by the user.

4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Display a second control; receiving a fourth operation of the user on the second control; In response to the fourth operation, the multimedia content corresponding to the target image is displayed in the second application.

5. The method according to any one of claims 2 to 4, characterized in that: The method further comprises: The target display style is sent to the server.

6. The method according to any one of claims 1 to 5, characterized in that: Before displaying the target image and the target text, the method further includes: The target image and the target text are received from the server.

7. The method according to claim 5 or 6, characterized in that: The method further comprises: Displaying a first control; receiving a second operation of the user on the first control; In response to the second operation, first information is sent to the server, wherein the first information is used to instruct the server to delete the target image and the target text in an image library corresponding to the user, wherein the image library stores multiple images and texts generated based on usage data of the user when browsing multimedia content and the multimedia content browsed by the user.

8. The method according to any one of claims 2 to 4, characterized in that: The method further comprises: Acquire the first data and the second data; The first data, the second data and the target display style are used as input, and the target image and the target text are generated through a target model, wherein the target model has the function of generating corresponding images and texts based on the multimedia content and display style browsed by the input user.

9. The method according to claim 8, characterized in that The target model includes an image generation model and a text generation model; the step of taking the first data, the second data and the target display style as input and generating the target image and the target text through the target model includes: Taking the first data and the target display style as input, generating the target image through the image generation model; The second data is taken as input and the target text is generated through the text generation model.

10. The method according to claim 9, characterized in that In the case where the first data at least includes the metadata, the method further includes: Obtaining a feature vector of the first data; Generate a first description word according to the metadata and the target display style, and generate a second description word according to the second data, wherein the first description word is used to describe the characteristics of the image to be generated, and the second description word is used to describe the characteristics of the text to be generated; The step of taking the first data and the target display style as input and generating the target image through the image generation model includes: Taking the feature vector of the first data and the first description word as input, generating the target image through the image generation model; The step of taking the second data as input and generating the target text through the text generation model includes: The second description word is used as input, and the target text is generated through the text generation model.

11. The method according to any one of claims 8 to 10, characterized in that: The method further comprises: Display a first control; receiving a second operation of the user on the first control; In response to the second operation, the target image and the target text are deleted in an image library corresponding to the user; the image library stores a plurality of images and texts generated based on usage data of the user when browsing multimedia content and the multimedia content browsed by the user.

12. The method according to any one of claims 1 to 11, characterized in that: The first application is a lock screen application; and the displaying of the target image and the target text includes: Displaying the target image on the lock screen interface; A third operation of the user in the lock screen interface is received, and the target text is displayed.

13. The method according to any one of claims 1 to 11, characterized in that: The displaying of the target image and the target text comprises: The first application is a desktop application, and the target image and the target text are used as desktop wallpaper and displayed on the desktop; or, The first application is a screen-off application, and the target image and the target text are displayed as widgets on a screen-off interface; or, The first application is a system application, and the target image and the target text are used as startup images and displayed on a startup interface of the first application.

14. An image generation method, characterized in that: Applied to a server, the method comprises: Acquire first data and second data of a terminal, wherein the first data is used to indicate multimedia content browsed by a user corresponding to the terminal using a second application, wherein the first data includes at least one of the following data: metadata of the multimedia content, media data of the multimedia content, the media data including at least one of text, audio, video and image of the multimedia content; and the second data is used to indicate usage data when the user uses the second application to browse the multimedia content; Generate a target image and a target text according to the first data and the second data; wherein the target image is at least related to the first data, and the target text is at least related to the second data; The target image and the target text are sent to the terminal.

15. The method according to claim 14, characterized in that The target image is further related to a target display style, and / or the target text is further related to the target display style.

16. The method according to claim 15, characterized in that The method further comprises: The target display style sent by the terminal is received, where the target display style is a display style selected by the user from a plurality of candidate display styles displayed by the terminal.

17. The method according to claim 15, characterized in that The target display style is a display style that has the highest matching degree with the user among preset display styles determined according to the multimedia content browsed by the user.

18. The method according to any one of claims 15 to 17, characterized in that: The step of generating the target image and the target text according to the first data and the second data includes: The first data, the second data and the target display style are used as input, and the target image and the target text are generated through a target model, wherein the target model has the function of generating corresponding images and texts based on the multimedia content and display style browsed by the input user.

19. The method according to claim 18, characterized in that The target model includes an image generation model and a text generation model; the step of taking the first data, the second data and the target display style as input and generating the target image and the target text through the target model includes: Taking the first data and the target display style as input, generating the target image through the image generation model; The second data is taken as input and the target text is generated through the text generation model.

20. The method according to claim 19, characterized in that In the case where the first data at least includes the metadata, the method further includes: Obtaining a feature vector of the first data; Generate a first description word according to the metadata and the target display style, and generate a second description word according to the second data, wherein the first description word is used to describe the characteristics of the image to be generated, and the second description word is used to describe the characteristics of the text to be generated; The step of taking the first data and the target display style as input and generating the target image through the image generation model includes: Taking the feature vector of the first data and the first description word as input, generating the target image through the image generation model; The step of taking the second data as input and generating the target text through the text generation model includes: The second description word is used as input, and the target text is generated through the text generation model.

21. The method according to any one of claims 14 to 20, characterized in that: The method further comprises: receiving first information from the terminal; In response to the first information, the target image and the target text are deleted in an image library corresponding to the user; the image library stores a plurality of images and texts generated based on the usage data of the user when browsing multimedia content and the multimedia content browsed by the user.

22. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the method according to any one of claims 1 to 13.

23. A server, characterized in that: include: a memory for storing instructions executed by one or more processors of the server; A processor, when the processor executes the instructions in the memory, can cause the server to execute the method described in any one of claims 14 to 21.

24. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the method according to any one of claims 1 to 21.

25. A chip, characterized in that: The chip includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory to execute the method according to any one of claims 1-21.

Citation Information

Cited By

  • Image display method, image generation method, and electronic device

    EP4742031A1

  • Image display method, image generation method, and electronic device

    WO2025102916A1