Cover picture generation method, system, equipment and medium
By connecting to the digital assistant to generate candidate cover images of multimedia content, users' diverse needs for multimedia content cover images are solved, and more intuitive cover image display is achieved, improving user experience.
Patent Information
- Application Number
- CN202410232611.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art is difficult to meet users' diverse needs for multimedia content cover images. Traditional methods such as using video frames or uploading images by users themselves are time-consuming and labor-intensive, making it difficult to intuitively distinguish different multimedia content.
By accessing a digital assistant, using its image generation ability, a candidate cover image of multimedia content is generated based on the description information, including receiving description items or reference content input by the user, and calling the image generation model to generate a candidate cover image.
The generated candidate cover image meets the diverse needs of users, can reflect multimedia content more intuitively, and improve user viewing experience.
Smart Images

Figure CN120561320A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a cover image generation method, system, electronic device, and computer-readable storage medium. Background Art
[0002] Multimedia is a combination of multiple media. Generally, multimedia content includes text content, sound content, image content, video content and other media content.
[0003] With the development of computer technology, computer systems produce and store a large amount of multimedia content. For example, in video conferencing scenarios, video conferencing systems often include a recording function to facilitate viewing of conference content by participants or other users. When a participant enables recording, the system records the conference and generates a recording file after the conference, which can be viewed by participants or other users. This recording file is multimedia content.
[0004] Users can gain a preliminary understanding of multimedia content through its cover image. For example, in a video conferencing scenario, conference recordings can be provided to users in a list format, displaying the names and cover images of the conference recordings. Video conferencing systems typically use a specific video frame from the conference as the cover image for the conference recording.
[0005] However, the above method is difficult to meet the diverse needs of users for cover images of multimedia content. Summary of the Invention
[0006] This application provides a method for generating a cover image. This method can generate candidate cover images that meet user needs, improving the user's experience of viewing multimedia content. This application also provides a system, electronic device, computer-readable storage medium, and computer program product corresponding to the above method.
[0007] In a first aspect, the present application provides a method for generating a cover image, the method comprising:
[0008] Determining descriptive information for a cover image of the first multimedia content;
[0009] Present at least one candidate cover image for the first multimedia content, where the at least one candidate cover image is generated by a digital assistant based on the description information.
[0010] In some possible implementations, the description information includes at least one description item related to the cover image, or the description information includes reference content indicating generation of the cover image.
[0011] In some possible implementations, the first multimedia content includes video content and text content, the text content is associated with the video content, and the reference content for generating the cover image includes the text content.
[0012] In some possible implementations, determining descriptive information for the cover image of the first multimedia content includes:
[0013] Displaying a dialogue page, wherein the dialogue page is used to interact with the digital assistant;
[0014] Receiving a target operation triggered in the conversation page;
[0015] Describe information of a cover image for the first multimedia content according to the target operation.
[0016] In some possible implementations, receiving a target operation triggered in the conversation page includes:
[0017] Receiving an information input operation triggered in the conversation page;
[0018] The determining, according to the target operation, descriptive information of the cover image for the first multimedia content includes:
[0019] Acquiring first information representing natural language content according to the information input operation;
[0020] Determine the first information as description information of a cover image of the first multimedia content; or,
[0021] Based on the first information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
[0022] In some possible implementations, receiving a target operation triggered in the conversation page includes:
[0023] Receiving a control triggering operation in the dialog page;
[0024] The determining, according to the target operation, descriptive information of the cover image for the first multimedia content includes:
[0025] According to the control triggering operation, obtaining second information indicated by the target control;
[0026] Determine the second information as description information of the cover image of the first multimedia content; or,
[0027] Based on the second information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
[0028] In some possible implementations, presenting the dialogue page includes:
[0029] In response to a triggering operation on a first control, displaying a dialog page, wherein the first control is used to trigger generation of a cover image for the first multimedia content; or
[0030] In response to a triggering operation on the digital assistant, a conversation page is displayed.
[0031] In some possible implementations, presenting the dialogue page includes:
[0032] displaying the conversation page in parallel with the display page of the first multimedia content; or,
[0033] In the display page of the first multimedia content, the conversation page is displayed in the form of a floating window component.
[0034] In some possible implementations, after receiving the target operation triggered in the conversation page, the method further includes:
[0035] Determining, according to the target operation, an operation intention indicated by the target operation;
[0036] In response to the operation intention including generating a cover image, a first processing tool is called to enable the first processing tool to generate at least one candidate cover image.
[0037] In some possible implementations, the first processing tool generates at least one candidate cover image, including:
[0038] The first processing tool sends the description information and the third information representing the cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model.
[0039] In some possible implementations, the description information includes reference content indicating generation of a cover image, the reference content indicating generation of the cover image includes text content in the first multimedia content, the first processing tool sends the description information and third information representing cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model, including:
[0040] Determining identification information of video content in the first multimedia content according to the description information;
[0041] Acquiring text content associated with the video content according to the identification information of the video content;
[0042] The description information, the text content and the third information representing the cover image generation requirements are sent to the image generation model, and at least one candidate cover image returned by the image generation model is received.
[0043] In some possible implementations, the method further includes:
[0044] In response to the fact that the operation intention does not include generating a cover image, a second processing tool is called to enable the second processing tool to perform natural language processing on the operation intention.
[0045] In some possible implementations, presenting at least one candidate cover image for the first multimedia content includes:
[0046] In the conversation page, at least one candidate cover image of the first multimedia content is presented.
[0047] In some possible implementations, the method further includes:
[0048] In response to a selection operation on a target cover image from the at least one candidate cover image, the target cover image is determined as the cover image of the first multimedia content.
[0049] In some possible implementations, the method further includes:
[0050] On a page associated with the first multimedia content, a cover image of the first multimedia content is presented.
[0051] In some possible implementations, the page associated with the first multimedia content includes at least one of the following: a viewing page of historical multimedia content, an instant messaging page after sharing the first multimedia content, and a preview page of the first multimedia content.
[0052] In a second aspect, the present application provides a cover image generation system, the system comprising:
[0053] A determination module, configured to determine description information of a cover image for the first multimedia content;
[0054] A presentation module is used to present at least one candidate cover image of the first multimedia content, where the at least one candidate cover image is generated by a digital assistant based on the description information.
[0055] In some possible implementations, the description information includes at least one description item related to the cover image, or the description information includes reference content indicating generation of the cover image.
[0056] In some possible implementations, the first multimedia content includes video content and text content, the text content is associated with the video content, and the reference content for generating the cover image includes the text content.
[0057] In some possible implementations, the determining module is specifically configured to:
[0058] Displaying a dialogue page, wherein the dialogue page is used to interact with the digital assistant;
[0059] Receiving a target operation triggered in the conversation page;
[0060] Describe information of a cover image for the first multimedia content according to the target operation.
[0061] In some possible implementations, the determining module is specifically configured to:
[0062] Receiving an information input operation triggered in the conversation page;
[0063] Acquiring first information representing natural language content according to the information input operation;
[0064] Determine the first information as description information of a cover image of the first multimedia content; or,
[0065] Based on the first information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
[0066] In some possible implementations, the determining module is specifically configured to:
[0067] Receiving a control triggering operation in the dialog page;
[0068] According to the control triggering operation, obtaining second information indicated by the target control;
[0069] Determine the second information as description information of the cover image of the first multimedia content; or,
[0070] Based on the second information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
[0071] In some possible implementations, the determining module is specifically configured to:
[0072] In response to a triggering operation on a first control, displaying a dialog page, wherein the first control is used to trigger generation of a cover image for the first multimedia content; or
[0073] In response to a triggering operation on the digital assistant, a conversation page is displayed.
[0074] In some possible implementations, the determining module is specifically configured to:
[0075] displaying the conversation page in parallel with the display page of the first multimedia content; or,
[0076] In the display page of the first multimedia content, the conversation page is displayed in the form of a floating window component.
[0077] In some possible implementations, the system further includes a calling module, wherein the calling module is configured to:
[0078] Determining, according to the target operation, an operation intention indicated by the target operation;
[0079] In response to the operation intention including generating a cover image, a first processing tool is called to enable the first processing tool to generate at least one candidate cover image.
[0080] In some possible implementations, the first processing tool generates at least one candidate cover image, including:
[0081] The first processing tool sends the description information and the third information representing the cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model.
[0082] In some possible implementations, the description information includes reference content indicating generation of a cover image, the reference content indicating generation of the cover image includes text content in the first multimedia content, the first processing tool sends the description information and third information representing cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model, including:
[0083] Determining identification information of video content in the first multimedia content according to the description information;
[0084] Acquiring text content associated with the video content according to the identification information of the video content;
[0085] The description information, the text content and the third information representing the cover image generation requirements are sent to the image generation model, and at least one candidate cover image returned by the image generation model is received.
[0086] In some possible implementations, the calling module is further configured to:
[0087] In response to the fact that the operation intention does not include generating a cover image, a second processing tool is called to enable the second processing tool to perform natural language processing on the operation intention.
[0088] In some possible implementations, the presentation module is specifically configured to:
[0089] In the conversation page, at least one candidate cover image of the first multimedia content is presented.
[0090] In some possible implementations, the determining module is further configured to:
[0091] In response to a selection operation on a target cover image from the at least one candidate cover image, the target cover image is determined as the cover image of the first multimedia content.
[0092] In some possible implementations, the presentation module is further configured to:
[0093] On a page associated with the first multimedia content, a cover image of the first multimedia content is presented.
[0094] In some possible implementations, the page associated with the first multimedia content includes at least one of the following: a viewing page of historical multimedia content, an instant messaging page after sharing the first multimedia content, and a preview page of the first multimedia content.
[0095] In a third aspect, the present application provides an electronic device comprising a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory to cause the electronic device to perform the cover image generation method according to the first aspect or any implementation of the first aspect.
[0096] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct an electronic device to execute the cover image generation method described in the above-mentioned first aspect or any implementation method of the first aspect.
[0097] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when run on an electronic device, enables the electronic device to execute the cover image generation method described in the first aspect or any one of the implementations of the first aspect.
[0098] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods.
[0099] It can be seen from the above technical solutions that this application has the following advantages:
[0100] The present application provides a cover image generation method, which first receives description information of a cover image for first multimedia content, and then presents at least one candidate cover image for the first multimedia content, wherein the at least one candidate cover image is generated by a digital assistant based on the description information.
[0101] This method addresses the generation of multimedia content cover images. By connecting to a digital assistant and leveraging its image-generating capabilities, the digital assistant is able to generate candidate cover images for the multimedia content based on the cover image's descriptive information. This method addresses the diverse user needs for multimedia content cover images by generating candidate cover images that meet those needs. Furthermore, the cover images more intuitively reflect the multimedia content, enhancing the user's viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] In order to more clearly illustrate the technical methods of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments.
[0103] Figure 1 A schematic diagram of a display page for a conference recording file provided in an embodiment of the present application;
[0104] Figure 2 A schematic diagram of a conference recording file list provided in an embodiment of the present application;
[0105] Figure 3 A schematic diagram of a flow chart of a cover image generation method provided in an embodiment of the present application;
[0106] Figures 4A to 4E A schematic diagram of a display page for a conference recording file provided in an embodiment of the present application;
[0107] Figure 5 A schematic structural diagram of a cover image generation system provided in an embodiment of the present application;
[0108] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0109] The terms "first" and "second" in the embodiments of this application are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.
[0110] First, some technical terms involved in the embodiments of this application are introduced.
[0111] Multimedia is a combination of multiple media. Generally, multimedia content includes text content, sound content, image content, video content and other media content.
[0112] In some examples, the multimedia content may be a recording of a video conference. A video conference is a meeting in which participants in two or more locations communicate via a communication device and a network. The communication format of a video conference is point-to-point, such as point-to-multipoint or multipoint-to-multipoint.
[0113] Typically, meeting participants can participate in a video conference through a video conferencing system. This system can also be referred to as video conferencing software or a video conferencing application (APP). With the continuous development of cloud computing, video conferencing systems can now perform data transmission, processing, and storage on cloud servers. Users only need to log in to the video conferencing system client to conduct efficient video conferencing, making the system more convenient to use.
[0114] Video conferencing systems typically include a recording feature to facilitate viewing of video conference content by conference participants or other users. Once a conference participant (e.g., a conference host with permission to enable recording) enables recording, the system will record the conference and generate a recording file after the conference, which can be viewed by conference participants or other users.
[0115] Specifically, the conference recording file may include multimedia content related to the video conference. Figure 1 The diagram shows a display page of a conference recording file. The conference recording file 10 includes a title bar 101, a video conference presentation area 102, a first information presentation area 103 and a second information presentation area 104.
[0116] Title bar 101 displays basic video conference attributes, such as the conference topic and start time. Furthermore, title bar 101 provides at least one control related to the conference recording, such as a control for sharing the conference recording. Title bar 101 allows users to quickly access basic conference recording attributes and trigger functions related to the conference recording.
[0117] The video conference presentation area 102 is used to present the video conference content. The video conference content can be the content of the meeting from the time the conference participants start recording to the time they finish recording (e.g., the end of the video conference). Users can view the video conference through the video conference presentation area 102. Furthermore, the video conference presentation area 102 can also provide functions related to viewing the video conference, such as play, pause, fast forward, rewind, double-speed playback, full-screen playback, and displaying subtitles, to meet the different viewing needs of users.
[0118] The first information presentation area 103 is used to present speaker information and conference information. Speaker information may include information about participants speaking in the video conference and the timeline of their speeches. Conference information may include information about the video conference, such as the creator and creation time of the video conference. Users can switch between viewing speaker information and conference information by clicking "Speaker" or "Conference Information."
[0119] The second information presentation area 104 is used to present meeting minutes and text records. Meeting minutes can be understood as a summary generated by summarizing the content of a video conference. For example, meeting minutes may include the topic of the video conference, the process of the video conference, and the content of the speakers' speeches during the video conference. Users can edit meeting minutes after the video conference ends. In some video conferencing systems, meeting minutes can also be automatically generated. Text records can be understood as records of speeches made by individual speakers during the video conference, such as speech records generated after voice recognition and classification of the video conference content. Users can switch between viewing the meeting minutes and text records of the video conference by clicking "Meeting Minutes" and "Text Records."
[0120] Considering that a user can participate in multiple video conferences using video conferencing software, the recording files of multiple video conferences can be provided to the user in the form of a list. Figure 2 The diagram shows a conference recording file list, where the conference recording file list 20 includes a header area 201 and a conference recording file presentation area 202 .
[0121] Header area 201 displays attributes related to meeting recordings, such as basic information about the recording, the owner of the recording, and the creation time of the recording. Users can change the "owner" of a recording to view recordings with permissions belonging to different owners. Users can also change the "creation time" to view recordings in a different order.
[0122] The conference recording file presentation area 202 is used to present the conference recording files in a list format. In the column where the "basic attribute information of the conference recording file" is located, the name and cover image of the conference recording file can be presented to facilitate users to quickly locate the desired conference recording file.
[0123] In the related art, video conferencing systems usually use a certain video frame (for example, the first video frame) in a video conference as the cover image of the conference recording file. In some video conferencing systems, users are supported to change the cover image of the conference recording file. Users can upload a local image to replace the cover image of the conference recording file with the local image. However, the above method is difficult to meet the diverse needs of users for cover images of multimedia content such as conference recording files. Using a certain video frame as the cover image of a conference recording file is difficult to intuitively distinguish different video conferences, and users uploading local images as the cover image of a conference recording file is time-consuming and labor-intensive.
[0124] In view of this, the present application provides a cover image generation method. The method first receives description information of a cover image for first multimedia content, and then presents at least one candidate cover image for the first multimedia content, wherein the at least one candidate cover image is generated by a digital assistant based on the description information.
[0125] This method addresses the generation of multimedia content cover images. By connecting to a digital assistant and leveraging its image-generating capabilities, the digital assistant is able to generate candidate cover images for the multimedia content based on the cover image's descriptive information. This method addresses the diverse user needs for multimedia content cover images by generating candidate cover images that meet those needs. Furthermore, the cover images more intuitively reflect the multimedia content, enhancing the user's viewing experience.
[0126] To facilitate understanding of the technical solutions provided in the embodiments of the present application, they will be described below with reference to the accompanying drawings.
[0127] See also Figure 3 The flowchart of a cover image generation method provided in an embodiment of the present application is shown, and the method specifically includes:
[0128] S301: Receive description information of a cover image for first multimedia content.
[0129] The first multimedia content can be understood as the multimedia content for which a cover image is to be generated. The first multimedia content can be composed of multiple media contents, such as video content, text content, image content, etc.
[0130] In some embodiments, the first multimedia content may be multimedia content created by the first user, such as a recording of a video conference created by the first user. In other embodiments, the first multimedia content may also be multimedia content that the first user has editing permissions for, such as a recording of a conference that the first user can view and edit, although this is not a limitation in the present application.
[0131] In some embodiments, the description information may include at least one descriptive item related to the cover image. The at least one descriptive item may describe the cover image from at least one dimension. For example, when describing the cover image from the perspective of style, the descriptive item may be oil painting style, sketch style, etc.; when describing the cover image from the perspective of theme, the descriptive item may be office, study, etc.; and when describing the cover image from the perspective of color, the descriptive item may be light color, dark color, etc.
[0132] In other embodiments, the description information may include reference content for generating the cover image. In other words, the description information may be the content upon which the cover image is generated. For example, the description information may be content A, in which case the cover image may be generated with reference to content A.
[0133] In some possible implementations, the first multimedia content may include video content and text content. The text content may be associated with the video content. For example, the text content may be a summary of the video content, or an introduction to the video content. In this case, the reference content for generating the cover image may include the text content. In other words, the cover image of the first multimedia content may be generated based on the text content. In this way, the generated cover image has a high degree of association with the first multimedia content, and the video content in the first multimedia content can be reflected through the cover image.
[0134] For example, the first multimedia content is a first conference recording file, which is a recording of a first video conference. The reference content for generating the cover image can include the minutes of the first video conference. In this way, the cover image generated by referencing the minutes is highly relevant to the video conference and better reflects the content of the video conference.
[0135] In an embodiment of the present application, the description information can be determined through a conversation page. In a specific implementation, the conversation page can be displayed, a target operation triggered in the conversation page is received, and then, based on the target operation, the description information for the cover image of the first multimedia content is determined.
[0136] The conversation page can be used to interact with a digital assistant. Digital assistants, also known as artificial intelligence (AI) assistants or conversational robots, typically have natural language processing capabilities and can analyze user input and generate corresponding responses, thereby enabling human-computer dialogue.
[0137] The digital assistant is accessed through the server, and a dialogue page is provided on the client, through which the user can conduct a human-computer dialogue with the digital assistant. The embodiments of the present application do not limit the deployment method of the server and the digital assistant. In some embodiments, the server and the digital assistant can both be independent software systems. In other embodiments, the server and the digital assistant can both be deployed in other systems (for example, a business platform), that is, the server and the digital assistant can be system modules in other systems.
[0138] The present application supports displaying the conversation page in different ways. In some embodiments, the client can display the conversation page alongside the first multimedia content display page. For example, in the client, the first multimedia content display page is displayed on the left, and the conversation page is displayed on the right. In this way, the user can interact with the digital assistant in a "separate conversation" manner.
[0139] In other embodiments, the client may display the conversation page in the form of a floating window component in the display page of the first multimedia content. For example, the conversation page may be displayed in the form of a floating window component above the cursor position or the position touched by the user in the display page of the first multimedia content. In this way, the user can interact with the digital assistant in a "floating window" manner.
[0140] Embodiments of the present application also support displaying a conversation page in different triggering modes. In some embodiments, the client can display a conversation page in response to a triggering operation on a first control. The first control is used to trigger the generation of a cover image for the first multimedia content.
[0141] In other words, the client can provide a first control for triggering the generation of a cover image. For example, the first control can be provided on a display page of the first multimedia content. After the user triggers the control, a conversation page can be displayed. Furthermore, since the first control is used to trigger the generation of a cover image for the first multimedia content, by triggering the control, the digital assistant can directly learn the user's operation intention (i.e., generating a cover image for the first multimedia content).
[0142] In other embodiments, the client may display a conversation page in response to a trigger operation on the digital assistant.
[0143] In other words, the client can provide an entry for triggering interaction with the digital assistant, for example, providing an icon corresponding to the digital assistant in the display page of the first multimedia content, and after the user triggers the icon, the dialogue page can be displayed.
[0144] It is understandable that since the user only triggers the digital assistant, the digital assistant cannot know the user's operation intention. Therefore, the server can determine the operation intention through subsequent user operations.
[0145] In the conversation page, descriptive information is determined in different ways. In some embodiments, an information input operation triggered in the conversation page can be received, and first information representing natural language content can be obtained based on the information input operation. Then, the first information is determined as descriptive information for the cover image of the first multimedia content, or reference content for instructing the generation of the cover image is determined based on the first information, and the reference content for instructing the generation of the cover image is determined as the descriptive information for the cover image of the first multimedia content.
[0146] In other words, the user can send the first information in the conversation page (for example, in the input box provided on the conversation page) by inputting natural language content, and the server can determine the descriptive information of the cover image for the first multimedia content based on the first information. For example, the user can enter "minimalist office style" in the conversation page. At this time, the information entered by the user can be directly determined as the descriptive information. For another example, the user can enter "generate cover image based on meeting minutes" in the conversation page. At this time, the reference content for indicating the generation of the cover image can be determined to be the meeting minutes based on the information entered by the user, and the reference content for indicating the generation of the cover image can be determined as the descriptive information.
[0147] In other embodiments, a control triggering operation may be received on a dialog page, and second information indicated by the target control may be obtained based on the control triggering operation. The second information may then be determined as descriptive information for a cover image of the first multimedia content, or reference content for instructing the generation of a cover image may be determined based on the second information, and the reference content for instructing the generation of a cover image may be determined as the descriptive information for the cover image of the first multimedia content.
[0148] In other words, the dialog page can provide multiple controls, and different controls can correspond to different second information. The server can determine the descriptive information of the cover image for the first multimedia content based on the second information corresponding to the target control triggered by the user. For example, the second information corresponding to the target control can be "a spacious and bright study". In this case, the second information corresponding to the target control can be directly determined as the descriptive information. For another example, the second information corresponding to the target control can be "the cover image is generated from the meeting minutes". In this case, the reference content indicating the generation of the cover image can be determined to be the meeting minutes based on the second information corresponding to the target control, and the reference content indicating the generation of the cover image can be determined as the descriptive information.
[0149] In some possible implementations, the second information corresponding to the control provided on the conversation page may be historical description information, or may be description information that is frequently used by other users, and this application does not impose any restrictions on this.
[0150] S302: Present at least one candidate cover image for the first multimedia content.
[0151] In some possible implementations, at least one candidate cover image for the first multimedia content may be presented in the conversation page.
[0152] At least one candidate cover image is generated by the digital assistant based on the description information. For example, when the description information includes at least one descriptive item related to the cover image, the digital assistant can generate a candidate cover image that satisfies the at least one descriptive item based on the at least one descriptive item. For another example, when the description information includes reference content indicating the generation of the cover image, the digital assistant can generate a candidate cover image related to the reference content based on the reference content for generating the cover image.
[0153] As mentioned above, when the user triggers the digital assistant and displays the dialogue page, the digital assistant cannot directly know the user's operation intention. Therefore, it can determine the operation intention indicated by the target operation based on the target operation.
[0154] In other words, the server can identify the user's intention for the target operation in the dialogue page and determine the user's operation intention. For example, the server can identify and analyze the user's input information in the dialogue page to determine the user's operation intention.
[0155] In an embodiment of the present application, the server can call different processing tools (also known as plug-ins) to process different operation intentions and implement different functions. In some embodiments, in response to the operation intention including generating a cover image, the server can call a first processing tool to cause the first processing tool to generate at least one candidate cover image.
[0156] The first processing tool may be capable of generating images based on text. For example, the first processing tool may be a processing tool based on generative artificial intelligence (AI) generated content. By calling the first processing tool, the candidate cover image is generated.
[0157] In specific implementation, the first processing tool can send the description information and the third information representing the cover image generation requirements to the image generation model, and receive at least one candidate cover image returned by the image generation model.
[0158] In other words, the first processing tool can generate a cover image using an image generation model. The image generation model has natural language processing and image generation capabilities. The first processing tool sends descriptive information and third information representing cover image generation requirements to the image generation model. The image generation model can identify and analyze the third information, understand the cover image generation requirements, and, based on the descriptive information, generate at least one candidate cover image for the first multimedia content that meets the cover image generation requirements.
[0159] In some possible implementations, the description information includes reference content indicating generation of the cover image, and the reference content indicating generation of the cover image includes text content in the first multimedia content.
[0160] The first processing tool can determine the identification information of the video content in the first multimedia content based on the description information, obtain the text content associated with the video content based on the identification information of the video content, and then send the description information, text content and third information representing the cover image generation requirements to the image generation model, and receive at least one candidate cover image returned by the image generation model.
[0161] That is, when the description information includes reference content indicating the generation of a cover image, the first processing tool can determine the identification information of the video content through the description information, such as the sender information carried in the description information, or the field in the description information indicating the video content. Then, using the identification information of the video content, the target processing tool can obtain text content associated with the video content. In this way, the target processing tool can obtain the text content and, in combination with the text content in the first multimedia content, generate a candidate cover image for the first multimedia content, thereby increasing the degree of association between the candidate cover image and the first multimedia content.
[0162] For example, the first multimedia content is a recording of a first video conference, and the reference content for generating a cover image is the meeting minutes of the first video conference. The first processing tool can obtain the meeting minutes of the first video conference based on the meeting identifier of the first video conference indicated by the descriptive information. The first processing tool then sends the descriptive information, the meeting minutes of the first video conference, and third information representing the cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model.
[0163] Among them, the minutes of the first video conference can be automatically generated by the video conferencing system based on the content of the video conference, or can be manually input by the user. This application does not impose any restrictions on this.
[0164] In other embodiments, in response to the operation intention not including generating a cover image, the server may call a second processing tool to enable the second processing tool to perform natural language processing on the operation intention.
[0165] The second processing tool may have the ability to process natural language tasks. For example, the second processing tool may be a processing tool based on a language model, and the language model may be a deep learning model trained based on text data.
[0166] By invoking the second processing tool, the language model can perform natural language processing based on the user's operational intent. Specifically, the language model can analyze the first message representing the natural language content input by the user and generate a response to the first message, thus achieving human-computer interaction. For example, if the user enters the first message "Help me summarize the meeting content," by invoking the second processing tool, the language model can generate a meeting summary and reply to the user with the meeting summary.
[0167] Furthermore, after presenting at least one candidate cover image, the target cover image can be determined as the cover image of the first multimedia content in response to a selection operation on the target cover image among the at least one candidate cover image.
[0168] That is, the user can select from the generated candidate cover images and select the target cover image that meets their needs as the cover image of the first multimedia content. For example, the user can click on a candidate cover image to select that candidate cover image as the cover image of the first multimedia content. In this way, the cover image of the first multimedia content can be generated or replaced without uploading a local image.
[0169] It can be understood that if none of the generated candidate cover images meet the user's needs or preferences, the user can also trigger an instruction to regenerate the cover image. For example, a retry control can be provided on the conversation page, and the user can trigger the control to regenerate a batch of candidate cover images. For another example, the user can also enter information representing natural language content on the conversation page, such as "help me regenerate", to regenerate a batch of candidate cover images.
[0170] After determining the cover image of the first multimedia content, the cover image of the first multimedia content may also be presented on a page associated with the first multimedia content.
[0171] Since the cover image is generated based on the description information, and the description information may be related to the first multimedia content, in a page associated with the first multimedia content, the user may quickly locate the first multimedia content through the cover image and learn about information related to the first multimedia content.
[0172] The page associated with the first multimedia content includes at least one of the following: a viewing page of historical multimedia content, an instant messaging page after sharing the first multimedia content, and a preview page of the first multimedia content.
[0173] The historical multimedia content viewing page (e.g., a multimedia content list) may present at least one multimedia content. For example, when the multimedia content is a meeting recording file, the historical multimedia content viewing page may present meeting recording files created by the current user, or meeting recording files for which the current user has viewing permission. The multimedia content presented on the historical multimedia content viewing page may include the first multimedia content, and the historical multimedia content viewing page may present a cover image of the first multimedia content.
[0174] The instant messaging (IM) page may be a chat page provided by the IM system. It is understandable that the first user can share the first multimedia content with other users of the IM system so that other users can view it. For example, the first user can share the first multimedia content with the second user of the IM system. In this case, the instant messaging page may be a chat page between the first user and the second user. For another example, the first user can share the first multimedia content with multiple users of the IM system (such as multiple users in the first group). In this case, the instant messaging page may be a chat page of the first group. After sharing the first multimedia content, the first multimedia content may be presented in the instant messaging page in the form of a card message. In other words, the cover image of the first multimedia content is presented in the instant messaging page, and other users can click on the cover image of the first multimedia content in the instant messaging page to view the first multimedia content.
[0175] The preview page of the first multimedia content can be used to preview the first multimedia content. For example, the preview page of the first multimedia content can provide multiple operation controls (such as a sharing control, a translation control, etc.) for the first multimedia content. The user selects the desired operation control from the multiple operation controls. The preview page of the first multimedia content can also present a cover image of the first multimedia content to inform the user that the relevant operation is currently being performed on the first multimedia content.
[0176] Based on the above description, an embodiment of the present application provides a cover image generation method. The method first receives description information of a cover image for a first multimedia content, and then presents at least one candidate cover image for the first multimedia content, wherein the at least one candidate cover image is generated by a digital assistant based on the description information.
[0177] This method addresses the generation of multimedia content cover images. By connecting to a digital assistant and leveraging its image-generating capabilities, the digital assistant is able to generate candidate cover images for the multimedia content based on the cover image's descriptive information. This method addresses the diverse user needs for multimedia content cover images by generating candidate cover images that meet those needs. Furthermore, the cover images more intuitively reflect the multimedia content, enhancing the user's viewing experience.
[0178] The foregoing describes the cover image generation method provided in this application. The following will introduce an application scenario in which the first multimedia content is a first conference recording file.
[0179] See also Figures 4A to 4E Schematic diagram of the display page of the conference recording file shown in FIG. Figure 4A As shown, the display page 40 of the conference recording file includes a title bar 401. The user can trigger the cover image generation operation for the first conference recording file through the first control in the title bar 401. Specifically, the title bar 401 may include a sharing control and an additional function control. After the user clicks the additional function control, a variety of function controls can be presented, such as a statistics view control, a feedback submission control, and an AI cover image change control. Among them, the AI cover image change control is the first control. After the user clicks the AI cover image change control, the conversation page can be displayed to interact with the digital assistant.
[0180] like Figure 4BAs shown, the conversation page 50 can be displayed side by side with the display page 40 of the first conference recording file. Since the user triggers the generation of the cover image for the first conference recording file by triggering the first control, the digital assistant can know the user's operating intention (i.e., generate a cover image for the first conference recording file). Therefore, the conversation page 50 can display the conversation content 501. The conversation content 501 can be provided by the digital assistant, such as the content sent by the digital assistant "Hi, I am your digital assistant, I can help you generate a cover image for the conference recording file", and the conversation content 501 can also provide multiple controls, such as a control for generating meeting minutes, a control for a simple office style, and a control for a spacious and bright study. Different controls can correspond to different second information. The user can quickly input the second information by clicking the target control. In this way, the description information for the first conference recording file is determined based on the second information.
[0181] In other embodiments, Figure 4C As shown, the display page 40 of the conference recording file includes a title bar 401, and the title bar 401 can provide an entrance 4012 of the digital assistant. The user can trigger the entrance 4012 of the digital assistant to display the conversation page and interact with the digital assistant.
[0182] like Figure 4D As shown, the conversation page 50 can be displayed side by side with the display page 40 of the first conference recording file. Since the user only triggers the digital assistant, the digital assistant cannot know the user's operation intention. Therefore, the user can trigger the information input operation in the conversation page. The conversation content 502 in the conversation page 50 may include information representing natural language content entered by the user. For example, if the user enters "Help me change the cover image" in the input box provided in the conversation page 50, the digital assistant can determine the operation intention based on the information entered by the user.
[0183] After the digital assistant determines the operation intention, it can generate a corresponding reply content. At this time, the dialogue page 50 can also include the digital assistant's reply content, such as "You can click the button below to confirm the description information, or you can enter the description information yourself." Similarly, the dialogue content 501 can also provide multiple controls, which will not be repeated here.
[0184] like Figure 4E As shown, the user inputs the first information by natural language input, and the conversation content 503 of the conversation page 50 may include the first information. After the digital assistant receives the first information, it can determine the description information and generate at least one candidate cover image based on the description information. Therefore, the conversation content 503 may also include the generated at least one candidate cover image, such as the candidate cover image. Figure 1 , candidate cover Figure 2 , candidate cover Figure 3As for candidate cover image 4, the user can automatically generate and replace the cover image of the first conference recording file by clicking on the cover image.
[0185] Combined with the above Figure 1 Figure 4 provides a detailed introduction to the cover image generation method provided in the embodiment of the present application. The system and equipment provided in the embodiment of the present application will be introduced below in conjunction with the accompanying drawings.
[0186] See also Figure 5 The schematic structural diagram of the cover image generation system shown in FIG. 5 includes:
[0187] Determining module 501, configured to determine description information of a cover image for a first multimedia content;
[0188] The presentation module 502 is used to present at least one candidate cover image of the first multimedia content, where the at least one candidate cover image is generated by the digital assistant based on the description information.
[0189] In some possible implementations, the description information includes at least one description item related to the cover image, or the description information includes reference content indicating generation of the cover image.
[0190] In some possible implementations, the first multimedia content includes video content and text content, the text content is associated with the video content, and the reference content for generating the cover image includes the text content.
[0191] In some possible implementations, the determining module 501 is specifically configured to:
[0192] Displaying a dialogue page, wherein the dialogue page is used to interact with the digital assistant;
[0193] Receiving a target operation triggered in the conversation page;
[0194] Describe information of a cover image for the first multimedia content according to the target operation.
[0195] In some possible implementations, the determining module 501 is specifically configured to:
[0196] Receiving an information input operation triggered in the conversation page;
[0197] Acquiring first information representing natural language content according to the information input operation;
[0198] Determine the first information as description information of a cover image of the first multimedia content; or,
[0199] Based on the first information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
[0200] In some possible implementations, the determining module 501 is specifically configured to:
[0201] Receiving a control triggering operation in the dialog page;
[0202] According to the control triggering operation, obtaining second information indicated by the target control;
[0203] Determine the second information as description information of the cover image of the first multimedia content; or,
[0204] Based on the second information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
[0205] In some possible implementations, the determining module 501 is specifically configured to:
[0206] In response to a triggering operation on a first control, displaying a dialog page, wherein the first control is used to trigger generation of a cover image for the first multimedia content; or
[0207] In response to a triggering operation on the digital assistant, a conversation page is displayed.
[0208] In some possible implementations, the determining module 501 is specifically configured to:
[0209] displaying the conversation page in parallel with the display page of the first multimedia content; or,
[0210] In the display page of the first multimedia content, the conversation page is displayed in the form of a floating window component.
[0211] In some possible implementations, the system further includes a calling module, wherein the calling module is configured to:
[0212] Determining, according to the target operation, an operation intention indicated by the target operation;
[0213] In response to the operation intention including generating a cover image, a first processing tool is called to enable the first processing tool to generate at least one candidate cover image.
[0214] In some possible implementations, the first processing tool generates at least one candidate cover image, including:
[0215] The first processing tool sends the description information and the third information representing the cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model.
[0216] In some possible implementations, the description information includes reference content indicating generation of a cover image, the reference content indicating generation of the cover image includes text content in the first multimedia content, the first processing tool sends the description information and third information representing cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model, including:
[0217] Determining identification information of video content in the first multimedia content according to the description information;
[0218] Acquiring text content associated with the video content according to the identification information of the video content;
[0219] The description information, the text content and the third information representing the cover image generation requirements are sent to the image generation model, and at least one candidate cover image returned by the image generation model is received.
[0220] In some possible implementations, the calling module is further configured to:
[0221] In response to the fact that the operation intention does not include generating a cover image, a second processing tool is called to enable the second processing tool to perform natural language processing on the operation intention.
[0222] In some possible implementations, the presenting module 502 is specifically configured to:
[0223] In the conversation page, at least one candidate cover image of the first multimedia content is presented.
[0224] In some possible implementations, the determining module 501 is further configured to:
[0225] In response to a selection operation on a target cover image from the at least one candidate cover image, the target cover image is determined as the cover image of the first multimedia content.
[0226] In some possible implementations, the presenting module 502 is further configured to:
[0227] On a page associated with the first multimedia content, a cover image of the first multimedia content is presented.
[0228] In some possible implementations, the page associated with the first multimedia content includes at least one of the following: a viewing page of historical multimedia content, an instant messaging page after sharing the first multimedia content, and a preview page of the first multimedia content.
[0229] The cover image generation system 50 according to the embodiment of the present application may correspond to the method described in the embodiment of the present application, and the above and other operations and / or functions of each module / unit of the cover image generation system 50 are respectively to achieve Figure 3 For the sake of brevity, the corresponding processes of the various methods in the illustrated embodiments are not described again here.
[0230] The embodiment of the present application also provides an electronic device. The electronic device is specifically used to implement Figure 5 The functions of the cover image generation system 50 in the illustrated embodiment.
[0231] Figure 6 A structural diagram of an electronic device 600 is provided. Figure 6 As shown, electronic device 600 includes bus 601, processor 602, communication interface 603 and memory 604. Processor 602, memory 604 and communication interface 603 communicate with each other via bus 601.
[0232] The bus 601 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0233] The processor 602 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0234] The communication interface 603 is used for communicating with the outside, for example, the communication interface 603 can be used for communicating with a terminal.
[0235] The memory 604 may include volatile memory, such as random access memory (RAM), or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0236] The memory 604 stores executable code, and the processor 602 executes the executable code to perform the aforementioned cover image generation method.
[0237] Specifically, in the implementation Figure 5 In the case of the embodiment shown, and Figure 5 When each module or unit of the cover image generation system 50 described in the embodiment is implemented by software, Figure 5 The software or program code required for the functions of each module / unit in the application may be partially or completely stored in the memory 604. The processor 602 executes the program code corresponding to each unit stored in the memory 604 and executes the aforementioned cover image generation method.
[0238] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned cover image generation method applied to the cover image generation system 60.
[0239] The present application also provides a computer program product comprising one or more computer instructions that, when loaded and executed on a computing device, fully or partially generate the process or function described in the present application.
[0240] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0241] When the computer program product is executed by a computer, the computer executes any of the aforementioned methods for generating a cover image. The computer program product may be a software installation package, and when any of the aforementioned methods for generating a cover image is required, the computer program product may be downloaded and executed on the computer.
[0242] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to the various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the prescribed logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0243] The units involved in the embodiments described in this application may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, constitute a limitation on the unit itself.
[0244] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0245] In the context of the present application embodiment, machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include but is not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0246] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0247] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0248] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0249] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0250] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A cover image generation method, characterized in that: The method comprises: Determining descriptive information for a cover image of the first multimedia content; Present at least one candidate cover image for the first multimedia content, where the at least one candidate cover image is generated by a digital assistant based on the description information.
2. The method according to claim 1, characterized in that The description information includes at least one description item related to the cover image, or the description information includes reference content indicating generation of the cover image.
3. The method according to claim 2, characterized in that The first multimedia content includes video content and text content, the text content is associated with the video content, and the reference content for generating the cover image includes the text content.
4. The method according to claim 1, wherein The determining of description information for the cover image of the first multimedia content includes: Displaying a dialogue page, wherein the dialogue page is used to interact with the digital assistant; Receiving a target operation triggered in the conversation page; Describe information of a cover image for the first multimedia content according to the target operation.
5. The method according to claim 4, characterized in that The receiving of a target operation triggered in the conversation page includes: Receiving an information input operation triggered in the conversation page; The determining, according to the target operation, descriptive information of the cover image for the first multimedia content includes: Acquiring first information representing natural language content according to the information input operation; Determine the first information as description information of a cover image of the first multimedia content; or, Based on the first information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
6. The method according to claim 4, characterized in that The receiving of a target operation triggered in the conversation page includes: Receiving a control triggering operation in the dialog page; The determining, according to the target operation, descriptive information of the cover image for the first multimedia content includes: According to the control triggering operation, obtaining second information indicated by the target control; Determine the second information as description information of the cover image of the first multimedia content; or, Based on the second information, reference content indicating generation of a cover image is determined, and the reference content indicating generation of a cover image is determined as descriptive information of the cover image for the first multimedia content.
7. The method according to claim 4, characterized in that The display dialogue page includes: In response to a triggering operation on a first control, displaying a dialog page, wherein the first control is used to trigger generation of a cover image for the first multimedia content; or In response to a triggering operation on the digital assistant, a conversation page is displayed.
8. The method according to claim 4, characterized in that The display dialogue page includes: displaying the conversation page in parallel with the display page of the first multimedia content; or, In the display page of the first multimedia content, the conversation page is displayed in the form of a floating window component.
9. The method according to claim 4, characterized in that After receiving the target operation triggered in the conversation page, the method further includes: Determining, according to the target operation, an operation intention indicated by the target operation; In response to the operation intention including generating a cover image, a first processing tool is called to enable the first processing tool to generate at least one candidate cover image.
10. The method according to claim 9, characterized in that The first processing tool generates at least one candidate cover image, including: The first processing tool sends the description information and the third information representing the cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model.
11. The method according to claim 10, characterized in that The description information includes reference content indicating generation of a cover image, the reference content indicating generation of the cover image includes text content in the first multimedia content, the first processing tool sends the description information and third information representing cover image generation requirements to the image generation model, and receives at least one candidate cover image returned by the image generation model, including: Determining identification information of video content in the first multimedia content according to the description information; Acquiring text content associated with the video content according to the identification information of the video content; The description information, the text content and the third information representing the cover image generation requirements are sent to the image generation model, and at least one candidate cover image returned by the image generation model is received.
12. The method according to claim 9, characterized in that The method further comprises: In response to the fact that the operation intention does not include generating a cover image, a second processing tool is called to enable the second processing tool to perform natural language processing on the operation intention.
13. The method according to claim 1, wherein The presenting of at least one candidate cover image for the first multimedia content includes: In the conversation page, at least one candidate cover image of the first multimedia content is presented.
14. The method according to any one of claims 1 to 13, characterized in that The method further comprises: In response to a selection operation on a target cover image from the at least one candidate cover image, the target cover image is determined as the cover image of the first multimedia content.
15. The method according to claim 14, characterized in that The method further comprises: On a page associated with the first multimedia content, a cover image of the first multimedia content is presented.
16. The method according to claim 15, characterized in that The page associated with the first multimedia content includes at least one of the following: a viewing page of historical multimedia content, an instant messaging page after sharing the first multimedia content, and a preview page of the first multimedia content.
17. A cover image generation system, characterized in that: The system comprises: A determination module, configured to determine description information of a cover image for the first multimedia content; A presentation module is used to present at least one candidate cover image of the first multimedia content, where the at least one candidate cover image is generated by a digital assistant based on the description information.
18. An electronic device, characterized in that: The electronic device includes a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that The method comprises instructions for instructing an electronic device to execute the method according to any one of claims 1 to 16.