Text generation method, device, electronic device and storage medium

By obtaining the target image and its content category and using a large model to generate text, the problem of single input data in the existing technology is solved, and diversified text generation methods are realized to meet users' text generation needs for different content categories.

CN118410779BActive Publication Date: 2025-09-23BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410480787.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-09-23
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

In the prior art, users can only input text through text and voice, which results in single input data and lacks diversified data, and cannot enrich the user's text generation method.

Method used

By obtaining the target image and its corresponding content category and sending a text generation request to the server, the large model is used to generate text of the target content category, supporting the generation of multiple image input methods and content categories.

Benefits of technology

It enriches the input methods of text generation, meets users' needs for text generation of different content categories, and improves the flexibility and diversity of text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118410779B_ABST
    Figure CN118410779B_ABST
Patent Text Reader

Abstract

This application discloses a text generation method, device, electronic device, and storage medium, and relates to the computer field, particularly to artificial intelligence fields such as deep learning and large models. A specific implementation scheme is as follows: obtaining a target image and a target content category corresponding to the target image; wherein the target content category is the content category of the text to be generated for the target image; in response to detecting a text generation operation, sending a text generation request to a server; wherein the text generation request includes prompt text corresponding to the target image and the target content category; receiving the target text sent by the server and displaying the target text; wherein the target text is text of the target content category generated by processing the target image and the prompt text using a large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, in particular to artificial intelligence fields such as deep learning and large models, and specifically to a text generation method, device, electronic device and storage medium. Background Art

[0002] Large models are machine learning models with large parameter sizes and complexity. They can not only generate natural language text, but also deeply understand the meaning of the text and handle various natural language tasks such as text summarization, question answering, and translation. Summary of the Invention

[0003] The present application provides a text generation method, device, electronic device and storage medium.

[0004] According to one aspect of the present application, a text generation method is provided, comprising:

[0005] Obtaining a target image and a target content category corresponding to the target image; wherein the target content category is a content category of a text to be generated from the target image;

[0006] In response to detecting the text generation operation, sending the text generation request to the server; wherein the text generation request includes the target image and the prompt text corresponding to the target content category;

[0007] Receive the target text sent by the server and display the target text; wherein the target text is a text of the target content category generated by processing the target image and the prompt text using a large model.

[0008] According to another aspect of the present application, a text generation method is provided, comprising:

[0009] Receive a text generation request sent by a client; wherein the text generation request includes a target image and a prompt text corresponding to a target content category, and the target content category is a content category of the text to be generated for the target image;

[0010] Input the target image and the prompt text into a large model to generate a target text of the target content category;

[0011] The target text is sent to the client.

[0012] According to another aspect of the present application, a text generation device is provided, comprising:

[0013] An acquisition module, configured to acquire a target image and a target content category corresponding to the target image; wherein the target content category is a content category of a text to be generated from the target image;

[0014] A sending module, configured to send the text generation request to the server in response to detecting the text generation operation; wherein the text generation request includes the target image and the prompt text corresponding to the target content category;

[0015] A receiving module is used to receive the target text sent by the server;

[0016] A display module is used to display the target text; wherein the target text is a text of the target content category generated by processing the target image and the prompt text using a large model.

[0017] According to another aspect of the present application, a text generation device is provided, comprising:

[0018] A receiving module, configured to receive a text generation request sent by a client; wherein the text generation request includes a target image and a prompt text corresponding to a target content category, wherein the target content category is a content category of the text to be generated for the target image;

[0019] A generation module, configured to input the target image and the prompt text into a macro model to generate a target text of the target content category;

[0020] The sending module is used to send the target text to the client.

[0021] According to another aspect of the present application, an electronic device is provided, including:

[0022] at least one processor; and

[0023] a memory communicatively connected to the at least one processor; wherein,

[0024] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the above embodiment.

[0025] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to the above embodiment.

[0026] According to another aspect of the present application, a computer program product is provided, including a computer program, which implements the steps of the method described in the above embodiment when executed by a processor.

[0027] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.

[0029] Figure 1 A flowchart of a text generation method provided in one embodiment of the present application;

[0030] Figure 2 Schematic diagram of the document assistant dialogue interface provided in this embodiment of the application Figure 1 ;

[0031] Figure 3 Schematic diagram of the photo trigger control and picture list display interface provided in an embodiment of the present application;

[0032] Figure 4 Schematic diagram of the photo taking interface provided in this embodiment of the application Figure 1 ;

[0033] Figure 5 A schematic diagram of the image display interface provided in an embodiment of the present application;

[0034] Figure 6 Schematic diagram of the document assistant dialogue interface provided in this embodiment of the application Figure 2 ;

[0035] Figure 7 Schematic diagram of the document assistant dialogue interface provided in this embodiment of the application Figure 3 ;

[0036] Figure 8 A schematic diagram of the client homepage provided in an embodiment of the present application;

[0037] Figure 9 Schematic diagram of the photo taking interface provided in this embodiment of the application Figure 2 ;

[0038] Figure 10 A schematic diagram of the image list display interface provided in an embodiment of the present application;

[0039] Figure 11 A flowchart of a text generation method provided in another embodiment of the present application;

[0040] Figure 12 Schematic diagram of the photo taking interface provided in this embodiment of the application Figure 3 ;

[0041] Figure 13A flowchart of a text generation method provided in another embodiment of the present application;

[0042] Figure 14 A flowchart of a text generation method provided in another embodiment of the present application;

[0043] Figure 15 A schematic diagram of the structure of a text generation device provided in one embodiment of the present application;

[0044] Figure 16 A schematic structural diagram of a text generation device provided in another embodiment of the present application;

[0045] Figure 17 It is a block diagram of an electronic device used to implement the text generation method of an embodiment of the present application. DETAILED DESCRIPTION

[0046] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0047] The text generation method, device, electronic device, and storage medium of the embodiments of the present application are described below with reference to the accompanying drawings.

[0048] In some embodiments, the model generates text based on user input. However, the input data is text, which has a single form. The user can only input instruction information to the model through text input or voice-to-text conversion.

[0049] Based on this, an embodiment of the present application provides a text generation method. Figure 1 A flowchart of a text generation method provided in one embodiment of the present application.

[0050] The text generation method of the embodiment of the present application can be executed by the text generation device of the embodiment of the present application, which can be configured in an electronic device. For example, the text generation device can be configured in a client, which can be installed in the electronic device.

[0051] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.

[0052] like Figure 1 As shown, the text generation method includes:

[0053] Step 101: Obtain a target image and a target content category corresponding to the target image.

[0054] In this application, a user can input an image through the client's image input portal. In response to detecting the user's image input operation, the client obtains a target image. The input operation can be an image capture operation, an image selection operation, etc. In other words, the target image can be generated by taking a photo or selected from local images, without limitation. The target image can be any image that is captured or selected.

[0055] For example, the client may display a photographing interface, and in response to detecting a triggering operation on a photographing control on the photographing interface, obtain a captured target image and display the target image.

[0056] For example, the client has a picture answer control that the user can trigger. Upon detecting the user triggering the picture answer control, the client displays a photo taking interface. If a triggering operation is detected on the photo taking control on the photo taking interface, the client retrieves and displays the target photo. Thus, the photo taking interface can be displayed by triggering the picture answer control, and the target photo can be retrieved by taking a photo, enriching the image input method.

[0057] For example, a picture answer control can be displayed on the client's homepage. Furthermore, the picture answer control can be displayed in a target area on the client's homepage, such as a diamond area. The diamond area can refer to a core functional area on the page. In some embodiments, the diamond area can also be called a diamond position.

[0058] For example, the client has a document assistant control. When a user triggers the document assistant control, the document assistant dialog interface is displayed. If a user triggers the image add control on the document assistant interface, the photo trigger control is displayed. If a user triggers the photo trigger control, the photo interface is displayed. Thus, the user can enter the photo interface through the document assistant dialog interface and then obtain the target image by taking a photo, enriching the image input method.

[0059] For example, a document assistant control may be displayed floating on the client interface. When the user triggers the display of the document assistant control, the client may display the document assistant dialogue interface.

[0060] For example, the client may display a picture list, and upon detecting a user's selection operation on a target picture in the picture list, the target picture may be acquired and a picture display interface of the target picture may be displayed.

[0061] For example, a user can trigger a picture answer control. In response to detecting the user's triggering operation on the picture answer control, the client displays a photo interface, which includes a picture import control. If the user triggers the picture import control on the photo interface, a picture list is displayed. Thus, by triggering the picture answer control to display a picture list, the user can then select a target picture from the picture list, enriching the picture input method.

[0062] For example, when a user triggers a document assistant control, the document assistant dialog interface is displayed. If a user triggers an image add control on the document assistant interface, a list of images is displayed. Thus, by triggering the image add control on the document assistant dialog interface, a list of images can be displayed, and then a target image can be selected from the list of images, enriching the image input method.

[0063] For example, if it is detected that the user triggers the image adding control on the document assistant interface, the photo trigger control is displayed. If it is detected that the user triggers the photo trigger control, the photo interface is displayed, and the photo interface displays the image import control. The user triggers the image import control, the client can display a list of images, and the user can select the target image from the image list.

[0064] In this application, multiple content categories can be displayed on the client's photo-taking interface or on the image display interface of the target image after selecting the target image. The user can select the corresponding content category for the target image as needed. In response to detecting the user's selection of the content category for the text to be generated for the target image, the client obtains the target content category corresponding to the target image. The target content category can refer to the content category of the text to be generated for the target image.

[0065] For example, content categories may include but are not limited to image analysis, image captioning, image generation into PPT, etc.

[0066] For example, if the target content category is "image-to-text," the target image's to-be-generated text will be the analyzed content of the target image. Another example: if the target content category is "image-to-text," the target image's to-be-generated text will be the copywriting content that matches the target image. Another example: if the target content category is "image-to-PowerPoint," the target image's to-be-generated text will be the PPT generated based on the target image.

[0067] For example, a user can capture a target image by taking a photo, and a content category can be displayed on the photo taking interface. When a user selects a target content category on the photo taking interface, the target content category corresponding to the target image is obtained. Thus, when the user captures a target image by taking a photo, they can select the target content category from the content categories displayed on the photo taking interface, making the operation simple and convenient.

[0068] It should be noted that the user can select the target content category before taking a photo, and then trigger the photo control to obtain the target image, or trigger the photo control to obtain the target image, and then select the target content category. There is no limitation on this.

[0069] For example, after a user selects a target image from a list of images, a picture display interface for the target image is displayed, where the picture display interface can display content categories. When a selection operation of a target content category on the picture display interface is detected, the target content category corresponding to the target image can be obtained. Thus, when a user selects a target image from a local picture list, they can select the target content category from the content categories displayed on the picture display interface for the target image, which is simple and convenient to operate.

[0070] For example, after taking a photo to obtain a target image or selecting a target image from a picture list, the size of the target image can be adjusted to meet the personalized needs of the user.

[0071] Step 102: In response to detecting a text generation operation, sending a text generation request to a server.

[0072] In this application, different prompt texts can be preset for different content categories, wherein the prompt text can be used to prompt the large model to perform the image-to-text task to generate text of the corresponding content category.

[0073] For example, the prompt text corresponding to the content category "Picture Caption" is "Generate a copy based on the picture content."

[0074] For example, the prompt text may include task requirements, output text requirements, etc. The task requirements may refer to what content category of text is generated based on the image, and the output text requirements may include text content requirements, word count requirements, language style requirements, etc.

[0075] For example, the prompt text is "Generate a life record copy based on the content of the picture. 1. Content: Understand the people or objects in the picture and share details of life experiences; 2. Number of words: No more than 100 words; 3. Language style should be colloquial; 4. Do not start with "just now"." The task requirement of the prompt text is "Generate a life record copy based on the content of the picture." The rest of the prompt text is the output text requirements.

[0076] In this application, after the user selects the target image and the content category corresponding to the target image, he can trigger the confirmation control on the client interface. Then the client detects the text generation operation and sends a text generation request to the server.

[0077] The text generation request may include a target image, prompt text corresponding to the target content category, and the like.

[0078] Step 103: Receive the target text sent by the server and display the target text.

[0079] The target text may be a text of the target content category generated by the server using a large model to process the target image and the prompt text. It is understandable that the target text is a text of the target content category.

[0080] For example, if the target content category selected by the user is press release, then the target text is press release.

[0081] For example, after taking a photo to obtain the target image or selecting the target image locally, a confirmation control is displayed on the image display interface of the target image. After the user triggers the confirmation control, the client can display a text display interface, wherein the text display interface can display the target image, the prompt text corresponding to the target image, etc.

[0082] For example, a cancel control may be displayed on the image display interface of the target image. If the user triggers the cancel control, the client may return to the previous interface. For example, after taking a photo to obtain the target image, the image display interface of the target image is displayed. If the user triggers the cancel control on the image display interface, the client may return to the photo taking interface.

[0083] For example, the client may also display the text generation status on the text display interface, such as displaying a prompt message "Document Assistant is assisting with creation."

[0084] Exemplarily, after the entire target text is generated, the server may send the entire target text to the client, so that the client displays the entire target text.

[0085] For example, the server may send the generated characters one by one to the client, and the client may display the generated characters one by one until the last character of the target text.

[0086] For ease of understanding, the following Figure 2-Figure 10 The above text generation process is explained.

[0087] Figure 2 The document assistant dialog interface shown in the figure shows a picture adding control. When the user triggers the picture adding control, it shows Figure 3 The interface shown in the figure shows a photo trigger control and a picture list. If the user triggers the photo trigger control, the client displays Figure 4The photo interface shown in the figure has the shooting area, photo control, and content categories such as image analysis, image text, and image generation PPT. Users can trigger the photo control to take a photo to obtain a picture, and can also select the content category corresponding to the picture on the interface. After the user triggers the photo control, the content category can be displayed. Figure 5 The picture display interface shown in the figure, users can Figure 5 Select or change the content category on the picture display interface shown. Figure 3 Select a picture from the picture list shown, and the client can display Figure 5 The picture display interface shown in Figure 5 On the picture display interface shown, the user can select a content category for the selected picture.

[0088] If the user triggers Figure 5 The confirmation control on the interface shown, the client can display Figure 6 The interface shown, Figure 6 The interface shown displays prompt text and pictures corresponding to the content category, and the client sends the content category and picture selected by the user to the server.

[0089] User trigger Figure 4 After clicking the photo control on the photo interface shown, the client displays Figure 5 The picture display interface shown in the figure, when the user triggers Figure 5 The cancel control on the interface shown, the client returns to the previous interface, which is displayed Figure 4 The photo taking interface shown.

[0090] The client can display the target text returned by the server in the document assistant dialogue interface, such as Figure 7 As shown, Figure 7 The interface shown may also display operation controls for the target text, such as a copy control, a save control, an edit control, and the like.

[0091] Figure 8 The interface shown is the client home page, which displays a picture answer control and a document question and answer control. When the user triggers the picture answer control, it displays Figure 9 The interface shown in the figure shows the shooting area, content category, photo control, picture import control, etc. If the user triggers the photo control, the client displays Figure 5 In the interface shown, if the user triggers the image import control, the client displays Figure 10 The interface shown, Figure 10 The interface shown shows a list of pictures. Users can select a picture from the list, and the client can then display Figure 5 The interface shown. When the user triggers Figure 5The client can display the confirmation control in the interface shown. Figure 6 The interface shown in the figure is displayed, and the prompts corresponding to the selected image and content category are sent to the server. When the client receives the target text sent by the server and displays the target text, Figure 7 shown.

[0092] In an embodiment of the present application, by obtaining a target image and a target content category corresponding to the target image, and when a text generation operation is detected, a text generation request is sent to the server, and the target image and prompt text corresponding to the target content category are sent to the server, so that the server uses a large model to generate a target text of the target content category based on the target image and the prompt text. Therefore, in the text generation scenario, not only can the input of images be supported, the form of input data is expanded, but also, for different content categories selected by the user, texts of different content categories can be generated for the user based on the prompt texts corresponding to different content categories, thereby meeting different text generation needs.

[0093] Figure 11 A flowchart of a text generation method provided in another embodiment of the present application.

[0094] like Figure 11 As shown, the text generation method includes:

[0095] Step 1101: Obtain a target image and a target content category corresponding to the target image.

[0096] In this application, step 1101 can be implemented in any of the embodiments of this application, so it will not be described here in detail.

[0097] Step 1102 : Display multiple first-level content categories.

[0098] In this application, after taking a photo to obtain a target image or selecting a target image locally, a picture display interface of the target image can be displayed, and multiple first-level content categories can be displayed on the picture display interface of the target image.

[0099] For example, the first-level content categories include image analysis, image captioning, and image-generated PPT.

[0100] Step 1103 : In response to detecting a selection operation on a target first-level content category among the plurality of first-level content categories, display a plurality of second-level content categories under the target first-level content category.

[0101] In this application, there may be multiple second-level content categories under the first-level content category, or there may be no second-level content category, which is not limited.

[0102] For example, if there are many first-level content categories and the interface fails to display all of them, if the first-level content categories are displayed horizontally, the user can view or select the first-level content categories by sliding left or right, or select the first level by triggering the operation.

[0103] Taking the first-level content category of picture caption as an example, there can be multiple display channels under the picture caption, such as application a1, application a2, application a3, etc. Since the text content applicable to different display channels may be different, each display channel can be used as the second-level content category of picture caption.

[0104] For example, under the picture caption, there are multiple second-level content categories such as application a1, application a2, press releases, entertainment articles, and picture writing.

[0105] In this application, if a target first-level content category among multiple first-level content categories has multiple second-level content categories, when the user selects the target first-level content category, multiple second-level content categories under the target first-level content category may be displayed.

[0106] For example, when a user selects a text to accompany a picture on the photo-taking interface, multiple second-level content categories under the picture can be displayed, such as application a1, application a2, press releases, entertainment articles, and picture writing.

[0107] For example, if there are many second-level content categories under a first-level content category, and the interface cannot display all second-level content categories, if the second-level content categories are displayed horizontally, the user can view or select the second-level content categories by sliding left or right, or select the second-level content categories through a trigger operation.

[0108] Step 1104 : In response to detecting a selection operation on any second-level content category, determining any second-level content category as a target content category.

[0109] In the present application, if a user selection operation on any second-level content category among multiple second-level content categories is detected, any second-level content category may be determined as a target content category corresponding to the target image.

[0110] For example, after a user takes a photo through the photo interface to obtain a target image, he or she first selects a caption for the image on the photo interface, and then selects a press release under the caption for the image. In this case, the press release is the target content category corresponding to the target image.

[0111] For ease of understanding, the following Figure 12 To explain, Figure 12On the photo taking interface shown, when the user selects the first-level content category of picture caption, the interface displays multiple second-level content categories under picture caption, such as application a1, press releases, and picture writing, from which the user can select one second-level content category.

[0112] It should be noted that the second-level content category may also include a third-level content category, the third-level content category may include a fourth-level content category, and so on, and there is no limitation to this.

[0113] Step 1105: In response to detecting a text generation operation, a text generation request is sent to the server.

[0114] In this application, step 1105 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0115] Step 1106: Receive the target text sent by the server and display the target text.

[0116] In the present application, step 1106 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0117] In an embodiment of the present application, if a user selects a target first-level content category from among multiple first-level content categories, multiple second-level content categories under the target first-level content category can be displayed. If the user selects a second-level content category, the second-level content category can be used as the target content category corresponding to the target image, so that the server can generate text based on the prompt text corresponding to the second-level content category. In this way, multiple levels of content categories can be provided for user selection, meeting the user's different text generation needs.

[0118] Figure 13 A flowchart of a text generation method provided in another embodiment of the present application.

[0119] like Figure 13 As shown, the text generation method includes:

[0120] Step 1301: Obtain a target image and a target content category corresponding to the target image.

[0121] In this application, step 1301 can be implemented in any of the embodiments of this application, so it will not be described here in detail.

[0122] Step 1302: In response to detecting a text generation operation, a text generation request is sent to a server.

[0123] In this application, step 1302 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0124] Step 1303: Receive the target text sent by the server and display the target text.

[0125] In this application, step 1303 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0126] Step 1304 : Determine candidate modification methods corresponding to the target text according to the target content category, and display the candidate modification methods.

[0127] In this application, candidate modification methods corresponding to the target text can be determined based on the text content characteristics of the target content category. For example, if the target content category is a press release, candidate modification methods may include "more formal tone" or "more concise expression".

[0128] In this application, there may be one or more candidate modification methods corresponding to the target text, and there is no limitation on this.

[0129] For example, the candidate modification methods can be displayed on the display interface of the target text. For example, Figure 7 In the , candidate modification methods can be displayed below the target text.

[0130] Step 1305 : In response to detecting a selection operation on a target modification method among candidate modification methods, the target modification method is sent to a server.

[0131] In this application, if it is detected that the user selects a target modification method from the candidate modification methods, the target modification method can be sent to the server. Furthermore, the target text and the target modification method can be sent to the server. Thus, the server modifies the target text according to the target modification method.

[0132] Exemplarily, the client may send a text modification request to the server, wherein the text modification request may include a target modification method corresponding to the target text.

[0133] Exemplarily, the client may send a text modification request to the server, wherein the text modification request may include a target text and a target modification method corresponding to the target text.

[0134] Step 1306: Receive the modified text sent by the server and display the modified text.

[0135] In this application, the client can receive the modified text sent by the server and display the modified text on the display interface of the target text.

[0136] The modified text may be obtained by the server modifying the target text according to the target modification method. For example, the modified text may be obtained by the server using a large model to modify the target text according to the target modification method.

[0137] In an embodiment of the present application, candidate modification methods for a target text can be displayed. If a user selects a target modification method from the candidate modification methods, the target modification method can be sent to a server, which then modifies the target text based on the target modification method. Thus, the user can be provided with candidate modification methods for the target text. When the user has modification requirements, the user can select the corresponding modification method, thereby satisfying the user's modification requirements for the generated target text.

[0138] Figure 14 A flowchart of a text generation method provided in another embodiment of the present application.

[0139] The text generation method of the embodiment of the present application can be executed by the text generation device of the embodiment of the present application, and the device can be configured in an electronic device, such as a server.

[0140] like Figure 14 As shown, the text generation method includes:

[0141] Step 1401: Receive a text generation request sent by a client.

[0142] The text generation request may include a target image and prompt text corresponding to the target content category.

[0143] The target image can be generated by taking a photo or selected from local images, and there is no limitation on this.

[0144] The target content category may refer to the content category of the text to be generated from the target image.

[0145] Exemplarily, the target content category may be a first-level content category, such as image analysis or image generation PPT.

[0146] For example, the target content category may be a second-level content category under the first-level content category. For example, the target content category may be a second-level content category under the first-level content category of picture captions, such as picture writing.

[0147] Step 1402: Input the target image and prompt text into the macro model to generate target text of the target content category.

[0148] In this application, the server can use a large model to process the target image and prompt text to generate the target text of the target content category.

[0149] The target text of the target content category may refer to the content category of the target text being the target content category.

[0150] For example, a large model can be used to extract target elements that match the target content category from the target image based on the prompt text, and then generate the target text based on the prompt text and the target elements. In this way, the target text can be made more in line with user needs.

[0151] For example, if the target content category is a press release, news-related elements can be extracted from the target image based on the features of the press release, and a press release can be generated based on the extracted news-related elements.

[0152] Exemplarily, a large model can be used to parse the target image using the image prompt text corresponding to the target image to obtain the parsed content of the target image, and based on the prompt text corresponding to the target content category, the target element matching the target content category can be extracted from the target image, and then the target text can be generated based on the prompt text and the target element.

[0153] For example, if the target content category is a presentation, the prompt text and target image corresponding to the presentation can be input into the large model. The large model is then used to process the target image based on the prompt text to generate a presentation text outline. The target text, i.e., the presentation, is then generated based on the presentation text outline. Thus, when the user selects the content category as a presentation, a presentation can be generated based on the image, thereby meeting the user's need for presentation generation.

[0154] Step 1403: Send the target text to the client.

[0155] Exemplarily, after the entire target text is generated, the server may send the entire target text to the client, so that the client displays the entire target text.

[0156] For example, the server may send the generated characters one by one to the client, so that the client displays the generated characters one by one until the last character of the target text.

[0157] In an embodiment of the present application, the server can receive a text generation request sent by the client, and input the target image and the prompt text corresponding to the target content category in the text generation request into the big model. The big model generates the target text of the target content category based on the target image and the prompt text, and sends it to the client for display. Therefore, in the picture-to-text scenario, for different content categories selected by the user, texts of different content categories can be generated for the user based on the prompt texts corresponding to different content categories, thereby meeting different text generation needs.

[0158] In one embodiment of the present application, based on the above embodiment, the server can receive the target modification method corresponding to the target text sent by the client, modify the target text according to the target modification method to obtain the modified text, and send the modified text to the client for display.

[0159] Exemplarily, the server may receive a text modification request sent by a client, wherein the text modification request may include a target modification method corresponding to the target text. Exemplarily, the server may receive a text modification request sent by a client, wherein the text modification request may include a target text and a target modification method corresponding to the target text.

[0160] Exemplarily, the server may input the target modification method and the target text into the large model, and use the large model to modify the target text according to the target modification method to obtain the modified text.

[0161] For example, if the target modification method is "adding a summary paragraph at the end", the server can summarize the target text according to the target modification method and add a summary paragraph at the end, thereby obtaining a modified text corresponding to the target text.

[0162] In the embodiment of the present application, the target text can be modified according to the target modification method to obtain a modified text, thereby meeting the user's modification needs for the generated target text.

[0163] In order to implement the above embodiment, the embodiment of the present application also proposes a text generation device. Figure 15 A schematic diagram of the structure of a text generation device provided in one embodiment of the present application.

[0164] like Figure 15 As shown, the text generation device 1500 includes:

[0165] The acquisition module 1510 is used to acquire a target image and a target content category corresponding to the target image; wherein the target content category is the content category of the text to be generated from the target image;

[0166] The sending module 1520 is configured to send a text generation request to the server in response to detecting a text generation operation; wherein the text generation request includes a target image and a prompt text corresponding to a target content category;

[0167] Receiving module 1530, for receiving the target text sent by the server;

[0168] The display module 1540 is used to display the target text; wherein the target text is a text of the target content category generated by processing the target image and prompt text using the large model.

[0169] Optionally, the acquisition module 1510 is configured to:

[0170] Display multiple first-level content categories;

[0171] In response to detecting a selection operation on a target first-level content category among the plurality of first-level content categories, displaying a plurality of second-level content categories under the target first-level content category;

[0172] In response to detecting a selection operation on any second-level content category, the any second-level content category is determined as a target content category.

[0173] Optionally, the acquisition module 1510 is configured to:

[0174] In response to detecting a triggering operation on the picture answer control, displaying a photo taking interface;

[0175] In response to detecting a triggering operation on a photographing control on the photographing interface, a photographed target image is acquired.

[0176] Optionally, the acquisition module 1510 is configured to:

[0177] In response to detecting a triggering operation on the document assistant control, displaying a document assistant dialog interface;

[0178] In response to detecting a trigger operation of adding a control to a picture on the document assistant dialogue interface, displaying a photo trigger control;

[0179] In response to detecting a trigger operation on a photo trigger control, displaying a photo interface;

[0180] In response to detecting a triggering operation on a photographing control on a photographing interface, a photographed target image is acquired.

[0181] Optionally, the acquisition module 1510 is configured to:

[0182] In response to detecting a selection operation on a target content category on the photographing interface, the target content category is acquired.

[0183] Optionally, the acquisition module 1510 is configured to:

[0184] In response to detecting a triggering operation on a picture answer control on the client interface, displaying a photo taking interface;

[0185] In response to detecting a triggering operation on a picture import control on a photo taking interface, displaying a picture list;

[0186] In response to detecting a selection operation on a target picture in the picture list, the target picture is acquired and a picture display interface of the target picture is displayed.

[0187] Optionally, the acquisition module 1510 is configured to:

[0188] In response to detecting a triggering operation on the document assistant control, displaying a document assistant dialog interface;

[0189] In response to detecting a triggering operation of adding a control to a picture on the document assistant dialogue interface, displaying a picture list;

[0190] In response to detecting a selection operation on a target picture in the picture list, the target picture is acquired and a picture display interface of the target picture is displayed.

[0191] Optionally, the acquisition module 1510 is configured to:

[0192] In response to detecting a selection operation on a target content category on the picture display interface, the target content category is acquired.

[0193] Optionally, the device further comprises:

[0194] A determination module, configured to determine candidate modification methods corresponding to the target text according to the target content category;

[0195] The display module 1540 is further used to display candidate modification methods;

[0196] The sending module 1520 is further configured to send the target modification method to the server in response to detecting a selection operation of the target modification method among the candidate modification methods;

[0197] The receiving module 1530 is further configured to receive the modified text sent by the server;

[0198] The display module 1540 is further configured to display the modified text; wherein the modified text is obtained by modifying the target text based on the target modification method.

[0199] It should be noted that the explanation of the aforementioned client-side text generation method embodiment is also applicable to the text generation device of this embodiment, so it will not be repeated here.

[0200] In an embodiment of the present application, by obtaining a target image and a target content category corresponding to the target image, and when a text generation operation is detected, a text generation request is sent to the server, and the target image and prompt text corresponding to the target content category are sent to the server, so that the server uses a large model to generate a target text of the target content category based on the target image and the prompt text. Therefore, in the text generation scenario, not only can the input of images be supported, the form of input data is expanded, but also, for different content categories selected by the user, texts of different content categories can be generated for the user based on the prompt texts corresponding to different content categories, thereby meeting different text generation needs.

[0201] In order to implement the above embodiment, the embodiment of the present application also proposes a text generation device. Figure 16 A structural diagram of a text generation device provided in another embodiment of the present application.

[0202] like Figure 16 As shown, the text generating device 1600 includes:

[0203] Receiving module 1610, configured to receive a text generation request sent by a client; wherein the text generation request includes a target image and a prompt text corresponding to a target content category, where the target content category is a content category of the text to be generated for the target image;

[0204] A generation module 1620 is used to input the target image and prompt text into the macro model to generate the target text of the target content category;

[0205] The sending module 1630 is configured to send the target text to the client.

[0206] Optionally, the generating module 1620 is configured to:

[0207] Using a large model, we extract target elements that match the target content category from the target image based on the prompt text.

[0208] Generate target text based on prompt text and target element.

[0209] Optionally, the generating module 1620 is configured to:

[0210] In response to the target content category being a presentation, the large model is used to process the target image according to the prompt text to generate a presentation text outline;

[0211] Generate target text based on the presentation text outline.

[0212] Optionally, the receiving module 1610 is further configured to receive a target modification method corresponding to the target text sent by the client;

[0213] The device further includes: a modification module for modifying the target text according to the target modification method using the large model to obtain a modified text;

[0214] The sending module 1630 is further configured to send the modified text to the client.

[0215] It should be noted that the explanation of the aforementioned server-side text generation method embodiment is also applicable to the text generation device of this embodiment, so it will not be repeated here.

[0216] In an embodiment of the present application, the server can receive a text generation request sent by the client, and input the target image and the prompt text corresponding to the target content category in the text generation request into the large model. The large model generates the target text of the target content category based on the target image and the prompt text, and sends it to the client for display. In this way, for different content categories selected by the user, texts of different content categories can be generated for the user based on the prompt texts corresponding to different content categories, thereby meeting different text generation needs.

[0217] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.

[0218] Figure 17 A schematic block diagram of an example electronic device 1700 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0219] like Figure 17 As shown, the device 1700 includes a computing unit 1701, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1702 or a computer program loaded from a storage unit 1708 into a RAM (Random Access Memory) 1703. Various programs and data required for the operation of the device 1700 can also be stored in the RAM 1703. The computing unit 1701, the ROM 1702, and the RAM 1703 are connected to each other via a bus 1704. An I / O (Input / Output) interface 1705 is also connected to the bus 1704.

[0220] Various components in device 1700 are connected to I / O interface 1705, including an input unit 1706, such as a keyboard, mouse, etc.; an output unit 1707, such as various types of displays, speakers, etc.; a storage unit 1708, such as a magnetic disk, optical disk, etc.; and a communication unit 1709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1709 allows device 1700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0221] Computing unit 1701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1701 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. Computing unit 1701 performs the various methods and processes described above, such as the text generation method. For example, in some embodiments, the text generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1700 via ROM 1702 and / or communication unit 1709. When the computer program is loaded into RAM 1703 and executed by computing unit 1701, one or more steps of the text generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1701 may be configured to execute the text generation method in any other appropriate manner (eg, by means of firmware).

[0222] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0223] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0224] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0225] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0226] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0227] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.

[0228] According to an embodiment of the present application, the present application also provides a computer program product, which, when an instruction processor in the computer program product executes, executes the text generation method proposed in the above embodiment of the present application.

[0229] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.

[0230] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A text generation method, comprising: Obtaining a target image and a target content category corresponding to the target image, wherein the target content category is a content category of a text to be generated for the target image, wherein the method for determining the target content category includes: displaying a plurality of first-level content categories; in response to detecting a selection operation of a target first-level content category among the plurality of first-level content categories, displaying a plurality of second-level content categories under the target first-level content category, wherein when the first-level content category is picture caption, the plurality of second-level content categories include press releases, entertainment articles, and / or picture-based writing; and in response to detecting a selection operation of any second-level content category, determining the selected second-level content category as the target content category; In response to detecting the text generation operation, sending the text generation request to the server, wherein the text generation request includes prompt text corresponding to the target image and the target content category, and the prompt text is used to prompt the large model to perform the image-to-text task; Receive the target text sent by the server and display the target text; wherein the target text is a text of the target content category generated by processing the target image and the prompt text using a large model.

2. The method according to claim 1, wherein The acquiring of the target image includes: In response to detecting a triggering operation on the picture answer control, displaying a photo taking interface; In response to detecting a triggering operation on a photo control on the photo interface, the captured target image is acquired.

3. The method according to claim 1, wherein The acquiring of the target image includes: In response to detecting a triggering operation on the document assistant control, displaying a document assistant dialog interface; In response to detecting a trigger operation of adding a control to a picture on the document assistant dialogue interface, displaying a photo trigger control; In response to detecting a trigger operation on the photo trigger control, displaying the photo interface; In response to detecting a triggering operation on a photo control on the photo interface, the captured target image is acquired.

4. The method according to claim 2 or 3, wherein: The method for determining the target content category further includes: In response to detecting a selection operation on the target content category on the photographing interface, the target content category is acquired.

5. The method according to claim 1, wherein The acquiring of the target image includes: In response to detecting a triggering operation on a picture answer control on the client interface, displaying a photo taking interface; In response to detecting a triggering operation on a picture import control on the photographing interface, displaying a picture list; In response to detecting a selection operation on the target image in the image list, the target image is acquired and an image display interface of the target image is displayed.

6. The method of claim 1, wherein: The acquiring of the target image includes: In response to detecting a triggering operation on the document assistant control, displaying a document assistant dialog interface; In response to detecting a triggering operation of adding a picture control on the document assistant dialogue interface, displaying a picture list; In response to detecting a selection operation on the target image in the image list, the target image is acquired and an image display interface of the target image is displayed.

7. The method according to claim 5 or 6, wherein: The obtaining of the target content category corresponding to the target image includes: In response to detecting a selection operation on the target content category on the picture display interface, the target content category is acquired.

8. The method of claim 1 , further comprising: Determining candidate modification methods corresponding to the target text according to the target content category, and displaying the candidate modification methods; In response to detecting a selection operation of a target modification method among the candidate modification methods, sending the target modification method to the server; Receive the modified text sent by the server and display the modified text; wherein the modified text is obtained by modifying the target text based on the target modification method.

9. A text generation method comprising: Receiving a text generation request determined according to the method according to any one of claims 1 to 8, sent by a client; wherein the text generation request includes a target image and a prompt text corresponding to a target content category, and the target content category is a content category of the text to be generated for the target image; Input the target image and the prompt text into a large model to generate a target text of the target content category; The target text is sent to the client.

10. The method of claim 9, wherein: The step of inputting the target image and the prompt text into a large model to generate the target text for the target content category pair includes: Using the large model, and according to the prompt text, extracting target elements matching the target content category from the target image; The target text is generated according to the prompt text and the target element.

11. The method of claim 9, wherein: The step of inputting the target image and the prompt text into a large model to generate a target text with the target content category includes: In response to the target content category being a presentation, using the large model, processing the target image according to the prompt text to generate a presentation text outline; The target text is generated according to the presentation text outline.

12. The method according to any one of claims 9 to 11, further comprising: receiving a target modification method corresponding to the target text sent by the client; Using the large model, modifying the target text according to the target modification method to obtain a modified text; The modified text is sent to the client.

13. A text generation device, comprising: An acquisition module is configured to acquire a target image and a target content category corresponding to the target image, wherein the target content category is a content category of a text to be generated for the target image, wherein a method for determining the target content category comprises: displaying a plurality of first-level content categories; in response to detecting a selection operation of a target first-level content category among the plurality of first-level content categories, displaying a plurality of second-level content categories under the target first-level content category, wherein when the first-level content category is picture caption, the plurality of second-level content categories include press releases, entertainment articles, and / or picture-based writing; and in response to detecting a selection operation of any second-level content category, determining the selected second-level content category as the target content category; a sending module configured to, in response to detecting a text generation operation, send the text generation request to the server, wherein the text generation request includes prompt text corresponding to the target image and the target content category, and the prompt text is used to prompt the large model to perform the image-to-text task; A receiving module is used to receive the target text sent by the server; A display module is used to display the target text; wherein the target text is a text of the target content category generated by processing the target image and the prompt text using a large model.

14. The apparatus of claim 13, wherein: The acquisition module is used to: In response to detecting a triggering operation on the picture answer control, displaying a photo taking interface; In response to detecting a triggering operation on a photo control on the photo interface, the captured target image is acquired.

15. The apparatus of claim 13, wherein: The acquisition module is used to: In response to detecting a triggering operation on the document assistant control, displaying a document assistant dialog interface; In response to detecting a trigger operation of adding a control to a picture on the document assistant dialogue interface, displaying a photo trigger control; In response to detecting a trigger operation on the photo trigger control, displaying the photo interface; In response to detecting a triggering operation on a photo control on the photo interface, the captured target image is acquired.

16. The device according to claim 14 or 15, wherein The acquisition module is used to: In response to detecting a selection operation on the target content category on the photographing interface, the target content category is acquired.

17. The apparatus of claim 13, wherein: The acquisition module is used to: In response to detecting a triggering operation on a picture answer control on the client interface, displaying a photo taking interface; In response to detecting a triggering operation on a picture import control on the photographing interface, displaying a picture list; In response to detecting a selection operation on the target image in the image list, the target image is acquired and an image display interface of the target image is displayed.

18. The apparatus of claim 13, wherein: The acquisition module is used to: In response to detecting a triggering operation on the document assistant control, displaying a document assistant dialog interface; In response to detecting a triggering operation of adding a picture control on the document assistant dialogue interface, displaying a picture list; In response to detecting a selection operation on the target image in the image list, the target image is acquired and an image display interface of the target image is displayed.

19. The device according to claim 17 or 18, wherein The acquisition module is used to: In response to detecting a selection operation on the target content category on the picture display interface, the target content category is acquired.

20. The apparatus of claim 13, further comprising: A determination module, configured to determine candidate modification methods corresponding to the target text according to the target content category; The display module is further configured to display the candidate modification methods; The sending module is further configured to send the target modification method to the server in response to detecting a selection operation of the target modification method among the candidate modification methods; The receiving module is further configured to receive the modified text sent by the server; The display module is further configured to display the modified text; wherein the modified text is obtained by modifying the target text based on the target modification method.

21. A text generation device, comprising: a receiving module, configured to receive a text generation request determined according to the method according to any one of claims 1 to 8, sent by a client; wherein the text generation request includes a target image and prompt text corresponding to a target content category, and the target content category is a content category of the text to be generated for the target image; A generation module, configured to input the target image and the prompt text into a macro model to generate a target text of the target content category; The sending module is used to send the target text to the client.

22. The apparatus of claim 21, wherein: The generating module is used to: Using the large model, and according to the prompt text, extracting target elements matching the target content category from the target image; The target text is generated according to the prompt text and the target element.

23. The apparatus of claim 21, wherein: The generating module is used to: In response to the target content category being a presentation, using the large model, processing the target image according to the prompt text to generate a presentation text outline; The target text is generated according to the presentation text outline.

24. The device according to any one of claims 21 to 23, The receiving module is further configured to receive a target modification method corresponding to the target text sent by the client; The device further comprises: A modification module, configured to modify the target text according to the target modification method using the large model to obtain a modified text; The sending module is further configured to send the modified text to the client.

25. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

26. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.

27. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.