Content generation method and electronic device

By displaying multiple pictures entered by the user in the content generation method, determining the target application and style, and generating and displaying related content, the problem of difficult to control and complicated copy quality in the prior art is solved, and high-quality content generation and convenient operation are achieved.

CN119941926APending Publication Date: 2025-05-06SAMSUNG GUANGZHOU MOBILE R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510014552.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing multimodal artificial intelligence content generation technology is difficult to control the quality of copywriting, and users are cumbersome to operate and cannot fully meet the needs of users.

Method used

It provides a content generation method, by displaying multiple pictures input by the user on the content generation function interface, determining the target application and target generation style, generating associated content in response to content generation instructions, and displaying multiple pictures and associated content in the specified interface of the target application.

Benefits of technology

It realizes that users can input multiple pictures at the same time and intelligently generate high-quality associated content, which is convenient to operate and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941926A_ABST
    Figure CN119941926A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content generation method and an electronic device, and relates to the field of artificial intelligence. The method comprises the following steps: displaying or partially displaying a plurality of pictures input by a user through a preset mode on a content generation function interface; determining a target application and a target generation style; in response to a content generation instruction, generating associated content with the target generation style of the plurality of pictures; and in response to a sending instruction, displaying the plurality of pictures and the associated content on a specified interface of the target application. Optionally, the method performed by the electronic device may be performed using an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more specifically, to a content generation method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the development of artificial intelligence technology, especially the development of multimodal artificial intelligence generated content (AIGC) technology, it has brought a lot of convenience to people's lives. At the same time, with the development of smart terminals and the continuous enrichment of applications, people are sharing or publishing graphic content through various applications (such as social media apps, review apps, food delivery apps, shopping apps, etc.) more and more. Therefore, using AIGC technology to automatically generate text based on pictures can improve people's work efficiency.

[0003] However, existing technologies still have certain limitations, such as difficulty in controlling the quality of copywriting and cumbersome user operations, making it difficult to fully meet user needs. Summary of the invention

[0004] According to a first aspect of the present disclosure, a content generation method is provided, which may include: displaying or partially displaying a plurality of pictures input by a user in a preset manner on a content generation function interface; determining a target application and a target generation style; generating associated content of the plurality of pictures having the target generation style in response to a content generation instruction; and displaying the plurality of pictures and the associated content on a designated interface of the target application in response to a sending instruction.

[0005] Optionally, the preset method for the user to input the multiple pictures includes one of the following: inputting the multiple pictures selected by the user into the content generation function interface by sharing; or inputting the multiple pictures selected by the user into the content generation function interface by dragging; or inputting the multiple pictures selected by the user into the content generation function interface by voice commands.

[0006] Optionally, determining the target generation style includes: determining the target generation style based on a user operation; or determining the target generation style based on an analysis of the multiple images.

[0007] Optionally, in response to a content generation instruction, associated content of the multiple pictures having the target generation style is generated, including: in response to the content generation instruction, preprocessing the multiple pictures to obtain a picture with a preset size; and generating associated content of the multiple pictures having the target generation style based on the picture of the preset size and the target generation style.

[0008] Optionally, the multiple pictures are preprocessed to obtain a picture of a preset size, including: splicing the multiple pictures, adjusting the spliced ​​picture to the preset size, and obtaining the picture of the preset size; or adjusting the size of each picture in the multiple pictures based on the preset size, and splicing each adjusted picture to obtain the picture of the preset size.

[0009] Optionally, after obtaining the picture of the preset size, the method further includes: adjusting the compression rate of the picture of the preset size to a preset compression rate; and generating the associated content based on the image of the preset size with the preset compression rate.

[0010] Optionally, in response to sending an instruction, displaying the multiple pictures and the associated content on a designated interface of a target application includes: simultaneously sending identification information corresponding to the multiple pictures and the associated content to the target application, so that the multiple pictures and the associated content are displayed on the designated interface of the target application.

[0011] Optionally, in response to sending an instruction, displaying the multiple pictures and the associated content on a designated interface of a target application includes: sending identification information corresponding to the multiple pictures to the target application, so that the multiple pictures are displayed on the designated interface of the target application; triggering a content editing function on the designated interface, and copying the associated content to the designated interface, so that the multiple pictures and the associated content are displayed on the designated target interface of the target application.

[0012] Optionally, in response to sending an instruction, displaying the multiple pictures and the associated content in a designated interface of a target application includes: creating a specific picture category for the multiple images; creating a virtual screen and launching the target application in the virtual screen; triggering a picture selection function of the target application in the virtual screen, and based on the specific picture category, displaying the multiple pictures in the designated interface of the target application in the virtual screen; triggering a content editing function on the designated interface in the virtual screen, and copying the associated content to the designated interface, so that the multiple pictures and the associated content are displayed in the designated interface of the target application in the virtual screen; switching the displayed content of the virtual screen to the main screen display, so that the multiple pictures and the associated content are displayed in the designated interface of the target application on the main screen.

[0013] Optionally, the image selection function of the target application is triggered in the virtual screen, and based on the specific image classification, the multiple images are displayed in the designated interface of the target application in the virtual screen, including: in response to triggering the image selection function of the target application in the virtual screen, entering the image selection interface of the target application; based on the specific image classification, selecting the multiple images from the image selection interface; triggering entry into the designated interface of the target application, and displaying the multiple images in the designated interface.

[0014] According to a second aspect of the present disclosure, a content generation device is provided, including: an input module: configured to: display or partially display multiple pictures input by a user in a preset manner on a content generation function interface; and determine a target application and a target generation style; a content generation module, configured to: generate associated content of the multiple pictures with the target generation style in response to a content generation instruction; and a sending module, configured to: display the multiple pictures and the associated content on a designated interface of the target application in response to a sending instruction.

[0015] Optionally, the input module is configured to: input the multiple pictures selected by the user to the content generation function interface by sharing; or input the multiple pictures selected by the user to the content generation function interface by dragging.

[0016] Optionally, the input module is configured to: determine the target generation style based on user operation; or determine the target generation style based on analysis of the multiple images.

[0017] Optionally, the content generation module is configured to: in response to a content generation instruction, pre-process the multiple images to obtain a picture with a preset size; and generate associated content with the target generation style for the multiple images based on the picture of the preset size and the target generation style.

[0018] Optionally, the content generation module is configured to: splice the multiple pictures, adjust the spliced ​​pictures to the preset size, and obtain the pictures of the preset size; or adjust the size of each picture in the multiple pictures based on the preset size, and splice each adjusted picture to obtain the pictures of the preset size.

[0019] Optionally, the content generation module is configured to: adjust the compression rate of the picture of the preset size to a preset compression rate; and generate the associated content based on the image of the preset size with the preset compression rate.

[0020] Optionally, the sending module is configured to: send identification information respectively corresponding to the multiple pictures and the associated content to the target application at the same time, so that the multiple pictures and the associated content are displayed on the designated interface of the target application.

[0021] Optionally, the sending module is configured to: send identification information corresponding to the multiple pictures respectively to the target application, so that the multiple pictures are displayed on the designated interface of the target application; trigger the content editing function on the designated interface, and copy the associated content to the designated interface, so that the multiple pictures and the associated content are displayed on the designated target interface of the target application.

[0022] Optionally, the sending module is configured to: create a specific image classification for the multiple images; create a virtual screen and start the target application in the virtual screen; trigger the image selection function of the target application in the virtual screen, and based on the specific image classification, display the multiple images in the designated interface of the target application in the virtual screen; trigger the content editing function on the designated interface in the virtual screen, and copy the associated content to the designated interface, so that the multiple images and the associated content are displayed in the designated interface of the target application in the virtual screen; switch the display content of the virtual screen to the main screen display, so that the multiple images and the associated content are displayed in the designated interface of the target application on the main screen.

[0023] Optionally, the sending module is configured to: enter the picture selection interface of the target application in response to triggering the picture selection function of the target application in the virtual screen; select the multiple pictures from the picture selection interface based on the specific picture classification; trigger entry into the designated interface of the target application, and display the multiple pictures in the designated interface.

[0024] According to a third aspect of the present disclosure, an electronic device is provided, which may include: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to execute a method of an exemplary embodiment of the present disclosure.

[0025] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program or instructions are stored. When the computer program or instructions are executed by at least one processor, the at least one processor executes the method of the exemplary embodiment of the present disclosure.

[0026] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method of the exemplary embodiments of the present disclosure when executed by a processor.

[0027] The embodiments of the present disclosure enable users to input multiple pictures at the same time and intelligently generate high-quality related content based on the multiple pictures, which is convenient to operate and improves user experience.

[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings herein are incorporated in and constitute a part of the specification, illustrate exemplary embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0030] Figure 1 is a flowchart of a content generating method according to an exemplary embodiment of the present disclosure.

[0031] Figures 2 to 5 is a schematic diagram of inputting multiple pictures into a content generation function interface according to an exemplary embodiment of the present disclosure.

[0032] Figures 6 to 11 is a schematic diagram of preprocessing multiple images according to an exemplary embodiment of the present disclosure.

[0033] Fig.12 is a schematic diagram of generating associated content according to an exemplary embodiment of the present disclosure.

[0034] Fig.13 and Fig.14 is a schematic diagram of a first sending operation according to an exemplary embodiment of the present disclosure.

[0035] Fig.15 and Fig.16 is a schematic diagram of a second sending operation according to an exemplary embodiment of the present disclosure.

[0036] Fig.17 and Fig.18 is a schematic diagram of a third sending operation according to an exemplary embodiment of the present disclosure.

[0037] Fig.19 is a schematic diagram illustrating publishing of multiple pictures and related contents according to an exemplary embodiment of the present disclosure.

[0038] Fig. 20 is a block diagram illustrating a content generating apparatus according to an exemplary embodiment of the present disclosure.

[0039] Fig.21is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure.

[0040] Fig. 22 A schematic structural diagram of an electronic device to which an embodiment of the present disclosure is applicable is shown. DETAILED DESCRIPTION

[0041] The following description with reference to the accompanying drawings is provided to facilitate a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should be considered as exemplary only. Therefore, one of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, for the sake of clarity and conciseness, descriptions of well-known functions and structures may be omitted.

[0042] The terms and expressions used in the following specification and claims are not limited to their dictionary meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustration purposes only and not for the purpose of limiting the present disclosure as defined in the appended claims and their equivalents.

[0043] It should be understood that the singular forms "a", "an", and "the" may also include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more such surfaces. When we refer to an element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or the one element and the other element may establish a connection relationship through an intermediate element. In addition, "connected" or "coupled" as used herein may include wireless connection or wireless coupling.

[0044] The term "include" or "may include" refers to the presence of the corresponding disclosed functions, operations or components that can be used in various embodiments of the present disclosure, rather than limiting the presence of one or more additional functions, operations or features. In addition, the term "include" or "have" may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components or combinations thereof, but should not be interpreted as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components or combinations thereof.

[0045] The term "or" used in various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple, or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" may be implemented as parameter A including A1 or A2 or A3, or may be implemented as parameter A including at least two of the three items A1, A2, and A3.

[0046] Unless defined differently, all terms (including technical terms or scientific terms) used in the present disclosure have the same meanings as understood by those skilled in the art to which the present disclosure belongs. Common terms as defined in dictionaries are interpreted as having meanings consistent with the context in the relevant technical field, and should not be interpreted ideally or overly formally unless explicitly defined in the present disclosure.

[0047] At least some functions of the device or electronic device provided in the embodiments of the present disclosure can be implemented by an AI model, such as at least one module among multiple modules of the device or electronic device can be implemented by an AI model. Functions associated with AI can be performed by non-volatile memory, volatile memory and processor.

[0048] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or pure graphics processing units, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor, such as a neural processing unit (NPU).

[0049] The one or more processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are provided by training or learning.

[0050] Here, providing by learning means obtaining a predefined operating rule or an AI model with desired characteristics by applying a learning algorithm to a plurality of learning data. The learning can be performed in the device or electronic device itself in which the AI ​​according to the embodiment is executed, and / or can be implemented by a separate server / system.

[0051] The AI ​​model may include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations by calculating between the input data of the layer (such as the calculation results of the previous layer and / or the input data of the AI ​​model) and the multiple weight values ​​of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.

[0052] A learning algorithm is a method of using a plurality of learning data to train a predetermined target device (e.g., a robot) to enable, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0053] The method provided in the present disclosure may involve one or more technical fields such as speech, language, image, video or data intelligence.

[0054] Optionally, when it comes to the field of speech or language, in the content generation method according to the present disclosure, a speech signal as an analog signal can be received via a speech input device (e.g., a microphone), and the speech portion can be converted into a computer-readable text using an automatic speech recognition (ASR) model. The user's speech intention can be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or the NLU model can be an artificial intelligence model. The artificial intelligence model can be processed by an artificial intelligence dedicated processor designed in a hardware structure specified for artificial intelligence model processing. Language understanding is a technology for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.

[0055] Optionally, when it comes to the field of images or videos, in the content generation method according to the present disclosure, output data can be obtained by using image data as input data of an artificial intelligence model. The method of the present disclosure may be related to the field of visual understanding of artificial intelligence technology, which is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.

[0056] Optionally, when it comes to the field of intelligent data processing, in the content generation method according to the present disclosure, in the inference or prediction stage, an artificial intelligence model can be used to perform predictions by using real-time input data. The processor of the electronic device can perform preprocessing operations on the data to convert it into a form suitable for use as an input to the artificial intelligence model. Inference prediction is a technology for logical reasoning and prediction by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning or recommendation.

[0057] In the present application, an artificial intelligence model can be obtained by training. Here, "obtained by training" means obtaining a predefined operating rule or artificial intelligence model configured to perform a desired feature (or purpose) by training a basic artificial intelligence model with multiple training data through a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and the neural network calculation is performed by calculating between the calculation result of the previous layer and the multiple weight values.

[0058] At present, the intelligent solution for generating copy based on multiple pictures has the problem of low quality of generated copy, especially when the picture content is relatively complex. The reason is that the current intelligent solution analyzes the entire publishing interface after uploading the picture, that is, the picture obtained by the model is a screenshot of the current entire interface or the picture in the current entire interface layout view, which will cause each uploaded picture to be small in size and low in clarity, and the picture quality will directly affect the quality of the copy generated by the model. In addition, the operation of the current intelligent solution is cumbersome and inconvenient for users.

[0059] Based on this, the present disclosure is designed to enable users to input multiple pictures at the same time and intelligently generate high-quality related content based on the multiple pictures. Figures 1 to 22 Embodiments of the present disclosure are described in detail.

[0060] Figure 1 is a flowchart of a content generating method according to an exemplary embodiment of the present disclosure. Figure 1 The method shown can be performed by the application / logic / circuit / module / unit for generating associated content based on a picture disclosed in the present invention. The application / logic / circuit / module / unit for generating associated content based on a picture disclosed in the present invention can be installed in any electronic device / terminal. In the following, the application / logic / circuit / module / unit for generating associated content based on a picture may be referred to as a content generation application or other names.

[0061] Reference Figure 1In step S101, multiple pictures input by the user in a preset manner are displayed or partially displayed on the content generation function interface. Here, the content generation function interface may be the interface of the content generation application of the present disclosure. The user may activate or enable the content generation application in some way and input multiple pictures selected from any other application interface into the content generation function interface. In the present disclosure, content may be used interchangeably with copywriting.

[0062] As an example, the preset method for a user to input multiple pictures into the content generation function interface may be: inputting multiple pictures selected by the user into the content generation function interface by sharing; or inputting multiple pictures selected by the user into the content generation function interface by dragging.

[0063] In the present disclosure, a user can select pictures from any application interface. For example, a user selects multiple pictures in a certain application interface and activates the sharing function to input the multiple pictures into the content generation function interface; or a user selects multiple pictures in a certain application interface and inputs the multiple pictures into the content generation function interface by dragging.

[0064] For example, a user can select multiple pictures that he wants to share in any application interface of the terminal, such as the photo album, my files, browser page, etc., and then send the selected pictures to the content generation application of the present invention through a specific interface / entrance and display them in the content generation function interface.

[0065] In the scenario where the user selects pictures from the album, the user can select multiple pictures in the album that need to be sent to the target application, click Share, and then multiple application icons will be displayed. The user can select the content generation application icon from them and enter the content generation function interface. In the content generation function interface, multiple pictures can be displayed or partially displayed, such as Figure 2 shown.

[0066] In the scenario where the user selects pictures from My Files, the user can select multiple pictures to be sent to the target application in My Files, click Share, and then multiple application icons will be displayed. Select the content generation application icon and enter the content generation function interface, where multiple pictures can be displayed or partially displayed, such as Figure 3 shown.

[0067] For another example, in the scenario of using screen search functions, select multiple pictures that need to be sent to the target application on the screen, click Share, and then icons of multiple applications will be displayed. Select the content generation application and enter the content generation function interface, such as Figure 4 As shown in the figure, users can start screen search applications through corresponding shortcut launch methods and perform a series of operations related to screen content, such as searching and sharing.

[0068] For another example, in the scenario of using the smart drag and drop function, the user selects multiple pictures to be sent to the target application and performs a drag / drop operation. At this time, the sidebar will display multiple application icons. Continue to drag the pictures to the content generation application icon to enter the content generation function interface, such as Figure 5 As shown in the figure, users can use the smart drag-and-drop function to quickly complete related operations, such as sharing and jumping, by dragging the content or part of the content on the screen to the edge of the screen.

[0069] For another example, in the scenario of using a voice assistant function, the user selects multiple pictures that need to be sent to the target application, sends the selected multiple pictures to the content generation application through voice commands, and enters the content generation function interface. For example, after the user selects multiple pictures that he wants to share in any application interface such as the photo album, my files, browser page, etc. in the terminal, the user can send a voice command such as "input the selected pictures to the content generation application" or other specific voice commands to the terminal, and the terminal can input the selected pictures into the content generation application in response to the voice command.

[0070] The above-mentioned manner of inputting multiple pictures into the content generation application is merely exemplary, and the present disclosure is not limited thereto.

[0071] In step S102, a target application and a target generation style are determined.

[0072] As an example, the user can select a target application from multiple applications through the content generation function interface. For example, in the content generation function interface, multiple icons or titles of selected applications can be displayed to the user in the form of a drop-down menu for the user to select.

[0073] When determining the target generation style, the target generation style may be determined based on user operations, or may be determined based on analysis of multiple input images.

[0074] As an example, the target generation style can be determined based on the user's selection of multiple preset generation styles in the content generation function interface. For example, in the content generation function interface, the names of multiple preset generation styles can be displayed to the user in the form of a drop-down menu for the user to choose. As another example, the user can customize the target generation style by input (text input or voice input). As another example, the content generation application can determine the target generation style based on the analysis of multiple input pictures, such as the content generation application can intelligently determine a target generation style based on the content of multiple input pictures. When there are at least two of the user selecting the target generation style, the user inputting the target generation style, and the intelligent determination of the target generation style, the target generation style selected by the user or the target generation style input by the user can be used as the final target generation style.

[0075] In step S103, in response to the content generation instruction, associated content of the plurality of images having a target generation style is generated.

[0076] The content generation application of the present disclosure may include a neural network model for generating content based on a picture. The neural network model may generate content associated with multiple pictures based on an input picture, a target generation style, and / or a target application.

[0077] After determining the target application and the target generation style, the content generation application may generate associated content based on the multiple images input by the user, the target application, and the target generation style.

[0078] In order to reduce the time of generating content, the input multiple pictures may be preprocessed first. According to an embodiment of the present disclosure, the preprocessing may include at least one of adjusting the size / resolution of the pictures, splicing the pictures, and adjusting the compression rate of the pictures. Adjusting the compression rate of the pictures refers to adjusting the file size of the pictures without changing the size / resolution of the pictures.

[0079] According to an embodiment of the present disclosure, multiple input images may be preprocessed to obtain an image of a preset size, and then associated content may be generated based on the image of the preset size, a target generation style and / or a target application.

[0080] The multiple input images can be preprocessed based on the resolution / size required by the neural network model to process the images, so that the images are input at the resolution / size that best suits the model.

[0081] The resolution required by the most suitable model refers to the optimal comprehensive effect of identifying the content in the picture, generating the quality of the copy, and the time consumption when the picture is input into the neural network model at this resolution. When training the neural network model, the corresponding training set will be used, and the image resolution of the training set is usually within a certain range. Therefore, if the resolution of an oversized picture is modified to the optimal resolution range, it will not have a significant impact on identifying the content in the picture and generating the quality of the copy, and at the same time it can reduce the time consumption required to generate the copy. The resolution required by the most suitable model for different neural network models is not the same, and can be adaptively determined according to the specific neural network model used.

[0082] There are many ways to preprocess image resolution / size. Different preprocessing methods can be used in different scenarios. The following is an example of preprocessing.

[0083] The size of each of the multiple input pictures can be adjusted based on the preset size, and each adjusted picture can be spliced ​​to obtain a picture of the preset size; or the multiple input pictures can be spliced, and the spliced ​​pictures can be adjusted to the preset size to obtain a picture of the preset size.

[0084] For example, all input images are modified in resolution (i.e., image size is adjusted), and then all images (or some of them) are spliced. The splicing method can be horizontal splicing, vertical splicing, nine-grid splicing, or different images can be assigned different layout weights and then spliced ​​based on the layout weights.

[0085] Assume that the optimal input resolution of the neural network model used is a×b. When the user selects 1 picture, if the length or width of the picture is greater than a and b respectively, the picture resolution is adjusted to a×b. If the length of the picture resolution is less than a, the picture length can be retained at the original size; similarly, if the width of the picture resolution is less than b, the picture width can be retained at the original size. When the user selects 2 pictures, the length of the original resolution of a picture is less than a / 2, the picture length can be retained at the original size; similarly, if the width of the original resolution of a picture is less than b / 2, the picture width can be retained at the original size. If the length and width of a picture are greater than a / 2 and b / 2 respectively, the picture resolution is adjusted to (a / 2)×(b / 2). And so on.

[0086] Figure 6 Two methods of modifying resolution according to an example embodiment of the present disclosure are shown. When the optimal input resolution of the neural network model is a×b, when the user inputs different numbers of pictures, the input pictures are modified and spliced ​​in different ways. Figure 6 The left side (a) shows a way to modify the resolution, and Figure 6 The right side (b) shows another way to modify the resolution. Figure 6 In , each block represents the maximum resolution of each image after resizing.

[0087] When the user selects 9 pictures (picture a to picture i), the resolution of each picture is different. The nine-square grid stitching method can be used. Assuming that the optimal input resolution of the neural network model used is a×b, the resolution of each picture is adjusted to a maximum of (a / 3)×(b / 3), and then stitched from left to right and from top to bottom, such as Figure 7 If the original resolution length of a picture is less than (a / 3), the picture length will remain the same; similarly, if the original resolution width of a picture is less than (b / 3), the picture width will remain the same. Figure 7 In , you can optionally adjust the compression rate of each image after adjusting the size of each image, and then stitch them into a single image.

[0088] For another example, the user selected 9 pictures (picture a to picture i), and the resolution of each picture is different. Use the nine-square stitching method to stitch each picture in its original size. For example, arrange pictures a, b, and c from left to right in the first row, arrange pictures d, e, and f in the second row, and arrange pictures g, h, and i in the third row, and stitch them in order. Then adjust the stitched pictures to the most suitable resolution / size a×b for the model, such as Figure 8 In addition, after adjusting the size of the spliced ​​pictures, the compression rate of the spliced ​​pictures can be selectively adjusted.

[0089] According to an embodiment, the size of each picture can be adjusted according to a preset size based on the number of multiple pictures input, the width and height values ​​of each picture in the multiple pictures, and at least one of the weights of each picture, and then each adjusted picture can be spliced ​​to obtain a picture of a preset size.

[0090] For example, layout weights can be assigned to multiple input images, and the resolution of each image can be modified according to the layout weights, and then spliced, such as Fig. 9 As shown. Fig. 9 In the example, picture a is assigned the largest layout weight, while pictures b to i are assigned the same smaller layout weight. Pictures with larger weights can occupy a larger proportion of the spliced ​​picture. Therefore, the resolution of picture a can be adjusted to a relatively larger resolution. After splicing the pictures, the compression rate of the spliced ​​pictures can be optionally adjusted.

[0091] For another example, several images can be selected from multiple input images for resolution modification and splicing processing, such as Fig.10 As shown. Fig.10 In the example, four images are selected from the nine input images for resolution modification and splicing. After splicing the images, the compression rate of the spliced ​​images can be optionally adjusted.

[0092] For another example, the resolution of images whose height or width is greater than a certain threshold can be modified by modifying the width value (or height value), then stitching them together, and then modifying the resolution of the stitched image to the target resolution. Fig.11 As shown, assuming that the optimal input resolution of the neural network model used is a×b, when the user inputs 9 pictures (picture a to picture i), the threshold for height can be a / 3, and the threshold for width can be b / 3. The height and width of pictures a to picture i are adjusted to a / 3 or b / 3.

[0093] After obtaining the spliced ​​image, the compression rate of the spliced ​​image can be selectively adjusted, for example, the compression rate of the image of a preset size is adjusted to a preset compression rate.

[0094] In addition, the compression rate of each image may be adjusted to a different value before stitching the images.

[0095] By adjusting the image resolution and compression rate, the time consumption requirements can be met while the quality requirements of the generated copy are met.

[0096] As an example, after the user selects the target application and target generation style for the generated text, the preprocessed image (or multiple original images) is passed to the neural network model, which combines the input image and the prompt words corresponding to the selected target application and target generation style to output the corresponding text. Fig.12 As shown, after selecting the target application (APP1) and the target generation style (Style 1), click the "Generate" button to output the corresponding copy. After the copy is generated, if the user is not satisfied with the content of the currently generated copy, the user can reselect the target generation style and then click the "Generate" button again to generate the related content again until the user is satisfied. Fig.12 The target generation styles and applications shown are merely exemplary, and the present disclosure is not limited thereto.

[0097] In step S104, in response to the sending instruction, a plurality of pictures and associated contents are displayed on a designated interface of the target application.

[0098] A corresponding sending operation may be adopted according to the type of the target application, so as to display the multiple pictures and related contents selected by the user on the designated interface of the target application.

[0099] For a target application that supports receiving multiple images + text, identification information (e.g., uniform resource identifier URI) corresponding to the input multiple images and the generated associated content can be sent to the target application at the same time, so that the multiple images and associated content are displayed on a specified interface of the target application, such as Fig.13 In the present disclosure, instead of sending the image to the target application, the identification information of the image (eg, uniform resource identifier URI) is sent to the target application, so that the target application can find the corresponding original image according to the image identification information, so that the image posted by the user is clear.

[0100] For example, when a user inputs a send command through the content generation function interface, the content generation application can start the target application through the application or logic (such as intent) used for communication between components, and send the URIs of multiple pictures selected by the user and the text to the target application in the intent. After receiving the pictures and text, the target application directly loads and displays them on the publishing interface of the target application. The user can publish to the social platform with one click, such as Fig.14 shown.

[0101] For a target application that supports receiving multiple images but cannot receive text at the same time, the images corresponding to the multiple images can be sent to the target application respectively, so that the multiple images are displayed on the specified interface of the target application; the content editing function on the specified interface is triggered, and the related content is copied to the specified interface, so that the multiple images and related content are displayed on the specified target interface of the target application, such as Fig.15 shown.

[0102] For example, the content generation application can start the target application through intent, and send the image URI selected by the user to the target application in the intent. After receiving the image, the target application directly loads and displays it on the publishing interface of the target application. The content generation application uses the simulated click function to click the edit text position on the publishing interface of the target application and paste the text into the edit text box. The user clicks publish to publish to the social platform with one click, such as Fig.16 shown.

[0103] For target applications that only support receiving a single image + text, you can create a specific image classification for the multiple images input, create a virtual screen and start the target application in the virtual screen; trigger the image selection function of the target application in the virtual screen, and based on the specific image classification, display multiple images in the specified interface of the target application in the virtual screen; trigger the content editing function on the specified interface in the virtual screen, and copy the related content to the specified interface, so that multiple images and related content are displayed in the specified interface of the target application in the virtual screen, and then switch the display content of the virtual screen to the main screen display, so that multiple images and related content are displayed in the specified interface of the target application on the main screen. Fig.17 As shown, the content generation application can call the smart classification application (such as a photo function that can provide classification for third-party application pictures to help users easily manage and find photos in the mobile phone album), and create a specific picture category for the multiple input pictures in the smart classification application. The target application can be simulated to click on the virtual screen to enter the picture selection interface, at which time the various picture categories output by the smart classification application (including the above-mentioned specific picture categories) can be displayed. The specific picture category can be simulated in the virtual screen, and the related content can be copied to the specified interface, so that the selected multiple pictures and related content can be displayed in the specified interface of the virtual screen. Next, the display content of the virtual screen can be switched to the main screen display, so that multiple pictures and related content are displayed in the specified interface of the target application on the main screen.

[0104] As another example, a specific picture classification may be created for multiple images, a virtual screen may be created and a target application may be started in the virtual screen, and a designated interface of the target application, such as a publishing interface, may be triggered in the virtual screen. A picture selection function may be triggered via the designated interface of the virtual screen, and multiple pictures may be displayed in a designated interface in the virtual screen based on the specific picture classification. A content editing function on a designated interface may be triggered in the virtual screen, and associated content may be copied to a designated interface, so that multiple pictures and associated content are displayed in a designated interface in the virtual screen. The displayed content of the virtual screen may be switched to the main screen display, so that multiple pictures and associated content are displayed in a designated interface of the target application on the main screen.

[0105] For example, the content generation application can use the smart classification application to create a specific image category for the images selected by the user, and store the URI of the images selected by the user under the specific image category. The specific image category here can be a temporary category. After the sending is completed, the smart classification application can delete the specific image category. The content generation application creates a virtual screen, simulates clicking on the target application and starts it on the virtual screen. Continue to simulate clicking, add the image button, and enter the add image interface (i.e., the image selection interface) of the target application. At this time, select the search button in the add image interface to display the image classification provided by the smart classification application. The image classification displays the above-mentioned specific image category, simulate clicking on the specific image category, and then simulate clicking / swiping to select all images in the category. When selecting images, you can click in sequence or click on all images at the same time, or simulate sliding actions to select all images. The content generation application simulates clicking on the text input box of the target application, and copy and paste the generated text into the text input box. When the above steps are completed, the content generation application displays the publishing interface of the target application on the main screen. In this way, the user can see the publishing interface of the target application, and the user can publish to the social platform with one click by clicking on the publish, such as Fig.18 shown.

[0106] Fig.19 is a schematic diagram illustrating publishing of multiple pictures and related contents according to an exemplary embodiment of the present disclosure.

[0107] Reference Fig.19, the user selects the pictures to be sent to the target application in the photo album or my files application, clicks share, and multiple application icons are displayed. The user selects the content generation application icon and enters the content generation function interface, where multiple pictures can be displayed or partially displayed. In the content generation function interface, select the style of text you want to generate (i.e., target generation) and the target application, click the generate button, and after the text is generated, click send to send the pictures and text directly to the editing interface of the target application. Click publish in the editing interface to complete the publishing of pictures and text. Users can operate simply and intuitively, saving energy and time.

[0108] Fig. 20 is a block diagram illustrating a content generating apparatus according to an exemplary embodiment of the present disclosure.

[0109] Reference Fig. 20 , the content generating device 200 may include an input module 201, a content generating module 202, and a sending module 203. The names and numbers of the above modules are only exemplary, and the present disclosure is not limited thereto.

[0110] The input module 201 can display or partially display multiple pictures input by the user in a preset manner on the content generation function interface.

[0111] The input module 201 may determine a target application and a target generation style.

[0112] The content generation module 202 may generate associated content of the plurality of images having the target generation style in response to the content generation instruction.

[0113] The sending module 203 may display the multiple pictures and the associated content on a designated interface of the target application in response to the sending instruction.

[0114] According to an embodiment, the input module 201 may input the multiple pictures selected by the user into the content generation function interface by sharing; or input the multiple pictures selected by the user into the content generation function interface by dragging; or input the multiple pictures selected by the user into the content generation function interface by voice command. For example, the multiple pictures may be input into the content generation function interface by the user selecting the multiple pictures in a specific application and activating the sharing function; or the multiple pictures may be input into the content generation function interface by the user selecting the multiple pictures in a specific application and dragging. For example, Figures 2 to 5 Shows how to input multiple pictures.

[0115] According to an embodiment, the input module 201 may determine the target generation style based on user operation, or determine the target generation style based on analysis of the plurality of images. For example, the user may select the target generation style from a plurality of predetermined generation styles, or the user may input the target generation style in text form or voice form.

[0116] According to an embodiment, the content generation module 202 may pre-process the multiple images in response to the content generation instruction to obtain an image with a preset size; and generate associated content with the target generation style for the multiple images based on the image with the preset size and the target generation style. Fig.12 Generate related content in a certain way.

[0117] According to an embodiment, the content generation module 202 may stitch the multiple pictures together, adjust the stitched pictures to the preset size, and obtain the pictures of the preset size; or adjust the size of each of the multiple pictures based on the preset size, and stitch each of the adjusted pictures together, and obtain the pictures of the preset size. Figures 6 to 11 The method shown is used to pre-process multiple pictures.

[0118] According to an embodiment, the content generation module 202 may adjust the compression rate of the image of the preset size to a preset compression rate; and generate the associated content based on the image of the preset size with the preset compression rate. Figures 6 to 11 The method shown is used to pre-process multiple pictures.

[0119] According to an embodiment, the sending module 203 may send the identification information corresponding to the plurality of pictures and the associated content to the target application at the same time, so that the plurality of pictures and the associated content are displayed on the designated interface of the target application. Fig.13 and Fig.14 The corresponding sending operation is performed in the manner shown.

[0120] According to an embodiment, the sending module 203 may send identification information corresponding to the plurality of pictures to the target application, so that the plurality of pictures are displayed on the designated interface of the target application; trigger the content editing function on the designated interface, and copy the associated content to the designated interface, so that the plurality of pictures and the associated content are displayed on the designated target interface of the target application. Fig.15 and Fig.16 The corresponding sending operation is performed in the manner shown.

[0121] According to an embodiment, the sending module 203 may create a specific image classification for the multiple images; create a virtual screen and start the target application in the virtual screen; trigger the image selection function of the target application in the virtual screen, and based on the specific image classification, display the multiple images in the designated interface of the target application in the virtual screen; trigger the content editing function on the designated interface in the virtual screen, and copy the associated content to the designated interface, so that the multiple images and the associated content are displayed in the designated interface of the target application in the virtual screen; switch the display content of the virtual screen to the main screen display, so that the multiple images and the associated content are displayed in the designated interface of the target application on the main screen.

[0122] According to an embodiment, the sending module 203 may enter the picture selection interface of the target application in response to triggering the picture selection function of the target application in the virtual screen; select the multiple pictures from the picture selection interface based on the specific picture classification; trigger entry into the specified interface of the target application, and display the multiple pictures in the specified interface. Fig.17 and Fig.18 The corresponding sending operation is performed in the manner shown.

[0123] Hereinabove, the content generating method according to the exemplary embodiment of the present disclosure has been described.

[0124] Next, an electronic device according to an embodiment of the present disclosure is briefly described. Fig.21 is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Fig.21 , the electronic device 1100 may include a memory 1101 and a processor 1102, wherein the processor 1102 is coupled to the memory 1101 and is configured to execute any of the methods described above.

[0125] Fig. 22 FIG. 2 shows a schematic diagram of the structure of an electronic device to which the embodiments of the present disclosure are applicable. Fig. 22 As shown, Fig. 22The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, each of the processor 4001, the memory 4003 and the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure. Optionally, the electronic device may be a first network node, a second network node or a third network node.

[0126] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0127] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 22 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0128] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.

[0129] The memory 4003 is used to store computer programs or executable instructions for executing the embodiments of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the above method embodiments.

[0130] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program or instructions stored thereon. When the computer program or instructions are executed by at least one processor, the steps and corresponding contents of the aforementioned method embodiment can be executed or implemented.

[0131] The embodiments of the present disclosure also provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.

[0132] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described in the text.

[0133] It should be understood that, although the flowchart of the embodiment of the present disclosure indicates each operation step by arrows, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios with different execution times, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present disclosure does not limit this.

[0134] The above text and drawings are provided only as examples to help readers understand the present disclosure. They are not intended and should not be interpreted as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, based on the contents disclosed herein, it is obvious to those skilled in the art that the embodiments and examples shown can be changed without departing from the scope of the present disclosure, and other similar implementation means based on the technical ideas of the present disclosure are adopted, which also fall within the protection scope of the embodiments of the present disclosure.

Claims

1. A content generation method, comprising: Displaying or partially displaying multiple pictures input by the user in a preset manner on the content generation function interface; Determine the target application and target generation style; In response to the content generation instruction, generating associated content of the plurality of images having the target generation style; and In response to the sending instruction, the plurality of pictures and the associated content are displayed on a designated interface of the target application.

2. The method according to claim 1, wherein: The preset method for the user to input the multiple pictures includes one of the following: Inputting the plurality of pictures selected by the user into the content generation function interface by sharing; or Inputting the multiple pictures selected by the user into the content generation function interface by dragging; or The multiple pictures selected by the user are input into the content generation function interface by means of voice commands.

3. The method according to claim 1, wherein: Determine the target generation style, including: Determining the target generation style based on user operation; or The target generation style is determined based on an analysis of the plurality of images.

4. The method according to claim 1, wherein: In response to the content generation instruction, generating associated content of the plurality of images having the target generation style includes: In response to the content generation instruction, preprocess the multiple pictures to obtain a picture with a preset size; Based on the pictures of the preset size and the target generation style, associated content of the multiple pictures with the target generation style is generated.

5. The method according to claim 4, wherein: Preprocessing the multiple images to obtain an image of a preset size includes: splicing the multiple images, adjusting the spliced ​​images to the preset size, and obtaining an image of the preset size; or The size of each of the multiple pictures is adjusted based on the preset size, and each adjusted picture is spliced ​​to obtain a picture of the preset size.

6. The method according to claim 4 or 5, after obtaining the picture of the preset size, the method further comprises: Adjusting the compression rate of the picture of the preset size to a preset compression rate; The associated content is generated based on the image of the preset size with the preset compression rate.

7. The method according to claim 1, wherein: In response to the sending instruction, displaying the plurality of pictures and the associated content on a designated interface of the target application includes: The identification information respectively corresponding to the plurality of pictures and the associated content are sent to the target application at the same time, so that the plurality of pictures and the associated content are displayed on the designated interface of the target application.

8. The method according to claim 1, wherein: In response to the sending instruction, displaying the plurality of pictures and the associated content on a designated interface of the target application includes: Sending identification information respectively corresponding to the plurality of pictures to the target application, so that the plurality of pictures are displayed on the designated interface of the target application; A content editing function on the designated interface is triggered, and the associated content is copied to the designated interface, so that the multiple pictures and the associated content are displayed on the designated target interface of the target application.

9. The method according to claim 1, wherein: In response to the sending instruction, displaying the plurality of pictures and the associated content on a designated interface of the target application includes: Creating a specific image classification for the plurality of images; Creating a virtual screen and launching the target application in the virtual screen; triggering a picture selection function of the target application in the virtual screen, and displaying the plurality of pictures in the designated interface of the target application in the virtual screen based on the specific picture classification; triggering a content editing function on the designated interface in the virtual screen, and copying the associated content to the designated interface, so that the plurality of images and the associated content are displayed in the designated interface of the target application in the virtual screen; Switch the display content of the virtual screen to the main screen display, so that the multiple pictures and the associated content are displayed in the designated interface of the target application on the main screen.

10. The method according to claim 9, wherein: Triggering a picture selection function of the target application in the virtual screen, and displaying the plurality of pictures in the designated interface of the target application in the virtual screen based on the specific picture classification, comprises: In response to triggering the image selection function of the target application in the virtual screen, entering the image selection interface of the target application; Based on the specific picture classification, selecting the plurality of pictures from the picture selection interface; Triggering entry into the designated interface of the target application, and displaying the plurality of pictures in the designated interface.

11. A content generating device, comprising: Input module: configured to: display or partially display multiple pictures input by the user in a preset manner on the content generation function interface; And determine the target application and target generation style; A content generation module is configured to: generate associated content of the plurality of images having the target generation style in response to a content generation instruction; as well as The sending module is configured to: in response to a sending instruction, display the multiple pictures and the associated content on a designated interface of a target application.

12. An electronic device comprising: at least one processor; as well as at least one memory storing computer executable instructions, When the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium storing instructions, wherein: When the instructions are executed by at least one processor, the at least one processor is prompted to perform the method according to any one of claims 1 to 10.