Information processing method and device, storage medium and program product
By displaying multiple pictures on the first page of the target application and automatically calling the picture back-pull model, the problems of cumbersome production process and inaccurate generation style in the existing technology are solved, and an efficient and automated production process is achieved.
Patent Information
- Application Number
- CN202510253419.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the process of generating pictures with similar styles to the reference picture is cumbersome, and the manual input prompt words are not accurate enough, resulting in a large difference between the styles of the generated pictures and the reference picture.
By displaying multiple pictures on the first page of the target application, each picture has a unique style. The user can initiate a picture generation instruction, automatically call the picture backward push model backward to obtain the picture prompt words and parameter information, and automatically jump to the second page to fill in the picture information items to realize an automated picture process.
The process of raw pictures is simplified, the efficiency and quality of raw pictures are improved, and the style similarity between the generated pictures and the reference pictures is higher.
Smart Images

Figure CN120198523A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer processing technologies, and in particular, to an information processing method, device, storage medium, and program product. Background Art
[0002] Currently, some image generation websites can generate an image with a style similar to a reference image based on a reference image. The process of generating an image with a style similar to the reference image on these websites is usually as follows: find a reference image on a website other than the image generation website, download it to the local, or directly select a reference image from the local image library, upload the reference image to the image generation website and manually enter a prompt, click to generate an image, and the image generation website can generate a new image.
[0003] However, the series of steps in the above method are relatively cumbersome, and the manually entered prompt may not be accurately described enough, resulting in a large difference in the style between the newly generated image and the reference image. Summary of the Invention
[0004] Embodiments of this application provide an information processing method, device, storage medium, and program product, which are used to simplify the image generation process, improve the image generation efficiency and quality, and improve the style similarity between the generated image and the reference image.
[0005] Embodiments of this application provide an information processing method, including: in response to an access operation to a target application, display a first page, where the first page includes multiple images, and each image has its own image style; in response to a image generation instruction initiated for a first image among the multiple images, call an image reverse inference model, reverse infer an image generation prompt and image generation parameter information according to the first image, and jump from the first page to a second page; fill the image generation prompt, image generation parameter information, and the first image as the reference image into corresponding multiple image generation information items on the second page; in response to an image generation operation initiated on the second page, call an image generation model to generate at least one second image with a style similar to the first image according to the image generation prompt, image generation parameter information, and the first image in the multiple image generation information items; display at least one second image on the second page.
[0006] Embodiments of this application also provide an electronic device, including: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps in the above methods.
[0007] Embodiments of this application also provide a computer-readable storage medium storing a computer program, which causes the processor to implement the steps in the above methods when the computer program is executed by the processor.
[0008] The embodiments of the present application also provide a computer program product. The computer program product includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the processor can implement the steps in the above-mentioned various methods.
[0009] In the embodiments of the present application, a picture generation instruction can be directly initiated for the first picture displayed on the first page of the target application. Based on this picture generation instruction, the picture generation prompt word and picture generation parameter information of the first picture can be automatically inferred by invoking a picture reverse inference model, and the page can be automatically jumped from the first page to the second page. Moreover, the picture generation prompt word, picture generation parameter information, and the first picture as a reference picture can be automatically filled in multiple corresponding picture generation information items on the second page. The entire process is automatically implemented, simplifying the picture generation process and improving the picture generation efficiency. Further, in response to the picture generation operation initiated on the second page, at least one second picture similar in style to the first picture can be generated by invoking a picture generation model according to the picture generation prompt word, picture generation parameter information, and the first picture in the multiple picture generation information items. Since the picture generation prompt word is inferred by the picture reverse inference model, it is more efficient, comprehensive, and accurate than the prompt word filled in manually, thereby further improving the picture generation efficiency and quality and increasing the style similarity between the second picture and the first picture. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0011] Figure 1 is a schematic flowchart of an information processing method provided by an exemplary embodiment of the present application;
[0012] Figure 2a is a schematic diagram of a page provided by an exemplary embodiment of the present application;
[0013] Figure 2b is a schematic diagram of another page provided by another exemplary embodiment of the present application;
[0014] Figure 2c is a schematic diagram of yet another page provided by yet another exemplary embodiment of the present application;
[0015] Figure 2d is a schematic diagram of yet another page provided by yet another exemplary embodiment of the present application;
[0016] Figure 2e is a schematic diagram of yet another page provided by yet another exemplary embodiment of the present application;
[0017] Figure 2f is a schematic diagram of yet another page provided by yet another exemplary embodiment of the present application;
[0018] Figure 2g A schematic diagram of another page provided for another exemplary embodiment of the present application;
[0019] Figure 2h A schematic diagram of another page provided for another exemplary embodiment of the present application;
[0020] Figure 2i A schematic diagram of another page provided for another exemplary embodiment of the present application;
[0021] Figure 3 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of the present application. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0023] It should be noted that in the case where the embodiments of the present application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection. In addition, various models involved in the present application (including but not limited to language models or large models) comply with relevant laws and standards.
[0024] In addition, it should be noted that in the case where the embodiments of the present application involve user access operations, interaction operations, or trigger operations, the user access operations, interaction operations, or trigger operations involved in the embodiments of the present application include but are not limited to: interaction operations in various ways such as touch operations, gesture operations, voice operations, head movement operations, eye movement operations, etc.; among them, touch operations include but are not limited to: click operations, double-click operations, long-press operations, swipe operations, pinch operations, mouse hover operations, or check operations, etc. Swipe operations include but are not limited to: linear swipes, curved swipes, etc.
[0025] Furthermore, it should be noted that in the case where the embodiments of the present application involve the jump between the first page and the second page, the jump methods involved in the embodiments of the present application include, but are not limited to: directly jumping from the first page to the second page, first jumping from the first page to the task page and then jumping to the second page when corresponding task operations are completed on the task page; the completion of corresponding task operations on the task page includes, but is not limited to: when the task page is implemented as a game page, completing game operations on the game page; when the task page is implemented as an identity authentication page, completing identity authentication on the identity authentication page; when the task page is implemented as a recharge page, completing recharge operations on the recharge page; and so on.
[0026] Regarding the technical problems of the cumbersome steps of generating pictures from pictures and the large style difference between the generated pictures and the reference pictures, in the embodiments of the present application, a picture generation instruction can be directly initiated for the first picture displayed on the first page of the target application. Based on this picture generation instruction, the picture reverse inference model can be automatically called to reverse infer the picture generation prompt words and picture generation parameter information of the first picture, and automatically jump from the first page to the second page. Moreover, the picture generation prompt words, picture generation parameter information, and the first picture as the reference picture can be automatically filled in multiple corresponding picture generation information items on the second page. The whole process is automatically realized, simplifying the picture generation process and improving the picture generation efficiency. Further, in response to the picture generation operation initiated on the second page, at least one second picture similar in style to the first picture can be generated by calling the picture generation model according to the picture generation prompt words, picture generation parameter information, and the first picture in the multiple picture generation information items. Since the picture generation prompt words are reverse inferred by the picture reverse inference model, they are more efficient, comprehensive, and accurate than the prompt words filled in manually, thereby further improving the picture generation efficiency and quality and increasing the style similarity between the second picture and the first picture.
[0027] The following will combine the accompanying drawings to provide a detailed description of a solution provided by the embodiments of the present application.
[0028] Figure 1 It is a schematic flowchart of the information processing method provided by the exemplary embodiment of the present application. As Figure 1 shown, the method includes:
[0029] 101. In response to an access operation to the target application, display a first page, where the first page includes multiple pictures, and each picture has its own picture style;
[0030] 102. In response to a picture generation instruction initiated for the first picture among the multiple pictures, call the picture reverse inference model, reverse infer the picture generation prompt words and picture generation parameter information according to the first picture, and jump from the first page to the second page;
[0031] 103. Fill the raw image prompt, raw image parameter information, and the first image serving as a reference image into multiple corresponding raw image information items on the second page;
[0032] 104. Respond to the raw image operation initiated on the second page, and call the raw image model to generate at least one second image with a style similar to that of the first image according to the raw image prompt, raw image parameter information, and the first image in the multiple raw image information items;
[0033] 105. Display at least one second image on the second page.
[0034] The embodiments of the present application do not limit the implementation form of the target application. For example, the target application can be implemented as an independently running software program, or can be implemented as a small program that depends on the above software program, and so on. The embodiments of the present application also do not limit the type of the target application. For example, the target application can be a browser application (abbreviated as browser), a multimedia application, and so on. It should be noted that the browser application is the focus in the following embodiments of the present application.
[0035] The target application can be installed and run on a terminal device. The embodiments of the present application do not limit the implementation form of the terminal device. For example, the terminal device can be a smart handheld device, such as a smart phone, a tablet computer, a notebook computer, or a desktop computer, and so on; for another example, the terminal device can also be a smart wearable device, such as a smart watch, a smart bracelet, and so on; for still another example, the terminal device can also be various smart home appliances with a display screen, such as a smart TV, a smart large screen, or a smart robot, and so on.
[0036] In the embodiments of the present application, regardless of the type of the target application, in response to an access operation on the target application, a page containing multiple pictures can be displayed. The multiple pictures have their respective picture styles, and the picture styles of different pictures among the multiple pictures can be the same or different. For the convenience of description and distinction, this page can be referred to as the first page. Among them, the access operation can be an operation directly initiated by the user based on the target application, or an operation indirectly initiated by the user on the target application based on other applications. This embodiment does not limit this. In addition, the picture style can be a realistic style, an abstract style, a cartoon style, an illustration style, a retro style, a minimalist style, a surreal style, a pop art style, a cyberpunk style, etc. This embodiment also does not limit this. Among them, the characteristics of the realistic style are high-fidelity to reality, rich details, true colors, and realistic lighting effects; the abstract style is not based on specific images, and expresses emotions or concepts through elements such as shapes, colors, and lines; the cartoon style is exaggerated, humorous, with simple images and bright colors, and is usually used in entertainment and children's content; the illustration style is created through painting or digital tools and is used to explain, decorate, or convey information; the retro style imitates past styles, has a darker tone, has a retro filter effect, and often has a sense of nostalgia; the minimalist style emphasizes white space and simple lines; the surreal style transcends reality, integrates dreamy and fantasy elements, and breaks conventional logic; the pop art style features bright colors, bold patterns, and exaggerated expression techniques, and often draws on elements of popular culture; the cyberpunk style combines high-tech and low-quality life scenarios, usually with neon lights, futuristic buildings, and a dystopian atmosphere. Pictures of different styles may vary in terms of picture colors, tones, lines of picture elements, rendering atmosphere, etc.
[0037] Among them, when the access operation is directly initiated by the user based on the target application, if the target application is a browser, the access operation may be a series of continuous interaction actions. At this time, the first page is a web page displayed through the browser. For example, in response to the user's trigger operation on the browser application, the browser home page is displayed; further, in response to the search operation for the input target URL on the browser home page, the page corresponding to the target URL (or called the network image page) is displayed, and the page includes multiple pictures. The page corresponding to the target URL can be used as the first page, and the target URL can be regarded as the URL of the network image page. Or, there are multiple web page entrances provided on the browser home page. In response to the access operation on any web page entrance, the corresponding page is displayed, and this page can be used as the first page, then this web page entrance is the entrance to the picture web page. Or, there are multiple web page entrances provided on the browser home page. In response to the access operation on any web page entrance, the corresponding page is displayed, and this page includes the page entrance to the network image page; in response to the trigger operation on this page entrance, the corresponding network image page is displayed, and this network image page can be regarded as the first page. If the target application is not a browser application, the access operation can be the user's direct trigger operation on the target application. For example, in response to the user's access operation on the target application, the home page of the target application is displayed, and the home page can be used as the first page; or, in response to the user's access operation on other pages in the target application, other pages are displayed, and other pages can be used as the first page.
[0038] Among them, when the access operation is indirectly initiated by the user to the target application based on other applications, the access operation can be the access operation triggered by the user on the access link displayed in other applications. In response to this access operation, the page address corresponding to the access link is obtained, and based on this page address, the corresponding page is jumped to. The corresponding page includes multiple pictures, then this page can be used as the first page. Among them, this first page can be a web page displayed in the browser, and the page address corresponding to the access link obtained in response to the trigger operation is the web page address. Or, this first page can also be a page of other applications other than the browser, and this page includes multiple pictures.
[0039] It should be noted that regardless of how the first page is obtained through the access operation method, the first page needs to contain multiple pictures and the technical solutions described in the following embodiments can be implemented based on these pictures. Moreover, when the first page is a web page opened by the browser, the first page can be any web page opened by the browser, that is, the first page is not a fixed page, as long as the first page contains multiple pictures and the technical solutions in the following embodiments can be implemented based on these pictures. In addition, the embodiments of the present application do not limit the website to which any web page belongs.
[0040] Further, after the first page is displayed, the user can view multiple pictures on the first page. If the multiple pictures support interactive functions, during the user's browsing process, the user can initiate an image generation instruction for any one of the multiple pictures, and any one of the pictures can be referred to as the first picture. Correspondingly, in response to the image generation instruction initiated for the first picture, the image reverse inference model can be called to reverse infer the image generation prompt and image generation parameter information based on the first picture, and then jump from the first page to the second page. Thus, based on the image generation instruction initiated for the first picture on the first page, a series of steps of calling the image reverse inference model, reverse inferring the image generation prompt and image generation parameter information based on the first picture, and jumping from the first page to the second page can be automatically executed, thereby simplifying the image generation process and improving the image generation efficiency. In addition, the picture described by the image generation prompt reverse inferred from the first picture by calling the image reverse inference model has a higher richness, a more abundant vocabulary, a more comprehensive and accurate description, and more accurate image generation parameter information, which can avoid the problem that the user's own editing of the image generation prompt and image generation parameters is inaccurate and affects the image generation quality. In addition, the process of calling the image reverse inference model to reverse infer the image generation prompt, image generation parameter information, and jumping from the first page to the second page based on the first picture can be regarded as being executed synchronously, with almost no time difference, saving time for the image generation process and further improving the image generation efficiency.
[0041] Among them, the image reverse inference model can be a text-to-image model based on AI. The type of the text-to-image model is not limited in this embodiment. For example, it can be an SD (Stable Diffusion) model. In addition, when selecting the SD model, a model with generalization ability after pre-training and high image clarity can be selected, such as the Stable Diffusion XL (abbreviated as SDXL) model. The SDXL model is a new version of the SD model series, and it has made significant improvements and optimizations in image generation quality. In addition, the image reverse inference model can be deployed on the server side or on the terminal device, and this embodiment does not limit this. When the image reverse inference model is deployed on the server side, in response to the image generation instruction, the image reverse inference model is called from the server side so that it can complete the reverse inference process of the image generation prompt and image generation parameter information of the first picture on the terminal device; or, in response to the image generation instruction, the first picture is sent to the server side so that the server side can call the image reverse inference model to reverse infer the image generation prompt and image generation parameter information of the first picture, and return the image generation prompt and image generation parameter information to the terminal device. When the image reverse inference model is deployed on the server side, in response to the image generation instruction, the reverse inference model is called locally from the terminal device so that it directly completes the reverse inference process of the image generation prompt and image generation parameter information of the first picture on the terminal device. Thus, the process of calling the image reverse inference model from the server side can be reduced, the process can be simplified, and the overall reverse inference efficiency of the image reverse inference model can be improved.
[0042] Further optionally, during the user's browsing process, the user can initiate an image generation instruction for multiple images among multiple images simultaneously. The multiple images can be referred to as the first images. Correspondingly, in response to the image generation instruction initiated for the first images, an image reverse inference model can be called to reverse infer the image generation prompt words and image generation parameter information for each of the first images respectively, and jump from the first page to the second page.
[0043] In some alternative embodiments, in response to the image generation instruction initiated for the first image among multiple images, calling the image reverse inference model, reverse inferring the image generation prompt words and image generation parameter information based on the first image, and jumping from the first page to the second page includes: in response to the selection operation for the first image, displaying a first image generation control, where the first image generation control is associated with the second page; in response to the triggering operation for the first image generation control, calling the image reverse inference model, reverse inferring the image generation prompt words and image generation parameter information based on the first image, and jumping from the first page to the second page.
[0044] In the above embodiments, regarding the content of displaying the first image generation control in response to the selection operation for the first image, the specific implementation manner of displaying the first image generation control in response to the selection operation for the first image is not specifically limited. The following is an example.
[0045] In an alternative embodiment, in response to the selection operation for the first image, displaying the first image generation control on the first image includes: in response to the hovering operation for the first image, displaying a first floating layer on the first image, and the first floating layer includes an image generation control, which can be referred to as the first image generation control. Among them, the hovering operation can be regarded as the triggering operation by the user for the first image, but the triggering operation method is not limited to the hovering operation. Other triggering operation methods can refer to the relevant descriptions in the above embodiments and will not be elaborated here. It should be noted that the first control can be the one that generates and displays the first image generation control associated with the first image in real time in response to the selection operation for the first image, or can pre-generate the first image generation control associated with each image. This embodiment does not make a limitation on this. Similarly, other controls in the following embodiments can also be generated in real time in response to the corresponding operations or can be pre-generated.
[0046] More specifically, in response to the selection operation for the first image, displaying the first image generation control on the first image includes: in response to the hovering operation for the first image, displaying a second image generation control on the first image; in response to the hovering operation for the second image generation control, displaying a first floating layer on the first image, and the first floating layer includes the first image generation control.
[0047] For example, after a user browses to a first picture on a first page displayed by a browser, if the user wants to generate a picture with a similar style to the first picture, the user can perform a hover operation on the first picture, such as controlling the mouse icon to hover over the first picture, and in response to the hover operation, displaying a raw image icon on the first picture and displaying a first floating layer in an area associated with the raw image icon, wherein the first floating layer includes a first raw image control. Figure 2a In the first picture shown, the "π" logo displayed on the first picture is a raw picture logo, the first floating layer can be displayed in a local area of the first picture, and the "one-click same style" control in the first floating layer is a first raw picture control. Displaying the raw picture logo can be used to indicate that a series of raw picture steps can be implemented based on the current picture.
[0048] For another example, after the user browses to the first picture in the first page displayed by the browser, if the user wants to generate a picture with a similar style to the first picture, the user can perform a hover operation on the first picture, such as the user can control the mouse marker to hover over the first picture, and in response to the hover operation, a raw image marker control is displayed at any position of the first picture, and any position can be the upper left corner, upper right corner, lower left corner, or lower right corner of the first picture, but is not limited to this. Further, the user can perform a hover operation on the raw image marker control, such as the user can control the mouse marker to hover over the raw image marker, and in response to the hover operation, a first floating layer is displayed in the area associated with the raw image marker in the first picture, and a menu list is displayed in the first floating layer, and the menu list includes the first raw image control. Figure 2a In the first picture shown, the "π" logo displayed on the first picture is a raw picture logo, the first floating layer can be displayed in a local area of the first picture, and the "one-click same style" control in the first floating layer is the first raw picture control. It should be noted that the above examples are only exemplary and do not constitute a limitation on the technical solution of this application.
[0049] In another optional embodiment, in response to a selection operation on the first image, a first raw image control is displayed on the first image, including: the first page includes a second raw image control or a raw image menu control, and in response to a triggering operation on the second raw image control or the raw image menu control, a picture upload page is displayed, wherein the picture upload page can be displayed in a local area of the first page, and the picture upload page can also be an independent page; in response to an operation of dragging the first image to the picture upload page, the first image is displayed on the picture upload page; wherein a first raw image control is displayed on the picture upload page, and the first raw image control can be provided by the picture upload page, or can be displayed in response to a dragging operation on the first image.
[0050] Let’s take an example. Figure 2bAs shown, an identification control with the word "π" is displayed in the lower right corner of the first page. This raw image identification control can be regarded as the second raw image control. A raw image menu control is displayed in the upper right corner of the first page. In response to a trigger operation on the second raw image control or the raw image menu control, a sidebar is displayed in the right area of the first page. This sidebar can be used as an image upload page, and the image upload page displays an image display area and a first raw image control. Further, if the user wants to generate an image similar in style to the first image, the user can, in response to the operation of dragging the first image to the image upload page, display the first image in the image display area of the image upload page. Additionally, as Figure 2c shown, in response to the drag operation on the first image, the image display area will be enlarged and displayed above the image upload page in the form of a floating layer; alternatively, the image upload page also includes an image display area magnification control. In response to the trigger operation on the image display area magnification control, the image display area is magnified and displayed above the first page in the form of a floating layer, so that the image can be accurately dragged into the image display area.
[0051] Additionally, the image upload page also includes: an image set control. In response to the trigger operation on the image set control, an image set page is displayed. All the images in the first page can be loaded and displayed in the image set page. These images can be displayed in the image set page in the page display order of the first page, or the display order of each image in the image set page can be adjusted according to the user's needs; in response to the selection operation on at least one image displayed in the image set page, the first raw image control is displayed. Among them, when the selected image is one, the first image is one, then one first image is displayed in the first image information item of the second page, so as to generate a second image similar in style to this image based on this first image. When the selected images are multiple, the first images are multiple, then multiple first images are displayed in the first image information item of the second page, so as to generate a second image similar in style to each first image respectively based on each first image, that is, multiple second images are obtained.
[0052] For example. As Figure 2d shown, an image set control is displayed in the sidebar of the first page. In response to the trigger operation on this image set control, all the images in the first page are loaded and displayed. Further, in response to the selection operation on any one image, a "One - key same style" control is displayed. This control can be used as the first raw image control; further, in response to the trigger operation on this control, the second page is displayed, and the above - mentioned one image is displayed in the second page. Or, as Figure 2e, in response to a selection operation on any N pictures, a "Batch Same Style" control will be displayed. This control can be used as the first image generation control, where N is a positive integer greater than or equal to 2. The "Batch Same Style" control can be used to batch generate multiple second pictures with a style similar to the first picture. For example, in response to a selection operation on any 2 pictures, the "Batch Same Style" control is displayed; further in response to a trigger operation on the "Batch Control", a second page is displayed, and the above two pictures are shown on the second page.
[0053] In another alternative embodiment, in response to a selection operation on the first picture, a first image generation control is displayed on the first picture, including: in response to a trigger operation on the first picture, a second floating layer is displayed on the first picture, and the second floating layer includes a third image generation control; in response to a trigger operation on the third image generation control, a third floating layer is displayed, and the third floating layer includes the first image generation control.
[0054] Illustrate with an example. After the user browses the first picture on the first page displayed by the browser, if the user wants to generate a picture with a style similar to the first picture, the user can perform a trigger operation on the first picture. For example, after the user hovers the mouse over the first picture and then right-clicks the mouse, a second floating layer is displayed on the first picture, and the second floating layer includes a third image generation control. In response to a trigger operation on the third image generation control, a third floating layer is displayed, and the third floating layer includes the first image generation control. As Figure 2f shown, in response to the user's right-click operation on the first picture with the mouse, a floating layer is displayed. This floating layer is the second floating layer, and the second floating layer includes an identification control with the word "π", and this identification control is the third image generation control. Further, in response to a trigger operation on the "π" identification control, another floating layer is displayed. This floating layer is the third floating layer, and the third floating layer includes a "One-key Same Style" control, and this control can be used as the first image generation control.
[0055] It should be noted that the ways of presenting the first image generation control in Embodiment 1, Embodiment 2, and Embodiment 3 can exist independently, or a combination of two ways can be presented, or all three ways can coexist. When two ways coexist or all three ways coexist, any one of the ways can be selected to display the first image generation control.
[0056] Further, after the first image generation control is displayed, a trigger operation on the first image generation control can be responded to, the image reverse inference model can be called, the image generation prompt word and the image generation parameter information can be obtained by reverse inference from the first image, and the first page can be jumped to the second page. More specifically, in response to the trigger operation on the first image generation control, an image reverse inference model call instruction and a page jump instruction are generated, so as to call the image reverse inference model based on the model call instruction, obtain the image generation prompt word and the image generation parameter information by reverse inference from the first image, and jump from the first page to the second page based on the page jump instruction. It can be understood that the image generation instruction mentioned in the embodiment of the present application can be regarded as an instruction generated in response to the trigger operation on the first image generation control. Or, a series of instructions including the first image generation control display instruction generated in response to the selection operation on the first image, and the model call instruction and the page jump instruction generated in response to the trigger operation on the first image generation control are regarded as the image generation instruction.
[0057] In some alternative embodiments, the image reverse inference model includes: an image element recognition network layer, a feature extraction network layer, and a prompt word generation network layer; then, in response to the trigger operation on the first image generation control, calling the image reverse inference model, and obtaining the image generation prompt word and the image generation parameter information by reverse inference from the first image, including: inputting the first image into the image element recognition network layer, segmenting each element included in the first image, and identifying the main elements included in the first image; inputting the main elements into the feature extraction network layer, extracting the visual features and the image generation parameter features corresponding to each main element, and taking the image generation parameter features as the image generation parameter information; inputting the first image, the visual features corresponding to each main element, the image generation parameter features, and the preset image generation prompt word structure into the prompt word generation network layer to generate the image generation prompt word corresponding to the first image, and outputting the image generation prompt word and the image generation parameter information. Among them, the main element refers to the main body of the picture that can affect the picture style. Taking the example that a person wearing ethnic clothing and a potted plant are shown in the picture, the person wearing ethnic clothing and the plant can be used as the main elements. The visual feature refers to the information such as the pattern, color, and line thickness included in the main element.
[0058] Among them, the image generation prompt word includes a positive prompt word, and the positive prompt word is used to guide the model to generate a prompt word that meets the expectation. The positive prompt word has a prompt word structure, and the prompt word structure of the positive prompt word can be, for example, "picture main body + picture details + picture style + picture parameters", but is not limited thereto. For example, Figure 2gAs shown, the exemplary prompt words are as follows: Please use the following description content as reference content to generate a new picture that conforms to its description style: "A young Asian woman in traditional Hanfu, with complex patterns and rich blue-green tones on the clothes. She has fair skin, black and shiny hair tied into a bun, decorated with an emerald green jade hairpin. She has a serene expression, closed eyes, with a faint smile, as if she is in deep thought or meditation. This is a digital CGI artwork with high-resolution HDR effect, and the background is blurred to highlight the subject." Among them, "A young Asian woman in traditional Hanfu" is the main body of the picture, "the clothes have complex patterns and rich blue-green tones, her skin is fair, black and shiny hair tied into a bun, decorated with an emerald green jade hairpin. She has a serene expression, closed eyes, with a faint smile, as if she is in deep thought or meditation" is the picture detail, "this is a digital CGI artwork" is the picture style, and "high-resolution HDR effect, the background is blurred to highlight the subject" is the picture parameter. In addition, the raw picture prompt words also include reverse prompt words, which are used to guide the model not to generate prompt words that do not meet expectations, and the reverse prompt words can also have a prompt word structure. An exemplary prompt is as follows: Do not include background plants when generating new images.
[0059] Further optionally, the picture upload page also includes: at least one picture editing control, which may be, for example, a screenshot control, a cutout control, and a picture beautification control, but is not limited thereto; after the first picture is displayed on the picture upload page, the first picture may be edited in response to a trigger operation on any of the picture editing controls to obtain a first picture after editing. The picture editing process includes, but is not limited to, screenshot processing, cutout processing, and picture beautification processing.
[0060] Correspondingly, in response to the triggering operation of the first raw image control, the page jumps to the second page, and the image reference information is filled with the first image after the photo is edited. Figure 2b As shown, taking the photo editing control as a screenshot control as an example, in response to a trigger operation on the screenshot space, a screenshot is taken of the first picture, and the area selected in the frame in the first picture is captured to obtain the first picture after the screenshot; in response to the trigger operation on the first raw picture control, the page jumps to the second page, and the first picture after the screenshot is filled in the picture reference information item.
[0061] Further optionally, in response to a selection operation on the first picture, displaying the first image generation control further includes: in response to the selection operation on the first picture, simultaneously displaying the first image generation control and the magnification control. The magnification control can be a high-definition magnification control, which can not only magnify the picture but also improve the resolution of the picture. Correspondingly, in response to a trigger operation on the magnification control, jump from the first page to the second page, and fill the first picture onto the second page; in response to a configuration operation of configuring magnification parameters on the second page, perform a magnification operation on the first picture according to the configured magnification parameters, and display the magnified first picture on the second page, as Figure 2h is a schematic diagram of the magnified page. Or, in response to a trigger operation on the magnification control, jump from the first page to the third page, and fill the first picture onto the third page; in response to a configuration operation of configuring magnification parameters on the third page, perform a magnification operation on the first picture according to the configured magnification parameters, and display the magnified first picture on the third page. The third page is different from the second page, and the second page and the third page can be web pages of the same website or different websites. In response to a trigger operation on the high-definition magnification control, jump from the first page to the second or third page, and fill the first picture onto the second or third page; in response to a configuration operation of configuring high-definition magnification parameters on the second or third page, perform a high-definition magnification operation on the first picture according to the configured high-definition magnification parameters, and display the magnified first picture on the second or third page. Additionally, on the second or third page, there is also a historical magnified picture display area, where historical high-definition magnified pictures can be displayed. Thus, the user can view a clearer first picture, improving the visual effect. Among them, when performing high-definition magnification on the picture, a high-definition magnification model can be called to process the picture. The high-definition magnification model can be, for example, a high-definition magnification model based on SD.
[0062] Further, in the case where the page displaying the magnified first picture is the third page different from the second page, in response to a hovering or trigger operation on any historical magnified picture, display the first image generation control; in response to a trigger operation on the first image generation control, a picture reverse-pushing model can be called to reverse-push the image generation prompt words and image generation parameter information based on the magnified first picture, and jump from the third page to the second page; fill the image generation prompt words, image generation parameter information, and the magnified first picture into the corresponding multiple image generation information items on the second page; in response to an image generation operation initiated on the second page, call an image generation model to generate at least one second picture with a style similar to the magnified first picture according to the image generation prompt words, image generation parameter information, and the magnified first picture in the multiple image generation information items; display at least one second picture on the second page. Since the picture of the first picture after high-definition magnification is clearer, the image generation prompt words and image generation parameter information reversed by the picture reverse-pushing model are more accurate, which is beneficial to improving the style similarity between the second picture and the first picture.
[0063] As can be seen from some of the above embodiments, in response to a trigger operation on the first image generation control, the image reverse inference model is called. While generating the image generation prompt and image generation parameter information by reverse inference from the first image, the first page jumps to the second page. Among them, the second page is a fixed image generation page. The process of calling the image reverse inference model to generate the image generation prompt, image generation parameter information by reverse inference from the first image and jumping from the first page to the second page can be regarded as being executed synchronously, with almost no time difference, saving time for the image generation process and further improving the image generation efficiency. It should be noted that when the target application is a browser, the second page is a web page, and the website to which the web page belongs is a fixed AI image generation website. When the target application is not a browser, the second page is the page of the target application itself, that is, the application page.
[0064] It should be noted that during the process of the first page jumping to the second page, the application types of the two applications before and after the jump are not restricted. The application types include browsers and other applications other than browsers. In other words, during the process of the first page jumping to the second page, the page types of the first page and the second page before and after the jump are not restricted. The page types include: web pages displayed by browsers and application pages of other applications other than browsers. In the same page jump process, the first page before the jump and the second page after the jump can be pages of the same type or different types. For example, if the first page is a web page and the second page is an application page, the corresponding page jump process is from the web page to the application page. Another example is that if the first page is an application page and the second page is a web page, the corresponding page jump process is from the application page to the web page. Another example is that if the first page is an application page and the second page is also an application page, the corresponding page jump process is from the application page to the application page. At this time, the two pages are usually application pages of different applications. Another example is that if the first page is a web page and the second page is also a web page, the corresponding page jump process is from the web page to the web page. At this time, the two web pages belong to web pages of the same website or web pages of different websites. The two web pages can be web pages displayed in different browsers or web pages displayed in the same browser. It should be noted that in the embodiments of the present application, the focus is on two web pages that are jumped and displayed in the same browser.
[0065] Furthermore, the second page contains multiple image generation information items, which include but are not limited to: reference picture information items, image generation prompt information items, and image generation parameter information items. Among them, the image generation parameter information items include but are not limited to: picture size information items, style degree information items, chaos degree information items, and style reference information items. Correspondingly, the image generation parameter information includes: size information, style degree information, chaos degree information, and style reference information. After the image generation prompt, size information, style degree information, chaos degree information, and style reference information corresponding to the first picture are deduced by the image reverse deduction model, the image generation prompt, size information, style degree information, chaos degree information, style reference information, and the first picture used as a reference picture are filled in the corresponding information items on the second page. Figure 2i It is an exemplary schematic diagram of the second page. Among them, the size information is the size information of the target picture to be generated. The style degree information is used to constrain the prompt. The larger the value of the style degree information, the more divergent the model will be when using the image generation model to generate pictures later, that is, the stronger the artistic degree of the generated picture, and the more deviated from the style picture described by the positive prompt. The smaller the value of the style degree information, the more divergent the model will be when using the image generation model to generate pictures later, that is, the weaker the artistic degree of the generated picture, and the closer it is to the style picture described by the positive prompt. The style reference information is used to constrain the generated target picture. The larger this value is, the greater the degree of reference to the reference picture when generating the target picture, and the more similar the style of the target picture is to the reference picture. The smaller this value is, the smaller the degree of reference to the reference picture when generating the target picture, and the more deviated the style of the target picture is from the reference picture. The chaos degree information is a constraint on the generation result, specifically referring to the similarity degree of the main elements in the target picture and the main elements in the reference picture. For example, if there are 3 reference pictures, then 3 pictures will be generated, each of which is similar in style to each reference picture, but the actions of the same person in each target picture are different, indicating a higher chaos degree. If the actions of the same person in each target picture are relatively close, it indicates a lower chaos degree.
[0066] Furthermore, the values of each image generation parameter information have default values. When filling in the information items, the default values can be filled. On this basis, each information item also has an adjustment function, and the user can adjust the values of each image generation parameter information according to the actual image generation requirements.
[0067] It should be noted that when there are multiple first pictures, on the second page, each picture has a corresponding image generation prompt information item and image generation parameter information item. Correspondingly, after calling the image reverse deduction model to infer the image generation prompt and image generation parameter information of each picture, the image generation prompt and image generation parameter information of each picture will be filled in the corresponding image generation prompt information item and image generation parameter information item of each picture respectively.
[0068] Further optionally, the generated image prompt obtained through inference can also be optimized. Correspondingly, the second page further includes an AI prompt optimization control. In response to a trigger operation on the AI prompt optimization control, based on the generated image prompt filled in the generated image prompt information item, an AI prompt optimization model is called to optimize the prompt. For example, Figure 2g the three Chinese prompts in [reference] are candidate prompts optimized based on the generated image prompt filled in the generated image prompt information item. Further, in response to a selection operation on any one of the optimized prompts, the prompt originally filled in the generated image information item is replaced with the selected candidate prompt, so as to generate a second image based on the optimized generated image prompt to improve the quality of the second image. Among them, the AI prompt optimization model can be, for example, the LLaMA (Large Language Model Meta AI) model.
[0069] Further optionally, the second page further includes an original image information item, which is used to fill in the original image. The original image is a base image used to generate an image similar in style to the first image.
[0070] Further, after filling the original image into the original image information item on the second page, in response to a generate-image operation initiated on the second page, at least one second image similar in style to the first image is generated by calling a generate-image model according to the generated image prompt, generated image parameter information, and the first image in multiple generated image information items. Among them, the generate-image model can be a hybrid model of an AI-based text-to-image model and an AI-based image-to-text model. The AI-based text-to-image model can be the SD model, and the AI-based image-to-text model can be the Inswapper model. For the relevant description of calling the generate-image model to generate at least one second image similar in style to the first image, refer to the relevant description of the following embodiments.
[0071] In some alternative embodiments, the second page includes a fourth generate-image control. In response to a generate-image operation initiated for the fourth generate-image control on the second page, at least one second image similar in style to the first image is generated by calling a generate-image model according to the generated image prompt, generated image parameter information, and the first image in multiple generated image information items.
[0072] In some alternative embodiments, the image generation model includes: an image element recognition network layer, a feature extraction network layer, a style adaptation network layer, an image generation network layer, and an image adjustment network layer; in response to an image generation operation initiated for a fourth image generation control on a second page, at least one second image similar in style to a first image is generated by calling the image generation model according to the image generation prompt word, image generation parameter information, and the first image among a plurality of image generation information items, including: inputting the original image into the image element recognition network layer to segment each element included in the original image and recognize each element included in the original image; inputting each element included in the original image into the feature extraction network layer to extract the visual features corresponding to each element; inputting the visual features corresponding to each element in the original image, the visual features corresponding to each main element included in each of at least one first image, each first image, and the image generation prompt word into the style adaptation network layer to respectively recognize the visual features corresponding to each element in the original image and the visual features corresponding to each main element included in the first image and / or the style of the first image, and in the case of a style mismatch, under the prompt of the image generation prompt word, respectively adjust the style of the visual features corresponding to each element in the original image according to the visual features corresponding to each main element included in each first image and / or the style of the first image to obtain the target visual feature information corresponding to each element in the original image and each first image respectively; inputting the target visual features corresponding to each first image and the image generation prompt word into the image generation network layer to generate at least one initial image based on the target visual feature information corresponding to each image respectively; inputting at least one initial image, the first image, and the prompt word into the image adjustment network layer, and respectively adaptively adjust at least one initial image based on the style of the first image under the prompt of the image generation prompt word to obtain at least one second image.
[0073] In some alternative embodiments, the target application may be a browser, and the first page and the second page are respectively the first web page and the second web page displayed through the browser. Among them, a target plugin is embedded in the browser. Through the target plugin, a first image generation control can be added to the first page or each picture on the first page, and based on this first image generation control, an image reverse inference model can be called to reverse infer an image generation prompt and image generation parameter information from the first picture, jump from the first page to the second page, and fill the image generation prompt, image generation parameter information, and the first picture as a reference picture into multiple corresponding image generation information items on the second page. Additionally, the target plugin can be developed through the plasmo framework and shadow dom to ensure style isolation. The plasmo framework is an open-source development framework for chrom plugins, through which chrom plugins can be developed more quickly. Shadow dom can isolate the style of the plugin, enabling the plugin to have its own exclusive format to avoid affecting the style of the original web page when the plugin is running. The browser also supports webconfig to configure different website policies to flexibly control the usage scope of the target plugin. For example, each website has a different way of writing pictures, so it is necessary to configure website information to adapt to different website policies and thus control the usage scope of the target plugin. On this basis, an image generation control can also be created for pictures through content script injection and dom listening. Content script is a script on the terminal device side of the target plugin used to handle the front-end page situation. That is, after the web page data browsed in the current browser is rendered by the corresponding website by transmitting data through the DOM backend, content script needs to monitor the change of the data, thereby identifying which picture the user stays on and creating a button for this picture. The side panel api is used to create a browser sidebar to achieve an immersive usage effect. The first picture is passed into the image reverse inference model or the high-definition magnification model by bypassing the anti-leeching link through canvas and base64. Canvas is a graphics library, and the picture is parsed through canvas. Base64 is a picture format that is convenient for the code to better recognize.
[0074] Optionally, call the image reverse inference model, reverse infer the image generation prompt and image generation parameter information from the first image, jump from the first page to the second page, and fill the image generation prompt, image generation parameter information, and the first image as a reference image into multiple corresponding image generation information items on the second page, including: using the target plugin, call the image reverse inference model, and reverse infer the image generation prompt and image generation parameter information from the first image; using the target plugin, jump from the first page to the second page; using the target plugin, fill the image generation prompt, image generation parameter information, and the first image as a reference image into multiple corresponding image generation information items on the second page. Among them, the second page includes an information configuration area and a display area, and the display area is used to display the target image.
[0075] Further optionally, the first web page and the second web page are web pages in the same website; or, the first web page and the second web page are web pages in different websites.
[0076] Further optionally, when the first web page and the second web page are web pages in different websites, run the target plugin to perform each operation in the method of the above embodiment, including: in response to the user's login operation, send an authentication request to the website server where the second web page is located, and the authentication request includes the information of the target plugin for the website server to authenticate. The login operation is a login operation for the target plugin; receive the authentication passed information returned by the website server, and run the target plugin to enable the target plugin to perform each operation in the method. Among them, when the user performs the login operation, the target plugin can use the SSO ticket of the image generation website to which the second page belongs to achieve functional interoperability with the image generation website, that is, the account of the logged-in target plugin is the same as the account of the image generation website, so as to achieve functional interoperability. Based on the image generation instruction, the image reverse inference model can be automatically called to reverse infer the image generation prompt and image generation parameter information of the first image, and automatically jump from the first page to the second page. According to the image generation prompt, image generation parameter information, and the first image in multiple image generation information items, call the image generation model to generate at least one second image with a style similar to the first image.
[0077] More specifically, when a user logs in to the target plugin or the image generation website through an account, the backend of the target plugin or the image generation website calls the Single Sign-On (SSO) service to verify the user's identity. After successful verification, the SSO service generates an SSO Ticket (such as a JWT token) and returns it to the target plugin or the image generation website. When the target plugin jumps to the image generation website, it attaches the SSO Ticket as a parameter to the URL (such as https: / / ******_ticket=token). After receiving the request, the image generation website parses the SSO Ticket in the URL and verifies its validity. If the SSO Ticket passes the verification, the image generation website establishes a user session based on the user information in the SSO Ticket, allowing the user to use the image generation function without logging in again. In this way, the user identity verification and function connection are achieved between the target plugin and the image generation website through a unified SSO Ticket. Through the above technical implementation, the target plugin can not only provide the function of generating pictures of the same style with one click for users when browsing web pages, but also achieve seamless connection with the image generation website through SSO technology, improving the user experience.
[0078] It should be noted that in addition to the target application being a browser application, when the first page jumps to the second page involves a jump between two non-browser applications, or when the first page jumps to the second page involves a jump between a browser application and a non-browser application, or when it is a jump between different browser applications, it can also be achieved by embedding a plugin. The specific implementation method can refer to the relevant description of the above target plugin and will not be elaborated here.
[0079] To facilitate the understanding of the technical solutions provided in the above embodiments, the following takes the target application as a browser, and the first web page and the second web page are web pages in different websites displayed by the browser as an example for detailed description.
[0080] First, a target plugin is embedded in the target browser. Through this target plugin, various controls can be implemented on the first web page displayed by the browser to call the image reverse-pushing model, reverse-push the image generation prompt and image generation parameter information from the first image, jump from the first web page to the second web page, and fill the image generation prompt, image generation parameter information, and the first image as a reference image into multiple corresponding image generation information items on the second web page. Users can log in to the target plugin and the image generation website through an account, so that the target plugin can provide the function of generating pictures of the same style with one click for users when browsing web pages in the browser, and can also achieve seamless connection with the image generation website through SSO technology, improving the user experience.
[0081] Further, in response to an access operation on the first web page in the target browser, the first web page is displayed. The first web page includes multiple pictures, and each picture has its own picture style. Then, in response to a hovering operation on the first picture, a first floating layer is displayed on the first picture, and the first floating layer includes a first image generation control; alternatively, the first web page includes a second image generation control or an image generation menu control. In response to a triggering operation on the second image generation control or the image generation menu control, an image upload page is displayed; in response to an operation of dragging the first picture to the image upload page, the first picture is displayed on the image upload page; wherein, a first image generation control is displayed on the image upload page. After the first picture is displayed, in response to a selection operation on the first picture, the first image generation control is displayed, and the first image generation control is associated with a second web page. In response to a triggering operation on the first image generation control, an image reverse inference model is called, and an image generation prompt word and image generation parameter information are obtained by reverse inference from the first picture, and the page jumps from the first web page to the second web page.
[0082] Further, the image upload page further includes: at least one image retouching control; then, after the first picture is displayed on the image upload page, in response to a triggering operation on any one of the image retouching controls, the first picture can be retouched to obtain a retouched first picture. Further, in response to a selection operation on the first picture, the first image generation control and a magnification control are simultaneously displayed; then, in response to a triggering operation on the magnification control, the page jumps from the first web page to the second web page, and the first picture is filled on the second web page; in response to a configuration operation of configuring magnification parameters on the second web page, the first picture is magnified according to the configured magnification parameters, and the magnified first picture is displayed on the second web page.
[0083] Further, the image generation prompt word, the image generation parameter information, and the first picture as a reference picture are filled in multiple corresponding image generation information items on the second web page. Then, in response to an image generation operation initiated on the second web page, at least one second picture similar to the style of the first picture is generated by calling an image generation model according to the image generation prompt word, the image generation parameter information, and the first picture in the multiple image generation information items, and at least one second picture is displayed on the second web page.
[0084] Based on the target plug-in, the technical solutions provided in the above embodiments of the present application can initiate an image generation instruction directly for the first image displayed on the first page of the target application. Based on this image generation instruction, the image reverse inference model can be automatically called to reverse infer the image generation prompt word and image generation parameter information of the first image, and automatically jump from the first page to the second page. Moreover, the image generation prompt word, image generation parameter information, and the first image as a reference image can be automatically filled in multiple corresponding image generation information items on the second page. The whole process is automatically realized, which simplifies the image generation process and improves the image generation efficiency. Further, in response to the image generation operation initiated on the second page, at least one second image similar in style to the first image can be generated by calling the image generation model according to the image generation prompt word, image generation parameter information, and the first image in the multiple image generation information items. Since the image generation prompt word is reverse inferred by the image reverse inference model, it is more efficient, comprehensive, and accurate than the prompt word filled in manually, thus further improving the image generation efficiency and quality and the style similarity between the second image and the first image.
[0085] Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. As Figure 3 shown, it includes: a memory 30a and a processor 30b; the memory 30a is used to store a computer program; the processor 30b is coupled to the memory 30a and is used to execute the computer program to implement the following steps:
[0086] In response to an access operation to the target application, display a first page, where the first page includes multiple pictures, and each picture has its own picture style; in response to an image generation instruction initiated for the first picture among the multiple pictures, call the image reverse inference model, reverse infer the image generation prompt word and image generation parameter information according to the first picture, and jump from the first page to the second page; fill the image generation prompt word, image generation parameter information, and the first picture as a reference picture in multiple corresponding image generation information items on the second page; in response to the image generation operation initiated on the second page, call the image generation model to generate at least one second picture similar in style to the first picture according to the image generation prompt word, image generation parameter information, and the first picture in the multiple image generation information items; display at least one second picture on the second page.
[0087] In this embodiment, when the processor responds to an image generation instruction initiated for the first picture among the multiple pictures, calls the image reverse inference model, reverse infers the image generation prompt word and image generation parameter information according to the first picture, and jumps from the first page to the second page, it is specifically used for: in response to a selection operation for the first picture, display a first image generation control, and the first image generation control is associated with the second page; in response to a trigger operation for the first image generation control, call the image reverse inference model, reverse infer the image generation prompt word and image generation parameter information according to the first picture, and jump from the first page to the second page.
[0088] Optionally, when the processor displays the first image generation control on the first image in response to a selection operation on the first image, it is specifically configured to: in response to a hovering operation on the first image, display a first floating layer on the first image, where the first floating layer includes the first image generation control; or the first page includes a second image generation control or an image generation menu control, and in response to a triggering operation on the second image generation control or the image generation menu control, display an image upload page; in response to an operation of dragging the first image to the image upload page, display the first image on the image upload page; where the first image generation control is displayed on the image upload page.
[0089] Further optionally, the image upload page further includes: at least one image retouching control; after the first image is displayed on the image upload page, the processor is further configured to: in response to a triggering operation on any image retouching control, perform image retouching processing on the first image to obtain the retouched first image.
[0090] Optionally, when the processor displays the first image generation control in response to a selection operation on the first image, it is specifically configured to: in response to a selection operation on the first image, simultaneously display the first image generation control and a magnification control; the method further includes: in response to a triggering operation on the magnification control, jump from the first page to the second page and fill the first image on the second page; in response to a configuration operation of configuring magnification parameters on the second page, perform a magnification operation on the first image according to the configured magnification parameters and display the magnified first image on the second page.
[0091] Further optionally, the target application is a browser, and the first page and the second page are the first web page and the second web page respectively; where a target plugin is embedded in the browser, then when the processor calls the image reverse inference model, reversely infers the image generation prompt word and the image generation parameter information according to the first image, and jumps from the first page to the second page; and fills the image generation prompt word, the image generation parameter information, and the first image as a reference image into multiple corresponding image generation information items on the second page, it is specifically configured to: use the target plugin to call the image reverse inference model and reversely infer the image generation prompt word and the image generation parameter information according to the first image; use the target plugin to jump from the first page to the second page; use the target plugin to fill the image generation prompt word, the image generation parameter information, and the first image as a reference image into multiple corresponding image generation information items on the second page.
[0092] Further optionally, the first web page and the second web page are web pages in the same website; or, the first web page and the second web page are web pages in different websites.
[0093] Further optionally, when the first web page and the second web page are web pages in different websites, when the processor runs the target plugin to perform each operation in the method, it is specifically used for: in response to the user's login operation, sending an authentication request to the website server where the second web page is located, where the authentication request includes information about the target plugin for the website server to perform authentication; receiving the authentication passed information returned by the website server, and running the target plugin to enable the target plugin to perform each operation in the method.
[0094] Further, as Figure 3 shown, the server further includes: other components such as a communication component 30c, a display 30d, a power supply component 30e, and an audio component 30f. Figure 3 Only some components are schematically shown, which does not mean that the electronic device only includes Figure 3 the components shown.
[0095] The detailed implementation manners and beneficial effects of the electronic device provided in the embodiments of the present application have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0096] An exemplary embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to implement the steps in the foregoing method embodiments.
[0097] An exemplary embodiment of the present application further provides a computer program product, which includes a computer program / instructions, which when executed by a processor enables the processor to implement the steps in the foregoing method embodiments.
[0098] The foregoing memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read only memory (EEPROM), an erasable programmable read only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disc.
[0099] The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0100] The above-mentioned display includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0101] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device where the power supply component is located.
[0102] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0103] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, Compact Disc Read-Only Memory (CD-ROM), optical memory, etc.) that contain computer-usable program code.
[0104] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0107] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and a memory.
[0108] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0109] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0110] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0111] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An information processing method, characterized in that: include: In response to an access operation to a target application, a first page is displayed, wherein the first page includes a plurality of pictures, each picture having a respective picture style; In response to a raw image instruction initiated for a first image among the multiple images, calling a picture inversion model, inverting the first image to obtain raw image prompt words and raw image parameter information, and jumping from the first page to a second page; Filling the raw image prompt words, the raw image parameter information and the first image as a reference image into the corresponding multiple raw image information items on the second page; In response to the image generation operation initiated on the second page, calling the image generation model to generate at least one second image with a similar style to the first image according to the image generation prompt words, the image generation parameter information and the first image in the plurality of image generation information items; The at least one second picture is displayed on the second page.
2. The method according to claim 1, characterized in that In response to a raw image instruction initiated for a first image among the multiple images, calling a picture inversion model, inverting the first image to obtain raw image prompt words and raw image parameter information, and jumping from the first page to the second page, including: In response to a selection operation on the first picture, displaying a first image generation control, wherein the first image generation control is associated with the second page; In response to the trigger operation on the first raw image control, the image inversion model is called, raw image prompt words and raw image parameter information are obtained based on the first image, and the page is jumped from the first page to the second page.
3. The method according to claim 2, characterized in that In response to a selection operation on the first picture, displaying a first image control on the first picture includes: In response to a hover operation on the first picture, displaying a first floating layer on the first picture, wherein the first floating layer includes the first image generation control; or The first page includes a second raw image control or a raw image menu control, and in response to a triggering operation of the second raw image control or the raw image menu control, a picture upload page is displayed; in response to an operation of dragging the first picture to the picture upload page, the first picture is displayed on the picture upload page; wherein the first raw image control is displayed on the picture upload page.
4. The method according to claim 3, characterized in that The picture upload page further includes: at least one picture editing control; after the first picture is displayed on the picture upload page, the following further includes: In response to a triggering operation on any photo editing control, the first image is edited to obtain a first image after editing.
5. The method according to claim 2, characterized in that: In response to the selection operation on the first picture, displaying a first raw image control includes: in response to the selection operation on the first picture, displaying the first raw image control and the zoom control at the same time; The method further comprises: In response to a triggering operation on the magnification control, jumping from the first page to a second page, and filling the first picture onto the second page; In response to a configuration operation of configuring the zoom parameters on the second page, a zoom operation is performed on the first image according to the configured zoom parameters, and the zoomed first image is displayed on the second page.
6. The method according to any one of claims 1 to 5, characterized in that: The target application is a browser, the first page and the second page are respectively a first web page and a second web page; The browser is embedded with a target plug-in, calling a picture inversion model, inverting the raw picture prompt words and raw picture parameter information according to the first picture, and jumping from the first page to the second page; and filling the raw picture prompt words, the raw picture parameter information and the first picture as a reference picture in the corresponding multiple raw picture information items on the second page, including: Using the target plug-in, calling the image inversion model, and inverting the raw image prompt words and raw image parameter information according to the first image; Using the target plug-in, jumping from the first page to the second page; By using the target plug-in, the raw image prompt words, the raw image parameter information and the first image as a reference image are filled into the corresponding multiple raw image information items on the second page.
7. The method according to claim 6, characterized in that The first web page and the second web page are web pages in the same website; or, the first web page and the second web page are web pages in different websites.
8. The method according to claim 7, characterized in that When the first web page and the second web page are web pages in different websites, running the target plug-in to perform various operations in the method includes: In response to a login operation by the user, initiating an authentication request to the website server where the second web page is located, wherein the authentication request includes information of the target plug-in for authentication by the website server; Receive authentication pass information returned by the website server, and run the target plug-in to enable the target plug-in to perform various operations in the method.
9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps in the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to implement the steps in the method according to any one of claims 1 to 8.
11. A computer program product, characterized in that The computer program product comprises a computer program / instruction, and when the computer program / instruction is executed by a processor, the processor is enabled to implement the steps in any one of the methods of claims 1-8.
Citation Information
Cited By
Interaction method and device for generating Web end by AI content, and electronic equipment
CN121433484A
Interaction system and method for generating Web end by AI content, and electronic equipment
CN121433485A
Multimedia content generation method and device, electronic equipment and storage medium
CN121527240A