Intention recognition method, image editing method, model training method and device

By identifying the intention of the sliding operation and performing image editing, the problem that users need to manually enter text prompt words in the prior art is solved, which improves the user experience and the convenience of image editing.

CN120161983APending Publication Date: 2025-06-17NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510068669.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, users need to manually input text prompt words to express image processing intentions, resulting in poor user experience.

Method used

By obtaining sliding operation information and images, the trained target generation language model recognizes the intention of sliding operation and realizes image editing.

Benefits of technology

Users can directly express image processing intentions through sliding operations, improve user experience, and achieve the convenience of image editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120161983A_ABST
    Figure CN120161983A_ABST
Patent Text Reader

Abstract

The invention provides an intention recognition method, an image editing method, a model training method and devices thereof, and relates to the technical field of artificial intelligence. First sliding operation information and a first image are obtained, and the first sliding operation information is used for indicating a sliding path of a first sliding operation acting on the first image; the first sliding operation information and the first image are input into a trained target generative language model, the target generative language model is used for outputting an intention recognition result based on the first sliding operation information and the first image, and the intention recognition result is used for indicating the intention of the first sliding operation acting on the first image, according to the method and the device, the intention recognition result output by the target generative language model is obtained, that is, the user can output the intention recognition result of the user through the sliding operation acting on the image, so that the user expresses the processing intention of the image through the sliding operation acting on the image, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an intention recognition method, an image editing method, a model training method, and an apparatus thereof. Background Art

[0002] Currently, a user can process an image on an electronic device, such as performing image editing. In the related art, the user needs to manually input a text prompt to express the user's intention for processing the image, so as to guide the electronic device to process the image according to the user's intention.

[0003] However, expressing the user's intention for processing the image by manually inputting a text prompt results in a poor user experience. Summary of the Invention

[0004] In view of this, an object of the present invention is to provide an intention recognition method, an image editing method, a model training method, and an apparatus thereof, which can enable a user to express the intention for processing an image by a sliding operation on the image, thereby improving the user experience.

[0005] In a first aspect, an embodiment of the present invention provides an intention recognition method, including: obtaining first sliding operation information and a first image, where the first sliding operation information is used to: indicate a sliding path of a first sliding operation acting on the first image; inputting the first sliding operation information and the first image into a trained target generative language model, where the target generative language model is used to: output an intention recognition result based on the first sliding operation information and the first image, and the intention recognition result is used to: indicate the intention of the first sliding operation acting on the first image; obtaining the intention recognition result output by the target generative language model.

[0006] In a possible implementation manner, the first image input into the target generative language model is superimposed with strokes of a sliding path, the first sliding operation information includes a plurality of position information, and the plurality of position information is at least used to: indicate a target area formed by the sliding path. Inputting the first sliding operation information and the first image into the trained target generative language model includes: inputting the plurality of position information and the first image into the trained target generative language model, where the target generative language model is used to: obtain the strokes of the sliding path and a regional image of the first image in the target area based on the plurality of position information, and output an intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

[0007] In a possible implementation, the multiple location information includes the first location information of the target area and the second location information of the target area. The distance between the first location information and the second location information is not less than the distance between any two points in the target area. Inputting the multiple location information and the first image into the trained target generative language model includes: inputting the first location information, the second location information, and the first image into the trained target generative language model. The target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the first location information and the second location information, and output an intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

[0008] In a possible implementation, the first sliding operation is carried by a canvas. The canvas includes a first edge and a second edge. The first location information includes a first coordinate value in a first direction and a second coordinate value in a second direction. The second location information includes a third coordinate value in the first direction and a fourth coordinate value in the second direction. The first direction is parallel to the first edge and the second direction is parallel to the second edge. Before inputting the first location information, the second location information, and the first image into the trained target generative language model, the method further includes: performing normalization processing on the first location information and the second location information respectively: performing normalization processing on the first coordinate value and the third coordinate value based on the length of the first edge, and performing normalization processing on the second coordinate value and the fourth coordinate value based on the length of the second edge; inputting the first location information, the second location information, and the first image into the trained target generative language model includes: inputting the normalized first location information, the normalized second location information, and the first image into the trained target generative language model. The target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the normalized first location information and the normalized second location information, and output an intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

[0009] In a possible implementation, before inputting the first sliding operation information and the first image into the trained target generative language model, the method further includes: performing image segmentation on the first image to obtain M first regional images. Each first regional image includes a first object, and M is a positive integer; inputting the first sliding operation information and the first image into the trained target generative language model includes: inputting the first sliding operation information and the first target regional image among the M first regional images into the trained target generative language model. The first target regional image includes the first regional image on which the first sliding operation acts. The target generative language model is used to: output an intention recognition result based on the first sliding operation information and the first target regional image among the M first regional images.

[0010] In a possible implementation, the first sliding operation information and the first image are input into a trained target generative language model, including: embedding the first sliding operation information into a prompt information template to obtain prompt information, and inputting the prompt information and the first image into the trained target generative language model. The target generative language model is used to: output an intention recognition result based on the prompt information, the first sliding operation information, and the first image. The prompt information template is used to: guide the target generative language model to output an intention recognition result. The target generative language model is used to: output an intention recognition result based on the prompt information, the first sliding operation information, and the first image.

[0011] In a possible implementation, embedding the first sliding operation information into a prompt information template to obtain prompt information, and inputting the prompt information and the first image into a trained target generative language model, including: embedding the first sliding operation information into a first prompt information template to obtain first prompt information, and inputting the first prompt information and the first image into the trained target generative language model. The target generative language model is used to: output an intention recognition result based on the first prompt information and the first image. The first prompt information template is used to: indicate that the intention recognition result output by the target generative language model relates to the editing parameters selected when the first sliding operation acts on the first image. The editing parameters include at least one of color, transparency, line type, or line width; or, embedding the first sliding operation information into a second prompt information template to obtain second prompt information, and inputting the second prompt information and the first image into the trained target generative language model. The target generative language model is used to: output an intention recognition result based on the second prompt information and the first image. The second prompt information template is used to: indicate the editing result of the first sliding operation acting on the first image.

[0012] In a second aspect, an embodiment of the present invention provides an image editing method, which provides a graphical user interface through a terminal device; a first image is displayed in the graphical user interface; the method includes: responding to a first sliding operation acting on the first image, obtaining an intention recognition result based on the first sliding operation information corresponding to the first sliding operation and the first image, and the intention recognition result is obtained based on the method in the first aspect; generating an image editing instruction based on the intention recognition result, and the image editing instruction is used to: indicate to edit the first image; editing the first image based on the image editing instruction to obtain a second image.

[0013] In a possible implementation, the first image includes a first object. Editing the first image based on an image editing instruction to obtain a second image includes: superimposing a second object on the first object based on the image editing instruction, where the second image includes the first object and the second object superimposed on the first object; or, editing the size of the first object based on the image editing instruction, where the second image includes the first object with the size edited; or, editing the shape of the first object based on the image editing instruction, where the second image includes the first object with the shape edited.

[0014] In a possible implementation, after editing the first image based on an image editing instruction to obtain a second image, the method further includes: in response to a re - editing instruction, updating the second image to a third image, where the re - editing instruction is used to: indicate re - editing the first image, and the third image is obtained by editing the first image based on the re - editing instruction.

[0015] In a possible implementation, the image editing instruction includes a first image editing instruction and a second image editing instruction, and the second image is obtained by editing the first image based on the first image editing instruction. Updating the second image to a third image in response to a re - editing instruction includes: displaying a first editing identifier and a second editing identifier through a graphical user interface, where the first editing identifier is used to: indicate the result of editing the first image according to the first image editing instruction; the second editing identifier is used to: indicate the result of editing the first image according to the second image editing instruction; in response to a first trigger operation acting on the second editing identifier, updating the second image to a third image, where the third image is obtained by editing the first image based on the second image editing instruction.

[0016] In a possible implementation, the first editing identifier includes a first thumbnail, which is a thumbnail after editing the first image according to the first image editing instruction, and the second editing identifier includes a second thumbnail, which is a thumbnail after editing the first image according to the second image editing instruction.

[0017] In a possible implementation, before generating an image editing instruction based on an intention recognition result, the method further includes: displaying an editing parameter identifier through a graphical user interface, where the editing parameter identifier is used to: indicate the editing parameters used for editing the first image, and the editing parameters include at least one of color, transparency, line type, or line width; receiving a second trigger operation acting on the editing parameter identifier; generating an image editing instruction based on the intention recognition result, including: generating an image editing instruction based on the editing parameters indicated by the editing parameter identifier and the intention recognition result.

[0018] In a third aspect, an embodiment of the present invention provides a model training method, including: obtaining a pre-trained initial generative language model; obtaining a training sample set, the training sample set including a plurality of training samples, each training sample including a sample image and a sample label corresponding to the sample image, the sample image including a fourth image and a target line superimposed on the fourth image, the target line being used for: representing a second sliding operation applied to the fourth image, and the sample label being used for: indicating the intention of processing the fourth image; fine-tuning the initial generative language model using the training sample set to obtain a trained target generative language model, the target generative language model being used for identifying the intention of a first sliding operation applied to a first image.

[0019] In a possible implementation manner, before fine-tuning the initial generative language model using the training sample set to obtain a trained target generative language model, the method further includes: performing image segmentation on the sample image to obtain N second region images, each second region image including a third object, where N is a positive integer; fine-tuning the initial generative language model using the training sample set to obtain a trained target generative language model, including: fine-tuning the initial generative language model using the second target region image in the sample image corresponding to the training sample and the sample label corresponding to the sample image, the second target region image including the second region image covered by the target line.

[0020] In a possible implementation manner, the sample image is obtained by the following method: obtaining a fifth image; performing image segmentation on the fifth image to obtain L third region images, each third region image including a fourth object, where L is a positive integer; for at least one third target region image among the L third region images, obtaining the contour of the fourth object in the third target region image; superimposing the contour of the fourth object on a plurality of different fourth images to obtain a plurality of sample images.

[0021] In a possible implementation manner, after superimposing the contour of the fourth object on a plurality of different fourth images to obtain a plurality of sample images, the method further includes: performing adjustment processing on the sample images to fine-tune the initial generative language model using the adjusted sample images, the adjustment processing including at least one of blending the fourth image and the contour with transparency or randomly perturbing the lines of the contour.

[0022] In a possible implementation manner, the method further includes: for each third region image, determining a density value of the edge pixels of the third region image, the density value being used for: indicating the ratio of the number of edge pixels to the total number of pixels of the third region image; based on the density values corresponding to the L third region images, determining the third target region image among the L third region images, the density value corresponding to the third target region image being not less than the images other than the third target region image among the L third region images.

[0023] Fourthly, an embodiment of the present invention provides an intention recognition device, including: a first acquisition module, configured to acquire first sliding operation information and a first image, where the first sliding operation information is used to: indicate a sliding path of a first sliding operation acting on the first image; an intention recognition module, configured to input the first sliding operation information and the first image into a trained target generative language model, where the target generative language model is used to: output an intention recognition result based on the first sliding operation information and the first image, and the intention recognition result is used to: indicate an intention of the first sliding operation acting on the first image; and obtain the intention recognition result output by the target generative language model.

[0024] Fifthly, an embodiment of the present invention provides an image editing device, which provides a graphical user interface through a terminal device; the graphical user interface displays a first image; the device includes: a response module, configured to respond to a first sliding operation acting on the first image, and obtain an intention recognition result based on the first sliding operation information corresponding to the first sliding operation and the first image, where the intention recognition result is obtained based on the method in the first aspect; an image editing module, configured to generate an image editing instruction based on the intention recognition result, where the image editing instruction is used to: indicate to edit the first image; and edit the first image based on the image editing instruction to obtain a second image.

[0025] Sixthly, an embodiment of the present invention provides a model training device, including: a second acquisition module, configured to acquire a pre-trained initial generative language model; acquire a training sample set, where the training sample set includes a plurality of training samples, and each training sample includes a sample image and a sample label corresponding to the sample image, the sample image includes a fourth image and a target line superimposed on the fourth image, and the target line is used to: represent a second sliding operation acting on the fourth image, and the sample label is used to: indicate an intention of processing the fourth image; a model training module, configured to fine-tune the initial generative language model by using the training sample set to obtain a trained target generative language model, where the target generative language model is used to recognize an intention of a first sliding operation acting on a first image.

[0026] Seventhly, an embodiment of the present invention provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above method.

[0027] Eighthly, an embodiment of the present invention provides a machine-readable storage medium, which stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above method.

[0028] In a ninth aspect, an embodiment of the present invention provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to execute the above method.

[0029] The embodiments of the present invention bring the following beneficial effects: By obtaining first sliding operation information and a first image, the first sliding operation information is used to indicate a sliding path of a first sliding operation acting on the first image. The first sliding operation information and the first image are input into a trained target generative language model. The target generative language model is used to output an intention recognition result based on the first sliding operation information and the first image. The intention recognition result is used to indicate the intention of the first sliding operation acting on the first image. The intention recognition result output by the target generative language model is obtained. That is to say, the user can output the intention recognition result of the user through the sliding operation acting on the image. In this way, the user can express the processing intention of the image through the sliding operation acting on the image, thereby improving the user experience.

[0030] Other features and advantages of the present invention will be described in the following specification, and some of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.

[0031] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a flowchart of an intention recognition method provided by an embodiment of the present invention; Figure 2 It is a flowchart of an image editing method provided by an embodiment of the present invention; Figure 3 It is a flowchart of a model training method provided by an embodiment of the present invention; Figure 4 It is a schematic diagram of the interface display of an image editing provided by an embodiment of the present invention; Figure 5Another schematic diagram of the interface display for image editing provided by the embodiments of the present invention; Figure 6 Schematic structural diagram of an intention recognition device provided by the embodiments of the present invention; Figure 7 Schematic structural diagram of an image editing device provided by the embodiments of the present invention; Figure 8 Schematic structural diagram of a model training device provided by the embodiments of the present invention; Figure 9 Schematic structural diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] In the related art, a user needs to manually input text prompts to express an intention, and then guide an electronic device to perform image processing. For example, to guide the electronic device to perform image editing. For example, the user manually inputs text prompts such as "draw a sun in the sky" to express the intention of the user to draw a sun in the sky of the image, and then guide the electronic device to add a sun in the sky of the image; another example is that the user manually inputs text prompts such as "draw a ring on the finger of the person" to express the intention of the user to draw a ring on the finger of the person in the image, and then guide the electronic device to add a ring on the finger of the person included in the image.

[0036] However, expressing the user's intention by manually inputting text prompts and then guiding the electronic device to perform image processing will result in poor convenience for the user to express the intention of image processing, and thus poor user experience.

[0037] Based on this, an intention recognition method, an image editing method, a model training method, and their devices provided by the embodiments of the present invention can enable a user to express the intention of image processing through a sliding operation on the image, and then process the image according to the user's processing intention. For example, the user can perform image editing through a sliding operation on the image, thereby improving the user experience.

[0038] In one embodiment of the present invention, the intention recognition method can be run on a local terminal device or a server. When the intention recognition method is run on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device (also called a local terminal device or a terminal device).

[0039] In an optional implementation, various cloud applications can be run under the cloud interaction system, such as cloud games. Taking cloud games as an example, cloud games refer to a game mode based on cloud computing. In the operation mode of cloud games, the operating body of the game program and the main body of the game screen presentation are separated. The storage and operation of the image editing method are completed on the cloud game server. The role of the client device is used for receiving and sending data and presenting the game screen. For example, the client device can be a display device with data transmission function close to the user side, such as a mobile terminal, a TV, a computer, a handheld computer, etc.; but the cloud game server in the cloud is used for information processing. When playing the game, the player operates the client device to send an operation instruction to the cloud game server. The cloud game server runs the game according to the operation instruction, encodes and compresses the game screen and other data, and returns it to the client device through the network. Finally, the client device decodes and outputs the game screen.

[0040] In an optional embodiment, taking a game as an example, a local terminal device stores a game program and is used to present a game screen. The local terminal device is used to interact with the player through a graphical user interface, that is, the game program is downloaded and installed by an electronic device and run conventionally. The local terminal device may provide the graphical user interface to the player in a variety of ways, for example, it may be rendered and displayed on a display screen of the terminal, or provided to the player through a holographic projection. For example, the local terminal device may include a display screen and a processor, the display screen is used to present a graphical user interface, the graphical user interface includes a game screen, and the processor is used to run the game, generate a graphical user interface, and control the display of the graphical user interface on the display screen.

[0041] To facilitate understanding of this embodiment, an intention recognition method disclosed in an embodiment of the present invention is first exemplified. The intention recognition method of this embodiment can be executed by a server, or by a client device, or by a server and a client device jointly. Figure 1 , Figure 1 The following is a flow chart of an intention recognition method provided by an embodiment of the present invention. Figure 1 The intent recognition method shown may include: S110: Acquire first sliding operation information and a first image, where the first sliding operation information is used to indicate a sliding path of the first sliding operation on the first image.

[0042] Among them, the first operation sliding information can be the information corresponding to the first sliding operation. The first sliding operation can be a sliding operation by the user on the first image. Among them, the sliding path is the sliding motion of an object on a specific track or path. In this embodiment, the sliding path is the path of the first sliding operation acting on the first image. Exemplarily, the first downward sliding operation can be, for example, a downward swipe, an upward swipe, or a sliding operation that constitutes the contour of an object, which is not limited herein. The sliding operation that constitutes the contour of an object can be, for example, a sliding operation of drawing a circle or a sliding operation of drawing a skirt.

[0043] It should be noted that the first sliding operation in this embodiment can be an operation directly by the user on the first image, such as an operation of the user's finger on the first image; in addition, it can also be an operation indirectly by the user on the first image through an external device, such as an operation indirectly by the user on the first image through a stylus or a mouse.

[0044] S120. Input the first sliding operation information and the first image into the trained target generative language model. The target generative language model is used to: output an intention recognition result based on the first sliding operation information and the first image. The intention recognition result is used to: indicate the intention of the first sliding operation acting on the first image.

[0045] Among them, the generative language model is a model that can generate text sequences that conform to language rules. It is based on the method of probability statistics and learns the vocabulary, grammar, and semantic rules in the language, so as to be able to generate new and reasonable text. Exemplarily, the generative language model in this embodiment can be, for example, a multi-modal large language model (MLLM). In this embodiment, the target generative language model is used to output an intention recognition result based on the first sliding operation information and the first image.

[0046] In this embodiment, the intention recognition result indicates the intention of the first sliding operation acting on the first image. That is to say, the intention recognition result can indicate what the intention of the first sliding operation by the user on the first image is. In this way, it is beneficial to process the first image using the intention recognition result. Optionally, the intention recognition result can indicate editing the first image, or indicating obtaining information of the first image, or indicating retrieving images similar to the first image, etc. This embodiment does not limit the specific intention recognition result.

[0047] Exemplarily, if the first sliding operation of the user on the first image is a sliding operation corresponding to a circle, the intention recognition result may be to edit the first image; if the first sliding operation of the user on the first image is a sliding operation corresponding to a triangle, the intention recognition result may be to obtain information about the first image; if the first sliding operation of the user on the first image is a sliding operation corresponding to a square, the intention recognition result may be to retrieve images similar to the first image, etc.

[0048] Exemplarily again, taking the intention recognition result of editing the first image as an example, it may be that the user makes a sliding operation corresponding to drawing a circle in the sky of the first image, and the intention recognition result may be: adding a sun in the sky; for another example, the user makes a sliding operation corresponding to drawing a circle in the night of the first image, and the intention recognition result may be: drawing a moon in the night; for another example, the user makes a sliding operation corresponding to drawing a circle in the night of the first image, and the intention recognition result may be: drawing a moon in the night; for another example, the user makes a sliding operation corresponding to drawing a circle on the person in the first image, and the intention recognition result may be: drawing a ring on the finger of the person.

[0049] Exemplarily again, taking the intention recognition result of editing the first image as an example, it may be that the user applies a downward sliding operation to the skirt in the first image, and the intention recognition result may be to lengthen the length of the skirt; for another example, the user applies an upward sliding operation to the skirt in the first image, and the intention recognition may be to shorten the length of the skirt, etc.

[0050] It should be understood that the intention recognition result that the target generative language model can output is related to the sample label of the training sample used during the training of the target generative language model. The sample label used during model training can be determined according to the intention to be recognized, and no limitation is made here.

[0051] S130. Obtain the intention recognition result output by the target generative language model.

[0052] The technical solution of this embodiment, by obtaining the first sliding operation information and the first image, where the first sliding operation information is used to indicate the sliding path of the first sliding operation acting on the first image, inputting the first sliding operation information and the first image into the trained target generative language model, the target generative language model is used to output an intention recognition result based on the first sliding operation information and the first image, and the intention recognition result is used to indicate the intention of the first sliding operation acting on the first image, and obtaining the intention recognition result output by the target generative language model. That is to say, the user can output the intention recognition result of the user through the sliding operation acting on the image. In this way, the user can express the processing intention of the image through the sliding operation acting on the image, thereby improving the user experience.

[0053] In a possible implementation, if there are multiple intention recognition results, the target generative language model also outputs the confidence levels of each intention recognition result among the multiple intention recognition results. The confidence level is used to indicate the reliability of the intention recognition result. The multiple intention recognition results can refer to the description in the above embodiments and will not be elaborated here.

[0054] For example, if the multiple intention recognition results include a first intention recognition result and a second intention recognition result, the target generative language model also outputs a first confidence level corresponding to the first intention recognition result and a second confidence level corresponding to the second intention recognition result.

[0055] Optionally, the target generative language model can output the intention recognition results with corresponding confidence levels higher than the confidence level threshold. For example, if the first confidence level is greater than the confidence level threshold and the second confidence level is less than the confidence level threshold, the first intention recognition result is output; or if both the first confidence level and the second confidence level are greater than the confidence level threshold, the first intention recognition result and the second intention recognition result are output. It should be understood that the size of the confidence level threshold can be set as needed, such as set to 60% or 80%, etc., and is not limited here.

[0056] Optionally, the target generative language model can output the intention recognition results with higher confidence levels, that is, output the intention recognition results with higher confidence levels. For example, output the intention recognition results ranked in the top 50% in terms of confidence level.

[0057] It should be noted that in this embodiment, the sum of the confidence levels of all intention recognition results can be 1, or can also be not 1, such as greater than 1 or less than 1.

[0058] In a possible implementation, strokes of a sliding path are superimposed on a first image input to the target generative language model. The first sliding operation information includes multiple position information, and the multiple position information is at least used to: indicate the target area formed by the sliding path. Inputting the first sliding operation information and the first image into the trained target generative language model includes: Inputting the multiple position information and the first image into the trained target generative language model, where the target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the multiple position information, and output an intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

[0059] In this embodiment, multiple location information and a first image are input into a trained target generative language model. Then, based on the multiple location information, the target generative language can determine the target area formed by the sliding path. And the first image input into the target generative language model has the strokes of the sliding path superimposed on it. Then, the target generative model can use this target area to obtain the strokes of the sliding path and the regional image of the first image in the target area, and further output an intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

[0060] In this embodiment, the first image with the strokes of the sliding path superimposed on it and multiple location information indicating the target area formed by the sliding path can be input into the target generative language model. In this way, it is possible to quickly locate the strokes of the sliding path and the regional image of the first image in the target area using the multiple location information, which can improve the efficiency of the obtained intention recognition result.

[0061] In another possible implementation, the multiple location information can be multiple consecutive location information on the sliding path. That is to say, the multiple location information forms a coordinate sequence of the sliding path. In this way, the target generative language model can obtain the strokes that the sliding path can form and the position where the sliding path acts on the first image through the multiple location information, and then output an intention recognition result using the strokes corresponding to the sliding path and the position of the strokes of the sliding path in the first image. In this embodiment, the first image input into the target generative language model may not have the strokes of the sliding path superimposed on it. That is to say, the input to the target generative language model can be the original first image.

[0062] Among them, the target area can be a regular area or an irregular area. Regular areas can include but are not limited to rectangular areas, circular areas, etc. In this embodiment, the target area can at least cover the area passed by the sliding path. Optionally, the target area can also cover the adjacent area of the area passed by the sliding path. In this way, when the target generative language model outputs an intention recognition result, it can obtain more image information, thereby improving the accuracy of the output intention recognition result. Optionally, the target area is smaller than the maximum area covered by the first image.

[0063] In this embodiment, multiple consecutive location information on the sliding path and the first image without the strokes of the sliding path superimposed on it can be input into the target generative language model. In this way, the amount of data input into the target generative language model can be reduced. In this way, when the client device and the server interact to implement image processing, for example, when the server implements intention recognition, the amount of information sent by the client device to the server is also less, reducing the transmission resources required for the cloud interaction system to implement intention recognition.

[0064] In a possible implementation, the multiple location information includes the first location information of the target area and the second location information of the target area. The distance between the first location information and the second location information is not less than the distance between any two points in the target area. Inputting the multiple location information and the first image into the trained target generative language model includes: Inputting the first location information, the second location information, and the first image into the trained target generative language model. The target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the first location information and the second location information, and output an intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

[0065] Optionally, the target area in this embodiment may be a rectangular area. Then, the first location information may be, for example, the location information of the upper left corner of the target area, and the second location information may be the location information of the lower right corner of the target area. In this embodiment, by inputting the location information of the two points with the farthest distance in the target area into the target generative language model, the area restored by the target generative language model based on the location information of the two points with the farthest distance may include the target area, and the result of intention recognition is also more accurate. And inputting the location information of the two points with the farthest distance in the target area into the target generative language model, in this way, the data volume of the location information input into the target generative language model is also less. Therefore, it is possible to reduce the data volume of the location information input into the target generative language model and improve the intention recognition result.

[0066] In another possible implementation, the multiple location information may further include the third location information and the fourth location information in the target area, etc., which is not limited here.

[0067] In a possible implementation, the first sliding operation is carried by a canvas. The canvas includes a first edge and a second edge. The first location information includes the first coordinate value in the first direction and the second coordinate value in the second direction. The second location information includes the third coordinate value in the first direction and the fourth coordinate value in the second direction. The first direction is parallel to the first edge and the second direction is parallel to the second edge. Before inputting the first location information, the second location information, and the first image into the trained target generative language model, the method further includes: Normalizing the first location information and the second location information respectively: normalizing the first coordinate value and the third coordinate value based on the length of the first edge, and normalizing the second coordinate value and the fourth coordinate value based on the length of the second edge.

[0068] In this embodiment, the canvas can be a rectangular canvas, and the first edge of the canvas and the second edge of the canvas can be perpendicular to each other. Normalizing the first coordinate value and the third coordinate value respectively based on the length of the first edge, the normalized first coordinate value can be positively correlated with the first coordinate value before normalization and negatively correlated with the length of the first edge; and, the normalized third coordinate value can be positively correlated with the third coordinate value before normalization and negatively correlated with the length of the first edge. The first edge can be the width of the canvas, and the second edge can be the height of the canvas.

[0069] Normalizing the second coordinate value and the fourth coordinate value respectively based on the length of the second edge, the normalized second coordinate value can be positively correlated with the second coordinate value before normalization and negatively correlated with the length of the second edge; and, the normalized fourth coordinate value can be positively correlated with the fourth coordinate value before normalization and negatively correlated with the length of the second edge.

[0070] Exemplarily, if the first position information before normalization is (x1, y1) and the second position information before normalization is (x2, y2), then the first position information after normalization is (x1 / W, y1 / H), and the second position information after normalization is (x2 / W, y2 / H), where W represents the width of the canvas and H represents the height of the canvas.

[0071] Correspondingly, inputting the first position information, the second position information, and the first image into the trained target generative language model includes: Inputting the normalized first position information, the normalized second position information, and the first image into the trained target generative language model, where the target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the normalized first position information and the normalized second position information, and output an intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

[0072] In this embodiment, based on the normalized first position information and the normalized second position information, and combined with the size of the canvas, the target generative language model can restore the first position information before normalization and the second position information before normalization, and then locate the target area to output the intention recognition result.

[0073] In this embodiment, by normalizing the first position information and the second position information, the data volume of the position information input into the target generative language model can be further reduced. In this way, the transmission resources for realizing intention recognition through the cloud interaction system can be further reduced.

[0074] In another possible implementation, it is also possible to directly input the first position information before normalization processing and the second position information before normalization processing into the target generative language model, which can reduce the processing process of the position information and reduce the computing resources required for intent recognition.

[0075] In one possible implementation, before inputting the first sliding operation information and the first image into the trained target generative language model, the method further includes: Performing image segmentation on the first image to obtain M first regional images, each first regional image including a first object, where M is a positive integer.

[0076] Among them, the quantity of M is related to the result of image segmentation and is not limited here. In this embodiment, an image segmentation algorithm can be used to perform image segmentation on the first image. The image segmentation algorithm can include but is not limited to a threshold-based segmentation algorithm, an edge-based segmentation algorithm, etc. The first object can be, for example, an object in the first image, which can be understood as the target presented in the first image. Optionally, the first object can include but is not limited to a person, an item (such as a skirt, a cup), a weather element (such as the sky, a cloud), etc.

[0077] Correspondingly, inputting the first sliding operation information and the first image into the trained target generative language model includes: Inputting the first sliding operation information and the first target regional image among the M first regional images into the trained target generative language model, where the first target regional image includes the first regional image on which the first sliding operation acts, and the target generative language model is used to: output an intent recognition result based on the first sliding operation information and the first target regional image among the M first regional images.

[0078] In this embodiment, among the M first regional images after image segmentation, the first target regional image on which the first sliding operation acts can be input into the target generative language model, so that the data volume of the input image is also smaller, thereby reducing the transmission resources required for the cloud interaction system to implement intent recognition.

[0079] In another possible implementation, it is also possible to input the global image of the first image into the target generative language model, that is, there is no need to perform image segmentation processing on the first image. In this way, the computing resources required for image segmentation can be reduced, and further the computing resources required for intent recognition can be reduced. Moreover, the first image retains the global information, so that when the target generative language model outputs an intent recognition result, it can also refer to the global information of the first image, which can improve the accuracy of the intent recognition result output by the generative language model.

[0080] In a possible implementation, inputting the first sliding operation information and the first image into a trained target generative language model includes: Embedding the first sliding operation information into a prompt information template to obtain prompt information, and inputting the prompt information and the first image into a trained target generative language model. The target generative language model is used to: output an intention recognition result based on the prompt information, the first sliding operation information, and the first image. The prompt information template is used to: guide the target generative language model to output an intention recognition result. The target generative language model is used to: output an intention recognition result based on the prompt information, the first sliding operation information, and the first image.

[0081] In this embodiment, the prompt information can be a prompt. Among them, a prompt can be input text or an instruction in a large language model or a deep learning model for guiding the model to generate a specific type of response. A prompt can be a question, a description, a task description, or even a part of the conversation history, etc. In this embodiment, the prompt information template can be a pre-determined template. Embedding the first sliding operation information into the prompt information template can obtain prompt information, and then using the prompt information to guide the target generative language model to output an intention recognition result.

[0082] In this embodiment, by embedding the first sliding operation information into the prompt information template to obtain prompt information and inputting the prompt information and the first image into a trained target generative language model, in this way, it is possible to better guide the target generative language model to output an intention recognition result, and further improve the accuracy of the intention recognition result output by the target generative language model.

[0083] In a possible implementation, embedding the first sliding operation information into a prompt information template to obtain prompt information and inputting the prompt information and the first image into a trained target generative language model includes: Embedding the first sliding operation information into a first prompt information template to obtain first prompt information, and inputting the first prompt information and the first image into a trained target generative language model. The target generative language model is used to: output an intention recognition result based on the first prompt information and the first image. The first prompt information template is used to: indicate that the intention recognition result output by the target generative language model involves the editing parameters selected when the first sliding operation acts on the first image. The editing parameters include at least one of color, transparency, line type, or line width; or, Embedding the first sliding operation information into a second prompt information template to obtain second prompt information, and inputting the second prompt information and the first image into a trained target generative language model. The target generative language model is used to: output an intention recognition result based on the second prompt information and the first image. The second prompt information template is used to: indicate the editing result of the first sliding operation acting on the first image.

[0084] Exemplarily, the first prompt message template can be, for example: "The user has uploaded an image containing a [color] contour. To assist in locating the contour, the normalized bounding box coordinates of the contour are: upper left corner coordinates: (), lower right corner coordinates: (). Please identify and describe the content within the contour with a single word or phrase." Then, if the user selects red, the first prompt message can be: "The user has uploaded an image containing a [red] contour. To assist in locating the contour, the normalized bounding box coordinates of the contour are: upper left corner coordinates: ({x1'}, {y1'}), lower right corner coordinates: ({x2'}, {y2'}). Please identify and describe the content within the contour with a single word or phrase." The first prompt message template can also be understood as a coloring stroke template.

[0085] Exemplarily, the second prompt message template can be, for example: "This is a 'draw and guess' game. I will upload an image containing strokes. To assist in locating the strokes, I will provide the normalized bounding box coordinates of the strokes: upper left corner coordinates: (), lower right corner coordinates: (). Please tell me with a single word or phrase what these strokes are trying to represent?" Then, the first prompt message can be: "I will upload an image containing strokes. To assist in locating the strokes, I will provide the normalized bounding box coordinates of the strokes: upper left corner coordinates: ({x1'}, {y1'}), lower right corner coordinates: ({x2'}). Please tell me with a single word or phrase what these strokes are trying to represent?" The second prompt message template can also be understood as a stroke template.

[0086] In this embodiment, different prompt message templates can be selected according to different situations to generate prompt messages. For example, when the user selects editing parameters, the first prompt message template is used to generate the first prompt message, and when the user does not select editing parameters, the second prompt message template is used to generate the second prompt message. In this way, the accuracy of intent recognition can be improved.

[0087] It should be noted that the intent recognition result can be represented in text, that is, the target generative language model can output intent text as the intent recognition result.

[0088] In a possible implementation manner, the intent text can also be optimized, for example, the output format is specified: limited to a single word or phrase. For example, simplify the complex description "a running brown puppy" to "puppy" to ensure the description is concise and clear. For example, "red sports shoes" instead of "a pair of sports shoes that look like red" to avoid complex modifiers. Optionally, "a very beautiful blue butterfly" can be simplified to "blue butterfly" to unify the expression format. Also, for color descriptions, uniformly use "red flower" instead of "red-colored flower".

[0089] Optionally, the methods for optimizing the results include, but are not limited to: removing redundant modifiers or extracting at least one of the core semantic components. Removing redundant modifiers, for example, can be optimizing "a cute little white rabbit" to "white rabbit". Extracting the core semantic components can be, for example, extracting "pink rose" from "the blooming pink rose".

[0090] In the above embodiments, the application of the trained target generative language model to intent recognition is described. Next, an exemplary description of the model training is provided.

[0091] The model training method of this embodiment can be executed by a server, or by a client device, or jointly by a server and a client device. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a model training method provided by an embodiment of this application. As Figure 2 shown, the model recognition method may include: S210. Obtain a pre-trained initial generative language model.

[0092] Among them, the initial generative language model can be a pre-trained generative language model, such as a general generative language model. Among them, a multimodal large language model with image understanding ability can be selected as the base model (initial generative language model). For example, it can be a vision language model based on the Transformer architecture, a bimodal model with image encoding and text generation capabilities, or a pre-trained model supporting image question answering. Among them, the model scale can be selected according to actual needs: for example, a lightweight model (3B - 7B parameters) is suitable for rapid deployment, a medium-scale model (10B - 20B parameters) balances performance and efficiency, and a large-scale model (30B+ parameters) pursues better results.

[0093] S220. Obtain a training sample set. The training sample set includes multiple training samples. Each training sample includes a sample image and a sample label corresponding to the sample image. The sample image includes a fourth image and target lines superimposed on the fourth image. The target lines are used to: represent a second sliding operation applied to the fourth image. The sample label is used to: indicate the intent to process the fourth image.

[0094] Among them, the intent to process the fourth image can be to extract information from the fourth image, retrieve an image matching the fourth image, or edit the fourth image, etc. The intent to process the fourth image can refer to the relevant description of the intent recognition result and will not be elaborated here.

[0095] It should be noted that in this embodiment, the target line is used to: represent the second sliding operation applied to the fourth image. It only represents the role of the target line in this solution, and does not mean that the target line is the line obtained by the user's actual second sliding operation on the fourth image.

[0096] S230. Fine-tune the initial generative language model using the training sample set to obtain a trained target generative language model, where the target generative language model is used to recognize the intention of the first sliding operation applied to the first image.

[0097] Among them, model fine-tuning refers to further training on the target data set based on the pre-trained model to make the model adapt to a specific task. Its principle is a specific implementation of transfer learning, that is, further adjusting and optimizing the parameters of the pre-trained model to adapt to the new task. The parameters of the pre-trained model have been obtained through training on a large amount of data, and only need to be fine-tuned on a small amount of data to quickly transfer learning to the target domain. Specifically, the fine-tuning methods can include but are not limited to full-scale fine-tuning or parameter-efficient fine-tuning. Full-scale fine-tuning uses specific task data to adjust all the parameters of the pre-trained model to fully adapt to the new task. Parameter-efficient fine-tuning realizes efficient transfer learning by minimizing the number of fine-tuned parameters and the computational complexity. Only part of the parameters in the model are updated, significantly reducing the training time and cost.

[0098] In this embodiment, fine-tuning is performed on the basis of the pre-trained initial generative language model to obtain a target generative language model that can adapt to the intention recognition task. The intention recognition task can include a hand-drawn understanding task. The hand-drawn understanding task is defined as: constructing a dialogue-based hand-drawn intention understanding system to convert the user's drawing actions into understandable editing instructions or output intention recognition results. How the target generative language model recognizes the intention of the first sliding operation applied to the first image can refer to the description of the above embodiment and will not be elaborated here.

[0099] The technical solution of this embodiment obtains a pre-trained initial generative language model, obtains a training sample set, where the training sample set includes multiple training samples, each training sample includes a sample image and a sample label corresponding to the sample image. The sample image includes a fourth image and a target line superimposed on the fourth image. The target line is used to: represent the second sliding operation applied to the fourth image. The sample label is used to: indicate the intention of processing the fourth image. The initial generative language model is fine-tuned using the training sample set to obtain a trained target generative language model, where the target generative language model is used to recognize the intention of the first sliding operation applied to the first image. In this way, the trained target generative language model can be used to recognize the intention of the first sliding operation applied to the first image. Thus, the user can express the intention of processing the image through the sliding operation on the image, thereby improving the user experience.

[0100] In this embodiment, a parameter-efficient fine-tuning method can be adopted to obtain the target generative language model. Specifically, low-rank adaptation technology can be selected, prompt engineering methods can be used, and quantization fine-tuning strategies can be adopted, etc. When performing fine-tuning, reasonable training parameters need to be designed. The learning rate range is recommended to be 0.0001 - 0.001, the training batch size is adjusted according to the video memory capacity, and the number of training epochs is determined through the validation set.

[0101] It should be noted that parameter-efficient fine-tuning can use parameter-efficient fine-tuning methods such as low-rank adaptation (LoRA), query-based low-rank adaptation (QLoRA), and potential tuning (P-Tuning). Taking LoRA as an example: The random number (rank) can be set to 8 for lightweight fine-tuning, the rank can be set to 16 to obtain better results, and the rank can be set to 32 to achieve stronger fitting ability. When setting training parameters, different optimizers and learning rates can be used. Taking the adaptive moment estimation optimizer with weight decay correction (AdamW) as an example: The learning rate can be set to 1e-4, the weight decay can be set to 0.01, and the number of warm-up steps can be set to 100. The "100 steps" here refers to the number of iterations or steps in the warm-up phase. During these 100 steps, the learning rate may gradually increase from a very small value to ensure that the model can transition smoothly at the initial stage of training and avoid problems such as gradient explosion or model divergence caused by too large a learning rate. After the warm-up phase ends, the learning rate will be adjusted according to the preset scheduling strategy to continue the model training process.

[0102] In this embodiment, fine-tuning the initial generative language model may be that the loss function meets a preset condition, for example, the training loss is less than the loss threshold. Design of the loss function: The principle of probability maximization can be adopted to optimize the prediction accuracy of the model for the input sequence: maximize(Θ_adapt) Σ log P(y_i | y_1,…,y_(i-1); {Θ_base, Θ_adapt}). Generally speaking, there is a pre-trained large model (parameters are Θ_base). When it is necessary to teach it to understand hand-drawn content, it is achieved by adjusting the new parameter Θ_adapt. Specifically, let the pre-trained model learn word by word (y_i is the current word, and y_1 to y_(i-1) are the previous words). Every time the pre-trained model guesses correctly, it is rewarded (the probability P is increased), and if it guesses wrong, it is corrected (the probability P is decreased). Through repeated practice (Σ summation), the model can learn to accurately understand the user's hand-drawn intention. Here, Θ_base represents the basic model parameters, and Θ_adapt represents the adaptive training parameters. That is to say, the Θ_base parameters remain unchanged, and the Θ_adapt parameters are updated during the fine-tuning process.

[0103] In a possible implementation manner, before fine-tuning the initial generative language model using the training sample set to obtain the trained target generative language model, the method further includes: Performing image segmentation on the sample image to obtain N second region images, where each second region image includes a third object, and N is a positive integer.

[0104] Among them, the sample image can be segmented by an image segmentation algorithm. The quantity of N is related to the actual image segmentation result and is not specifically limited here. The third object can refer to the description of the first object and will not be elaborated here.

[0105] Correspondingly, fine-tuning the initial generative language model using the training sample set to obtain the trained target generative language model includes: Fine-tuning the initial generative language model using the second target region image in the sample image corresponding to the training sample and the sample label corresponding to the sample image, where the second target region image includes the second region image covered by the target line.

[0106] In this embodiment, by performing image segmentation on the sample image to obtain N second region images, and fine-tuning the initial generative language model using the second region image covered by the target line in the sample image corresponding to the training sample and the sample label corresponding to the sample image, the data volume of the image input to the initial generative model can be reduced during model fine-tuning, and the model training efficiency can be improved.

[0107] In another possible implementation, it is also possible not to perform image segmentation on the sample image, that is, to fine-tune the initial generative language model using the global information of the sample image. In this way, the global information of the sample image can be obtained during fine-tuning, which is beneficial to improving the accuracy of intent recognition of the trained target generative language model.

[0108] In one possible implementation, the sample image is obtained in the following way: Obtain a fifth image; perform image segmentation on the fifth image to obtain L third-region images, each third-region image including a fourth object, where L is a positive integer; for at least one third target-region image among the L third-region images, obtain the contour of the fourth object in the third target-region image; superimpose the contour of the fourth object on a plurality of different fourth images to obtain a plurality of sample images.

[0109] Among them, the contour of the fourth object can be represented by the edge of the fourth object, for example, represented by the edge pixels of the fourth object. In this embodiment, the purpose of image segmentation may include: extracting the edges of objects in the image. The edge detection methods that can be used to obtain the contour of the fourth object in the third target region image in this embodiment may include, but are not limited to, pixel difference networks (PiDiNet), edge detection algorithms (Canny algorithm), Sobel operators, or holistically-nested edge detection (HED). PiDiNet neural network: suitable for accurate edge extraction in complex scenes. Canny algorithm: a classic gradient-based edge detection, suitable for simple scenes. Sobel operator: high computational efficiency, suitable for real-time processing. Deep learning methods can provide multi-scale edge information. In this embodiment, multiple edge detection methods can be used to process the input image, and a suitable edge detection method can be selected for edge detection according to the actual situation. The fourth object can refer to the description of the third object and will not be elaborated here. In this embodiment, the contour of the fourth object is superimposed on multiple different fourth images to obtain multiple sample images, and the fourth image can be understood as the background on which the contour of the fourth object is superimposed. It should be noted that the fourth image can be randomly generated as the background of the contour. Optionally, a high-quality random background can be generated based on the brush Net model based on stablediffusion (Brush Net model based on XLSDXL), or precise area redrawing can be achieved using the image inpainting model (Inpainting model of Stable Diffusion), or large-area background reconstruction can be performed using the Large Mask Inpainting (LaMa) model, or a natural-transition texture background can be generated using the mask-aware transformer for large hole image inpainting (MAT) model.

[0110] Specifically, after extracting the edge maps (also known as the contours of objects) from the image, the contours of those edge maps can be used as a simulation of the sliding operation. However, currently these edge maps are extracted based on the original image. For example, the background of the edge map of a building is indeed the building. Such data will lead to poor training effects and cannot be directly used for training. Therefore, it is hoped that the background corresponding to the edge map of this building is random, so that it can simulate the content drawn by the user in a random scene. For example, in a background of blue sky and white clouds, the user draws the contour (edge map) of a building.

[0111] In this embodiment, by performing image segmentation on the fifth image, L third-region images are obtained. For at least one third target-region image among the L third-region images, the contour of the fourth object in the third target-region image is acquired, and the contour of the fourth object is superimposed on multiple different fourth images to obtain multiple sample images. In this way, sample images can be generated for fine-tuning, which is conducive to improving the convenience of model fine-tuning and reducing the difficulty of obtaining training samples.

[0112] In another possible implementation, it can also be to collect the target line through the user's sliding operation.

[0113] In one possible implementation, after superimposing the contour of the fourth object on multiple different fourth images to obtain multiple sample images, the method further includes: Performing adjustment processing on the sample images to fine-tune the initial generative language model using the adjusted sample images. The adjustment processing includes at least one of blending the fourth image and the contour with transparency or randomly perturbing the lines of the contour.

[0114] Among them, transparency blending blends the object with the scene to produce a realistic visual effect. In transparency blending, the output color of each pixel is a mixture of the object color and the background color in a certain proportion, and this proportion is usually controlled by the transparency value (alpha value) of the object. Randomly perturbing the lines of the contour can be to adjust and change the lines of the contour. In this embodiment, blending the fourth image and the contour with transparency or randomly perturbing the lines of the contour can better simulate the user's sliding operation, thereby improving the effect of model fine-tuning and making the result of model fine-tuning closer to the actual situation.

[0115] In this embodiment, superimposing the previously extracted edge map (also known as the contour) on the fourth image can simulate the user's sliding operation on the image, such as simulating the effect of the user using a paintbrush tool for scribbling. Optionally, the add Weighted function of Open CV can be used to implement the transparency blending of the edge map. Specifically, the edge map can be converted to the RGBA format, the default value of the alpha channel is set to 0.7, and then the blending formula: result = alpha * edge+(1 - alpha)* background can be used. In addition, the cv2.Gaussian Blur function can be used to blur the edge pixels with a 3x3 kernel size.

[0116] Specifically, the add Weighted function of Open CV is used to blend the edge map and the background map according to the specified transparency (alpha value). Here, the edge map is first converted to the RGBA format, and the default value of the alpha channel is set to 0.7, which means that the edge map will cover the background map with a transparency of 70%. Before blending, the edge map can be blurred using cv2.Gaussian Blur to reduce the jaggedness of the edges and make the edges smoother.

[0117] In this embodiment, random perturbation of the edge lines can be achieved using NumPy. Optionally, a random offset value within the range of [-2, 2] can be added to the original edge coordinates, and the Bezier curve function of the scipy.interpolate module can be used for smoothing, and the line thickness can be randomly generated using random.rand int(2, 5).

[0118] Specifically, a random offset value is added to the original edge coordinates to increase the natural and hand-drawn feel of the lines. This randomness simulates the irregularity of the lines during hand-drawing. The perturbed lines are smoothed to further reduce the jaggedness of the lines and make them more in line with the characteristics of hand-drawn lines. In this embodiment, various stroke styles can be achieved in the following ways: for solid lines, cv2.line can be used for direct drawing; for dashed lines, cv2.line can be used in combination with line type=cv2.LINE_AA; for dotted lines, slicing operations of numpy can be used for rendering at intervals.

[0119] In this embodiment, the following scheme can be adopted to simulate the pressure effect: Min Max Scaler can be used to normalize the edge intensity to [0, 1], linear interpolation can be used to map the intensity to the alpha value range of [0.3, 0.9], and the line thickness can be calculated proportionally according to the normalized intensity value.

[0120] Specifically, Min Max Scaler is used to normalize the edge intensity to the range of [0, 1] to unify the measurement standard. The normalized intensity value is mapped to the alpha value range of [0.3, 0.9] through linear interpolation to simulate the pressure change during hand-drawing. The greater the intensity, the greater the alpha value, which means the line is more opaque, simulating the effect that the harder the hand-drawing, the more obvious the line. The line thickness is calculated proportionally according to the normalized intensity value to further enhance the hand-drawn feel.

[0121] In the above manner, a large number of training samples simulating real hand-drawn editing scenarios can be generated. Specifically, for each key region, multiple training samples with different background versions are generated to ensure that the generated samples have a good distribution in terms of background diversity, stroke style, edge quality, etc., and a complete training triple containing images, strokes, and semantic labels can be constructed, as well as supporting data augmentation and sample balancing in subsequent model training.

[0122] Among them, Open CV: Open Source Computer Vision Library. Add Weighted: A weighted summation function for image blending. RGBA: Red, Green, Blue, Alpha color space. Alpha channel: The transparency channel used to control the transparency of an image. cv2.Gaussian Blur: A Gaussian blur function for smoothing images. NumPy: A powerful Python library for scientific computing. scipy.interpolate: The interpolation module in the SciPy library for data smoothing or curve fitting. Bezier curve: A curve commonly used in computer graphics for drawing smooth curves. cv2.line: A function in Open CV for drawing a straight line. Line type = cv2.LINE_AA: Anti-aliased line type. Min Max Scaler: A normalization method that scales data between a given minimum and maximum value. Linear interpolation: A simple and commonly used interpolation method for estimating unknown values between two known data points.

[0123] In another possible implementation, it is also possible not to perform adjustment processing on the sample images, which can reduce the computing power resources required for model training.

[0124] In one possible implementation, the method further includes: For each third-region image, determine the density value of the edge pixels of the third-region image, where the density value is used to indicate the ratio of the number of edge pixels to the total number of pixels of the third-region image; based on the density values corresponding to each third-region image, determine the third target-region image among the L third-region images, and the density value corresponding to the third target-region image is not less than that of the images other than the third target-region image among the L third-region images.

[0125] In this embodiment, for each region in the image, the density of edge pixels is calculated to obtain an edge density distribution map. Based on the edge density distribution, high-density regions containing rich structural details are selected. These regions usually correspond to the main objects or scene elements in the image. One or more region images with the highest edge density are selected from each image as the key region images (also known as the third target region images), and the semantic labels (also known as sample labels) corresponding to each key region image are extracted. The semantic labels can be, for example, "building", "car", "person", etc. Optionally, the labels can be cleaned by only retaining the core noun components and removing modifiers and redundant descriptions to ensure the simplicity and consistency of the labels. Exemplarily, the density value in this embodiment = the number of edge pixels in the region / the total number of pixels in the region.

[0126] In this embodiment, for each third region image, the density value of the edge pixels of the third region image is determined. The density value is used to: indicate the ratio of the number of edge pixels to the total number of pixels of the third region image; based on the density values corresponding to the third region images, the third target region image among the L third region images is determined. The density value corresponding to the third target region image is not less than the images among the L third region images other than the third target region image, that is, high-density regions containing rich structural details are selected. These regions usually correspond to the main objects or scene elements in the image, and thus the accuracy and effectiveness of the extracted object can be improved, which is beneficial to improving the accuracy of sample image generation.

[0127] It should be understood that the application scenarios for the above model to perform intent recognition can include, but are not limited to, image processing, such as image information extraction, image editing, etc. Below, an application where intent recognition is applied to image editing will be described by way of example.

[0128] In a possible implementation manner, an embodiment of the present invention provides a method for image editing. A graphical user interface is provided through a terminal device. The terminal device can be the aforementioned local terminal device or the client device in the aforementioned cloud interaction system. Through the terminal device, a graphical user interface is provided. The content displayed on the graphical user interface can be based on the type of the launched application program. For example, a game scene screen, a communication interaction window, an image editing window, etc. In this embodiment, an image editing window is displayed in the graphical user interface. The image editing window can be a window for implementing image editing, such as a window of an image editing application program or a window related to the image editing function in a game interface.

[0129] Please refer to Figure 3 , Figure 3Schematic flowchart of an image editing method provided by an embodiment of the present invention. In this embodiment, a graphical user interface can be provided through a terminal device; a first image is displayed in the graphical user interface; as Figure 3 shown, the method may include: S310. Respond to a first sliding operation on the first image, and obtain an intention recognition result based on the first sliding operation information corresponding to the first sliding operation and the first image.

[0130] The intention recognition result is obtained based on the method of the above embodiment and will not be elaborated here.

[0131] S320. Generate an image editing instruction based on the intention recognition result, where the image editing instruction is used to: indicate editing of the first image.

[0132] Image editing is a part of image processing, involving various operations on an image to change its appearance or content. In this embodiment, image editing can be operations such as various transformations, restorations, enhancements, or compositions on the image.

[0133] S330. Edit the first image based on the image editing instruction to obtain a second image.

[0134] In this embodiment, the user can perform a first sliding operation on the first image. Then, the client device can respond to the first sliding operation on the first image, obtain an intention recognition result based on the first sliding operation information corresponding to the first sliding operation and the first image, then generate an image editing instruction based on the intention recognition result, and further edit the first image based on the image editing instruction to obtain a second image. In this way, the user can express the editing intention through the sliding operation, and then edit the first image to obtain the second image. This enables the user to express the editing intention of the image through the sliding operation on the image, and then edit the image according to the user's editing intention, thereby improving the user experience.

[0135] In a possible implementation manner, the first image includes a first object. Editing the first image based on the image editing instruction to obtain a second image includes: Overlaying a second object on the first object based on the image editing instruction, where the second image includes the first object and the second object overlaid on the first object; or, Editing the size of the first object based on the image editing instruction, where the second image includes the first object with the size edited; or, Editing the shape of the first object based on the image editing instruction, where the second image includes the first object with the shape edited.

[0136] Among them, superimposing a second object on a first object can be, for example, superimposing a ring on a person's finger or a sun in the sky. Editing the size of the first object can be, for example, editing the length of a skirt. Editing the shape of the first object can be, for example, changing the shape of the skirt, which can be understood as changing the style of the skirt, etc., and there is no limitation here.

[0137] In a possible implementation manner, after editing a first image based on an image editing instruction to obtain a second image, the method further includes: In response to a re - editing instruction, updating the second image to a third image, where the re - editing instruction is used to: indicate re - editing the first image, and the third image is obtained by editing the first image based on the re - editing instruction.

[0138] In this embodiment, if the user is not satisfied with the editing result of the first image, for example, the user believes that the editing result is not what they want, they can input a re - editing instruction, so that the client device, in response to the re - editing instruction, edits the first image based on the re - editing instruction to obtain a third image, and thus updates the second image to the third image for display. In this way, the user can re - edit when they are not satisfied with the editing result, which can improve the user experience.

[0139] In a possible implementation manner, the user can input a text prompt as the re - editing instruction.

[0140] In a possible implementation manner, the image editing instruction includes a first image editing instruction and a second image editing instruction. The second image is obtained by editing the first image based on the first image editing instruction. In response to the re - editing instruction, updating the second image to a third image includes: Displaying a first editing identifier and a second editing identifier through a graphical user interface. The first editing identifier is used to: indicate the result of editing the first image according to the first image editing instruction; the second editing identifier is used to: indicate the result of editing the first image according to the second image editing instruction; In response to a first trigger operation on the second editing identifier, updating the second image to a third image, where the third image is obtained by editing the first image based on the second image editing instruction.

[0141] Among them, the first image editing instruction can be generated based on a first intention recognition result, and the second image editing instruction can be generated based on a second intention recognition result. Optionally, if the confidence level of the first intention recognition result is higher than that of the second intention recognition result, the first image is edited according to the editing instruction obtained from the intention recognition result with the highest confidence level to obtain the second image. The first trigger operation can be, for example, a click operation or a long - press operation, etc., and there is no limitation here.

[0142] In this embodiment, the first editing identifier and the second editing identifier are displayed through the graphical user interface. Then, the user can know what the editing result of the current second image is through the first editing identifier. Next, if the user is not satisfied and is satisfied with the editing result corresponding to the second editing identifier, the second image can be updated to the third image through the first triggering operation on the second editing identifier.

[0143] In a possible implementation manner, the first editing identifier includes a first thumbnail, which is a thumbnail of the first image edited according to the first image editing instruction, and the second editing identifier includes a second thumbnail, which is a thumbnail of the first image edited according to the second image editing instruction.

[0144] Among them, a thumbnail is a smaller picture, usually a reduced version of the original picture, used to quickly preview the image content in a file browser, web page or application without loading or viewing the complete picture file. In this embodiment, the editing result of the image is represented by a thumbnail. This embodiment can use the thumbnail as the editing identifier, so that the user can quickly judge which thumbnail is satisfactory through the thumbnail, which can improve the efficiency of the user to update the image editing result.

[0145] In another possible implementation manner, the text corresponding to the editing result can also be used as the editing identifier, which can improve the simplicity of the interface display.

[0146] In a possible implementation manner, before generating the image editing instruction based on the intention recognition result, the method further includes: Displaying an editing parameter identifier through the graphical user interface, where the editing parameter identifier is used to: indicate the editing parameters used to edit the first image, and the editing parameters include at least one of color, transparency, line type or line width; receiving a second triggering operation on the editing parameter identifier.

[0147] Correspondingly, generating an image editing instruction based on the intention recognition result includes: Generating an image editing instruction based on the editing parameters indicated by the editing parameter identifier and the intention recognition result.

[0148] Among them, the description of the second triggering operation can refer to the description of the first triggering operation and will not be elaborated here. Exemplarily, in this embodiment, if the editing parameter includes color, then when the user makes a sliding operation to draw a circle on the person in the first image, the intention recognition result can be: draw a ring of the corresponding color on the finger of the person. For example, if the color is silver, the intention recognition result can be: draw a silver ring on the finger of the person.

[0149] If the editing parameter includes transparency, when the user makes a sliding operation corresponding to drawing a circle on a person in the first image, the intention recognition result can be: drawing a ring corresponding to this transparency on the finger of the person.

[0150] If the editing parameter includes line width, when the user makes a sliding operation corresponding to drawing a circle on a person in the first image, the intention recognition result can be: drawing a ring corresponding to this line width on the finger of the person. For example, the larger the line width, the larger the width of the ring.

[0151] If the editing parameter includes line type, when the user makes a sliding operation corresponding to drawing a circle on a person in the first image, the intention recognition result can be: drawing a ring corresponding to this line type on the finger of the person. For example, if the line types are different, the styles of the rings are also different. For example, the pattern of the ring is this line type. The line type can include but is not limited to solid lines, dotted lines, etc., or straight lines, wavy lines, spiral lines, etc.

[0152] In this embodiment, an image editing instruction can be generated based on the editing parameter indicated by the editing parameter identifier. That is to say, the result of image editing based on the image editing instruction also takes into account the editing parameter. In this way, the result of image editing is more in line with the user's intention, further improving the user experience.

[0153] In another possible implementation manner, the intention recognition result output by the target generative language model can take into account the editing parameter. In this way, the intention recognition result can be used as an image editing instruction. Since the target generative language model directly outputs the image editing instruction, the efficiency of image editing can be improved.

[0154] Generally speaking, this embodiment can display the recognition result immediately. For example, the preliminary recognition result is given within 100 milliseconds after the user finishes drawing. And multiple candidate solutions are provided. Optionally, the top three results with confidence levels can be displayed: such as "cat (90%)", "kitten (85%)", "pet (75%)". In addition, manual fine-tuning and correction are supported. For example, the user can select a candidate result through a drop-down menu or directly edit the text box to modify the recognition result. In addition, a mechanism for feedback and optimization can be added. Specifically, record the correction history of the user's image editing, that is, record the results of the user's secondary editing (re-editing) of the image and the results of the user's primary editing (editing before re-editing) of the image for model optimization. For example, the user often corrects "dog" to "Shiba Inu", etc. In this embodiment, the prediction probabilities of different intent categories can be calculated to obtain multiple intent recognition results, and a dynamic threshold can be set for screening the intent recognition results. In addition, for the intent recognition results with low confidence, the image needs to be edited according to the image editing instructions corresponding to the intent recognition result only after the user's secondary confirmation.

[0155] Exemplarily, an example is given for modifying the clothing style based on the hand-drawn intent.

[0156] 1. Initial scenario.

[0157] Input image: A photo of a girl wearing a short skirt (size: 1024x1024), user intent: change the short skirt to a long red dress, model configuration: use LLaVA-1.5-13B as the base model and adopt LoRA fine-tuning (rank = 16).

[0158] 2. User interaction process.

[0159] 2.1. The user selects a red brush (RGB: 255, 0, 0) in the image editing interface and draws the outline of the long dress extending downward in the area of the short skirt. 2.2. The system records the stroke information: stroke type: Add Brush. 2.3. Normalized coordinates: starting position: (0.45, 0.48) / / the position of the hem of the short skirt, ending position: (0.45, 0.85) / / extending to the ankle position. Stroke attributes: color: RGB(255, 0, 0), transparency: 0.7, stroke width: 5px.

[0160] 3. System processing flow.

[0161] 3.1. Data preprocessing: Original image processing: Scale the 1024x1024 image to the standard input size (512x512).

[0162] 3.2. Perform normalization: pixel_values = (image - mean) / std. Here, this formula is used for normalizing the image, aiming to convert the pixel values of the image into a standard distribution form for subsequent image processing or machine learning model processing. Here, image represents the original image data, and mean and std represent the average value and standard deviation of the image pixel values respectively.

[0163] 3.3. Stroke processing: Generate a stroke mask (M): Convert the user-drawn trajectory into a binary mask. In this embodiment, after stroke processing, the pixels drawn by the user can be marked.

[0164] 3.4. Extract the bounding box: bbox = [0.45, 0.48, 0.45, 0.85]. 3.5. Construct the model input. Construct an inference template: This is a "you guess what I draw" game. I will upload an image containing strokes. To help locate the strokes, the normalized bounding box coordinates of the strokes are: - Upper left corner: (0.45, 0.48) - Lower right corner: (0.45, 0.85). Please describe the content expressed by the strokes with a word or phrase.

[0165] 4. Model inference process.

[0166] 4.1. Model output: Long dress.

[0167] 4.2. Result processing.

[0168] 4.3. The system combines the model output with the color information selected by the user: Red long dress / / Composed of "long dress" and RGB(255, 0, 0).

[0169] 4.4. Pass the combined prompt word to the image generation model for local redrawing.

[0170] This embodiment shows how to accurately understand the user's hand-drawn intention through the trained MLLM model, only output the core object name "long dress", and then combine it with the color attribute selected by the user to form a complete editing instruction. The whole process is simple and direct, without complex semantic understanding and description generation.

[0171] For ease of understanding, the following embodiment exemplarily illustrates the image editing process with an exemplary interface display.

[0172] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the interface display for image editing provided by the embodiment of the present invention. As Figure 4In the interface shown in (a), a first image and editing parameter identifiers are displayed. Among them, the first image can include, for example, image elements such as the sky and clouds. The editing parameter identifiers are located below the first image. Optionally, the editing parameter identifiers can include but are not limited to color identifiers, line type identifiers, line width identifiers, transparency identifiers, etc. Among them, the color identifier indicates the color to be edited, the line type identifier indicates the type of the line to be edited, the line width identifier indicates the width of the line to be edited, and the transparency identifier indicates the transparency of the line to be edited.

[0173] Then, as Figure 4 shown in (b), the terminal device can detect a trigger operation on the color identifier, and then the line to be edited later is drawn in the color corresponding to the color identifier. For example, the color identifier can be a white identifier or a yellow identifier, which is not limited here. Exemplarily, the sun generated along the trajectory corresponding to the white identifier is softer, while the sun generated along the trajectory corresponding to the yellow identifier is more dazzling.

[0174] Then, as Figure 4 shown in (c), the terminal device can detect a first sliding operation on the first image. As Figure 4 shown in (c), the first sliding operation is a sliding operation similar to a circle, then a sliding trajectory of the first sliding operation can be drawn at the position where the first sliding operation acts on the first image, and the sliding trajectory is similar to a circle.

[0175] At this time, after normalizing the coordinate information of the upper left corner and the lower right corner of the target area where the sliding trajectory is located, the terminal device inputs the normalized coordinate information of the upper left corner, the normalized coordinate information of the lower right corner, and the first image into the target generative language model. Then, the target generative language model can use the normalized coordinate information of the upper left corner, the normalized coordinate information of the lower right corner, and the canvas to restore the coordinates of the target area, and then extract the sliding trajectory and the regional image of the first image in the target area based on the coordinates of the target area. Furthermore, the target generative language model outputs a first intention recognition result and a second intention recognition result by using the sliding trajectory and the regional image of the first image in the target area, where the confidence level of the first intention recognition result is higher than that of the second intention recognition result. Exemplarily, the first intention recognition result can be to draw a sun, and the second intention recognition result can be, for example, to draw a moon. Then, the terminal device can generate a first image editing instruction based on the first intention recognition result and a second image editing instruction based on the second intention recognition result.

[0176] It should be understood that for how to output the intention recognition result, reference can be made to the description of the above embodiments, which will not be elaborated here.

[0177] In another possible implementation, the target generative language model can also directly output a first image editing instruction corresponding to the first intention recognition result and a second image editing instruction corresponding to the second intention recognition result, thereby directly editing the first image.

[0178] It should be noted that the target generative language model can be deployed on a terminal device or on a server, and there is no limitation here.

[0179] Then, as shown in (d) of Figure 4 , the terminal device can edit the first image based on the first image editing instruction to obtain a second image. In the second image, a sun image element is added to the sky. And, as shown in the upper right corner of the interface in (d) of Figure 4 , a first thumbnail and a second thumbnail are also displayed. Among them, the first thumbnail is the thumbnail corresponding to the second image, and the second thumbnail is the thumbnail corresponding to the image obtained by editing the first image based on the second image editing instruction. The first thumbnail is in a selected state, representing that the currently displayed image is the image enlarged from the first thumbnail.

[0180] Then, if the user is not satisfied and believes that the second thumbnail is the editing result they want, then as shown in (e) of Figure 4 , when the terminal device detects a trigger operation on the second thumbnail, it can display the interface shown in (f) of Figure 4 .

[0181] In the interface shown in (f) of Figure 4 , the second thumbnail is in a selected state, and the image shown in (f) of Figure 4 is a third image. In the third image, a moon image element is added to the sky. The third image is the image enlarged from the second thumbnail.

[0182] It should be noted that the operation in (b) of Figure 4 may not be required either.

[0183] In another possible implementation, as shown in (c) of Figure 4 , the sliding trajectory may not be displayed either. The terminal device can input consecutive multiple position information of the sliding operation into the target generative language model.

[0184] Please refer to Figure 5 , Figure 5 which is another schematic diagram of the interface display for image editing provided by an embodiment of the present invention. As shown in Figure 5As shown in (a) therein, a first image and an editing parameter identifier are displayed. Among them, the first image may, for example, include image elements such as grassland, trees, and sky. The editing parameter identifier is located below the first image. And the terminal device can detect a first sliding operation acting on the first area. Then, the coordinate information of the target area covering the sliding trajectory of the first sliding operation and the first image can be input into the target generative language model, and the target generative language model can output an image editing instruction, which instructs to add a football as this image element.

[0185] Then, based on the image editing instruction, the first image can be edited to obtain a second image as shown in (b) therein. Figure 5 In the second image, a football is added as this image element.

[0186] By comparing Figure 4 and Figure 5 it can be found that although the first sliding operation is the same, when acting on different images, the edited results are different. That is to say, the edited result is related to the image on which the first sliding operation acts. In this way, the first image can be edited according to the information of the first image, and the edited result is also more in line with the scene of the first image, thereby improving the user experience.

[0187] Generally speaking, this embodiment provides a method for real-time recognition of image local redrawing intention based on a multimodal large language model. By constructing training data simulating a hand-drawn scene and using a multimodal large language model for fine-tuning, the accurate understanding of the hand-drawn intention can be achieved. The end user can convey the intention to the model in real time in a hand-drawn way, and the model will automatically generate corresponding editing instructions, thereby realizing efficient image editing. Specifically, it involves the following content: 1. Training dataset construction module.

[0188] Extract and construct training samples simulating a hand-drawn scene from a large-scale image database. Use an edge detection algorithm to extract image structure features. Generate a simulated hand-drawn effect through image restoration and edge superposition techniques. Establish a training dataset containing diverse scenes and rich labels.

[0189] 2. MLLM model fine-tuning module.

[0190] Select a suitable multimodal large language model as the basic architecture. Adopt an efficient parameter fine-tuning strategy for model optimization. Enhance the model's understanding and semantic mapping ability of hand-drawn content. Establish an accurate conversion from visual features to semantic descriptions.

[0191] 3. Real-time intention recognition module.

[0192] Design a dedicated intent understanding task framework. Support intent capture in multiple interaction modes, including coloring, adding content brushes, etc. A context-aware intelligent reasoning mechanism. Optimize the reasoning performance to achieve real-time response.

[0193] 4. Model inference process.

[0194] After the model is trained, it can be used for inference and application. In actual applications, usually, three modules of hand-drawn stroke capture, coordinate normalization, and inference output processing need to be connected in series to form a complete processing flow. First, it is necessary to capture the user's hand-drawn strokes. Generally, an interactive canvas is provided for the user to draw on. The user can use a mouse or a stylus to draw on the canvas. For each stroke drawn, information such as its coordinates, color, and transparency needs to be captured. In addition, the content drawn by the user is based on the original image, so the image transmitted to the model includes the original image and the strokes superimposed on the original image. In this way, the model can understand which part of the original image the user's drawing is based on, thereby enhancing the model's understanding of the user's intent.

[0195] Next, the effects of this embodiment will be described uniformly.

[0196] Technical effects: Improve the accuracy of hand-drawn intent recognition in local redrawing scenarios. And by constructing high-quality simulated hand-drawn training data, significantly improve the model's recognition ability for incomplete, deformed, or abstract hand-drawn content. And utilize the semantic understanding ability of the multi-modal large language model to accurately parse the user's random scribbles and simple lines. And based on the training scheme of edge map extraction and region annotation, enable the model to understand diverse stroke styles. And support the recognition of multiple interaction intents, including: Inserting new elements: adding new objects through simple contour lines, Color modification: indicating the area that needs to change color through coloring strokes. Reduce recognition latency: adopt parameter-efficient fine-tuning technology to optimize the model's inference performance. Support an instant feedback mechanism to achieve the continuity and smoothness of the editing process. Achieve real-time intent prediction during the stroke process: continuously analyze stroke features during the user's drawing process, and update the intent prediction results in real time according to the stroke trajectory; control the display timing of the prediction results through a confidence threshold. Enhance the context understanding ability, combine the overall context of the image for intent understanding, avoid generating content that is inconsistent with the scene, and accurately grasp the relevance between the editing area and the surrounding content. And, through the semantic mapping ability of the multi-modal model, understand the spatial and semantic relationships between objects in the image. Intelligently understand complex editing intents: analyze the spatial relationship between strokes and existing content, understand the combined intent of the user's consecutive multiple strokes, infer the most reasonable editing operation according to the scene context, and automatically adjust the generated content to maintain the overall coordination of the image.

[0197] Application effects: Improve the user's editing efficiency. And there is no need to manually input text prompts, and the intention can be directly expressed by hand drawing. And the user's intention is recognized in real time to avoid repeated adjustments and waiting. And it supports batch processing of redrawing tasks for multiple local areas. And reduce the number of repeated operations through intelligent understanding. Improve the interaction experience: maintain a natural creative process without interrupting the user's train of thought; preview the recognition results in real time and correct incorrect understandings in a timely manner; support multiple hand-drawing styles to adapt to different user habits; the generated content is automatically coordinated with the scene to reduce post-adjustment. Lower the usage threshold: there is no need to master complex prompt writing skills, and simple doodles can express the editing intention. It can intelligently understand incomplete or rough hand-drawn content and automatically optimize the generated results, reducing the professional skill requirements. Improve the local redrawing effect: accurately understand the area range that the user wants to modify, intelligently maintain the visual coherence with the surrounding content, automatically adjust the lighting and texture details of the generated content, and support multiple iterations of optimization until the ideal effect is achieved.

[0198] The method embodiments are described above, and the following embodiments describe the product embodiments.

[0199] Please refer to Figure Figure 6 , Figure 6 which is a schematic structural diagram of an intention recognition device provided by an embodiment of the present invention. As Figure 6 shown, the device may include: A first acquisition module 610, configured to acquire first sliding operation information and a first image, where the first sliding operation information is used to: indicate the sliding path of the first sliding operation acting on the first image; an intention recognition module 620, configured to input the first sliding operation information and the first image into a trained target generative language model, where the target generative language model is used to: output an intention recognition result based on the first sliding operation information and the first image, and the intention recognition result is used to: indicate the intention of the first sliding operation acting on the first image; obtain the intention recognition result output by the target generative language model.

[0200] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an image editing device provided by an embodiment of the present invention. As Figure 7 shown, the device may include: A response module 710, configured to respond to a first sliding operation acting on the first image, and acquire an intention recognition result based on the first sliding operation information corresponding to the first sliding operation and the first image, where the intention recognition result is obtained based on the above method; an image editing module 720, configured to generate an image editing instruction based on the intention recognition result, where the image editing instruction is used to: indicate editing of the first image; edit the first image based on the image editing instruction to obtain a second image.

[0201] Please refer to Figure 8 ,Figure 8 The structural schematic diagram of a model training device provided by an embodiment of the present invention. As Figure 8 shown, the device may include: A second acquisition module 810, configured to acquire a pre-trained initial generative language model; acquire a training sample set, the training sample set includes a plurality of training samples, each training sample includes a sample image and a sample label corresponding to the sample image, the sample image includes a fourth image and a target line superimposed on the fourth image, the target line is used for: representing a second sliding operation applied to the fourth image, and the sample label is used for: indicating the intention of processing the fourth image; a model training module 820, configured to fine-tune the initial generative language model by using the training sample set to obtain a trained target generative language model, and the target generative language model is used to identify the intention of a first sliding operation applied to a first image.

[0202] For the device of this embodiment, reference may be made to the description of the above method embodiment, and details are not described herein again.

[0203] This embodiment also provides an electronic device, including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above image editing method. The electronic device may be a server or a terminal device.

[0204] See Figure 9 As shown, the electronic device includes a processor 100 and a memory 101, the memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the steps of the above method.

[0205] Further, Figure 9 the electronic device shown further includes a bus 102 and a communication interface 103, and the processor 100, the communication interface 103 and the memory 101 are connected through the bus 102.

[0206] Wherein, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which may be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 may be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 9 only a bidirectional arrow is used in

[0207] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0208] This embodiment also provides a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the steps of the above method.

[0209] This embodiment also provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above method.

[0210] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0211] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0212] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0213] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0214] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for intention recognition, characterized in that: include: Acquire first sliding operation information and a first image, wherein the first sliding operation information is used to: indicate a sliding path of the first sliding operation on the first image; Inputting the first sliding operation information and the first image into a trained target generative language model, wherein the target generative language model is used to: output an intention recognition result based on the first sliding operation information and the first image, wherein the intention recognition result is used to: indicate the intention of the first sliding operation to act on the first image; Obtain the intent recognition result output by the target generative language model.

2. The method according to claim 1, characterized in that The first image input into the target generative language model is superimposed with the strokes of the sliding path, the first sliding operation information includes a plurality of position information, and the plurality of position information is at least used to indicate a target area formed by the sliding path, and the first sliding operation information and the first image are input into the trained target generative language model, comprising: The multiple position information and the first image are input into a trained target generative language model, and the target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the multiple position information, and output the intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

3. The method according to claim 2, characterized in that The multiple position information includes first position information of the target area and second position information of the target area, a distance between the first position information and the second position information is not less than a distance between any two points in the target area, and inputting the multiple position information and the first image into a trained target generative language model includes: The first position information, the second position information and the first image are input into a trained target generative language model, and the target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the first position information and the second position information, and output the intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

4. The method according to claim 3, characterized in that The first sliding operation is carried by a canvas, the canvas includes a first edge and a second edge, the first position information includes a first coordinate value in a first direction and a second coordinate value in a second direction, the second position information includes a third coordinate value in the first direction and a fourth coordinate value in the second direction, the first direction is parallel to the first edge and the second direction is parallel to the second edge, and before inputting the first position information, the second position information and the first image into a trained target generative language model, the method further includes: Normalizing the first position information and the second position information respectively: normalizing the first coordinate value and the third coordinate value respectively based on the length of the first edge, and normalizing the second coordinate value and the fourth coordinate value respectively based on the length of the second edge; The step of inputting the first position information, the second position information, and the first image into a trained target generative language model includes: The normalized first position information, the normalized second position information and the first image are input into a trained target generative language model, and the target generative language model is used to: obtain the strokes of the sliding path and the regional image of the first image in the target area based on the normalized first position information and the normalized second position information, and output the intention recognition result based on the strokes of the sliding path and the regional image of the first image in the target area.

5. The method according to claim 1, characterized in that Before inputting the first sliding operation information and the first image into a trained target generative language model, the method further includes: Performing image segmentation on the first image to obtain M first region images, each of which includes a first object, where M is a positive integer; The step of inputting the first sliding operation information and the first image into a trained target generative language model includes: The first sliding operation information and the first target area image among the M first area images are input into a trained target generative language model, where the first target area image includes the first area image acted upon by the first sliding operation. The target generative language model is used to output an intention recognition result based on the first sliding operation information and the first target area image among the M first area images.

6. The method according to any one of claims 1 to 5, characterized in that Inputting the first sliding operation information and the first image into a trained target generative language model includes: The first sliding operation information is embedded in a prompt information template to obtain prompt information, and the prompt information and the first image are input into a trained target generative language model, wherein the target generative language model is used to output an intent recognition result based on the prompt information, the first sliding operation information, and the first image, and the prompt information template is used to guide the target generative language model to output an intent recognition result, and the target generative language model is used to output an intent recognition result based on the prompt information, the first sliding operation information, and the first image.

7. The method according to claim 6, characterized in that The step of embedding the first sliding operation information into a prompt information template to obtain prompt information, and inputting the prompt information and the first image into a trained target generative language model includes: embedding the first sliding operation information into a first prompt information template to obtain first prompt information, inputting the first prompt information and the first image into a trained target generative language model, wherein the target generative language model is used to: output an intention recognition result based on the first prompt information and the first image, and the first prompt information template is used to: indicate that the intention recognition result output by the target generative language model involves an editing parameter selected when the first sliding operation acts on the first image, and the editing parameter includes at least one of color, transparency, line type, or line width; or, The first sliding operation information is embedded in a second prompt information template to obtain second prompt information, and the second prompt information and the first image are input into a trained target generative language model. The target generative language model is used to: output an intention recognition result based on the second prompt information and the first image, and the second prompt information template is used to: indicate an editing result of the first sliding operation acting on the first image.

8. An image editing method, characterized in that: Providing a graphical user interface via a terminal device; The graphical user interface displays a first image; the method comprises: In response to a first sliding operation applied to the first image, acquiring an intention recognition result based on first sliding operation information corresponding to the first sliding operation and the first image, wherein the intention recognition result is obtained based on the method according to any one of claims 1 to 7; Generate an image editing instruction based on the intention recognition result, wherein the image editing instruction is used to: instruct to edit the first image; The first image is edited based on the image editing instruction to obtain a second image.

9. The method according to claim 8, characterized in that The first image includes a first object, and the editing the first image based on the image editing instruction to obtain a second image includes: superimposing a second object in the first object based on the image editing instruction, the second image including the first object and the second object superimposed on the first object; or, The size of the first object is edited based on the image editing instruction, and the second image includes the first object after the size is edited; or The shape of the first object is edited based on the image editing instruction, and the second image includes the first object after the shape is edited.

10. The method according to claim 8, characterized in that After editing the first image based on the image editing instruction to obtain the second image, the method further includes: In response to a re-editing instruction, the second image is updated to a third image, wherein the re-editing instruction is used to instruct re-editing of the first image, and the third image is obtained by editing the first image based on the re-editing instruction.

11. The method according to claim 10, characterized in that The image editing instruction includes a first image editing instruction and a second image editing instruction, the second image is obtained by editing the first image based on the first image editing instruction, and in response to the re-editing instruction, updating the second image to a third image includes: Displaying a first editing mark and a second editing mark through the graphical user interface, wherein the first editing mark is used to indicate the result of editing the first image according to the first image editing instruction; and the second editing mark is used to indicate the result of editing the first image according to the second image editing instruction; In response to a first trigger operation acting on the second editing identifier, the second image is updated to a third image, where the third image is obtained by editing the first image based on the second image editing instruction.

12. The method according to claim 11, characterized in that The first editing mark includes a first thumbnail, which is a thumbnail of the first image edited according to the first image editing instruction. The second editing mark includes a second thumbnail, which is a thumbnail of the first image edited according to the second image editing instruction.

13. The method according to any one of claims 8 to 12, characterized in that: Before generating the image editing instruction based on the intention recognition result, the method further includes: Displaying an editing parameter identifier in the graphical user interface, the editing parameter identifier is used to: indicate an editing parameter used to edit the first image, the editing parameter including at least one of color, transparency, line type or line width; receiving a second trigger operation acting on the editing parameter identifier; The generating the image editing instruction based on the intention recognition result includes: An image editing instruction is generated based on the editing parameter indicated by the editing parameter identifier and the intention recognition result.

14. A model training method, characterized in that: include: Get the pre-trained initial generative language model; Acquire a training sample set, the training sample set comprising a plurality of training samples, each training sample comprising a sample image and a sample label corresponding to the sample image, the sample image comprising a fourth image and a target line superimposed on the fourth image, the target line being used to indicate a second sliding operation acting on the fourth image, and the sample label being used to indicate an intention to process the fourth image; The initial generative language model is fine-tuned using the training sample set to obtain a trained target generative language model, where the target generative language model is used to identify the intention of the first sliding operation acting on the first image.

15. The method according to claim 14, characterized in that Before fine-tuning the initial generative language model using the training sample set to obtain a trained target generative language model, the method further includes: Performing image segmentation on the sample image to obtain N second region images, each of which includes a third object, where N is a positive integer; The step of fine-tuning the initial generative language model using the training sample set to obtain a trained target generative language model includes: The initial generative language model is fine-tuned using a second target area image in a sample image corresponding to the training sample and a sample label corresponding to the sample image, wherein the second target area image includes a second area image covered by the target line.

16. The method according to claim 14 or 15, characterized in that The sample image is obtained by: acquiring a fifth image; Performing image segmentation on the fifth image to obtain L third region images, each of which includes the fourth object, where L is a positive integer; For at least one third target area image among the L third area images, acquiring a contour of the fourth object in the third target area image; The contour of the fourth object is superimposed on a plurality of different fourth images to obtain a plurality of sample images.

17. The method according to claim 16, characterized in that After superimposing the contour of the fourth object onto a plurality of different fourth images to obtain a plurality of sample images, the method further includes: The sample image is adjusted to fine-tune the initial generative language model using the adjusted sample image, wherein the adjustment includes at least one of blending the fourth image with the contour with transparency or randomly perturbing the lines of the contour.

18. The method according to claim 16, characterized in that The method further comprises: For each third region image, determining a density value of edge pixels of the third region image, wherein the density value is used to: indicate a ratio of the number of edge pixels to the total number of pixels of the third region image; The third target area image among the L third area images is determined based on the density values ​​corresponding to each third area image, and the density value corresponding to the third target area image is not less than that of images other than the third target area image among the L third area images.

19. An intention recognition device, characterized in that: include: A first acquisition module, configured to acquire first sliding operation information and a first image, wherein the first sliding operation information is used to indicate a sliding path of the first sliding operation on the first image; an intention recognition module, configured to input the first sliding operation information and the first image into a trained target generative language model, wherein the target generative language model is configured to: output an intention recognition result based on the first sliding operation information and the first image, wherein the intention recognition result is configured to: indicate an intention of the first sliding operation to act on the first image; Obtain the intent recognition result output by the target generative language model.

20. An image editing device, characterized in that: Providing a graphical user interface via a terminal device; The graphical user interface displays a first image; the device comprises: a response module, configured to respond to a first sliding operation on the first image, and obtain an intention recognition result based on first sliding operation information corresponding to the first sliding operation and the first image, wherein the intention recognition result is obtained based on the method according to any one of claims 1 to 7; An image editing module is used to generate an image editing instruction based on the intention recognition result, and the image editing instruction is used to: instruct to edit the first image; edit the first image based on the image editing instruction to obtain a second image.

21. A model training device, characterized in that: include: a second acquisition module, configured to acquire a pre-trained initial generative language model; acquire a training sample set, wherein the training sample set includes a plurality of training samples, each training sample includes a sample image and a sample label corresponding to the sample image, the sample image includes a fourth image and a target line superimposed on the fourth image, the target line is used to indicate a second sliding operation acting on the fourth image, and the sample label is used to indicate an intention to process the fourth image; A model training module is used to fine-tune the initial generative language model using the training sample set to obtain a trained target generative language model, wherein the target generative language model is used to identify the intention of the first sliding operation to act on the first image.

22. An electronic device, characterized in that: It includes a processor and a memory, the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the intention recognition method described in any one of claims 1-7, the image editing method described in any one of claims 8-13, or the model training method described in any one of claims 14-18.

23. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions prompt the processor to implement the intention recognition method described in any one of claims 1-7, the image editing method described in any one of claims 8-13, or the model training method described in any one of claims 14-18.

24. A computer program product, characterized in that The computer program product includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer device, the computer device executes the intention recognition method described in any one of claims 1 to 7, the image editing method described in any one of claims 8 to 13, or the model training method described in any one of claims 14 to 18.

Citation Information

Cited By

  • Multi-modal content generation method and device, intelligent agent and electronic equipment

    CN121541807A